PubMed HealthSearch

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Reliability of transient-evoked otoacoustic emissions.

OBJECTIVE: This investigation addressed four factors affecting transient-evoked otoacoustic emission (TEOAE) reliability: 1) The effect of evoking-stimulus level, 2) the effect of analyzing bandwidth, 3) the effect of slight-mild hearing loss, and 4) the effect of variability in the stimulus spectrum. DESIGN: TEOAEs at 80, 74, 68, and 62 dB pSPL evoking-stimulus levels were measured in 25 ears spanning a range of hearing levels from normal to mild hearing loss for a minimum of 10 test sessions. Reliability was assessed for 1/6-, 1/3-, 1/2-, and 1-octave analyzing bandwidths. RESULTS: Evoking-stimulus level, hearing loss, and center frequency did not significantly affect reliability. With decreasing analyzing bandwidth, reliability decreased. Intrasubject test-retest standard deviations were 1.2 dB for a broadband analyzing bandwidth and 1.4, 1.5, 1.6, and 1.8 dB for 1-, 1/2-, 1/3-, and 1/6-octave analyzing bandwidths, respectively. Stimulus variability within narrower bandwidths was of sufficient magnitude to influence test-retest reliability, and attempts to correct for the variations in stimulus spectrum were unsuccessful. Slopes of the input-output functions differed across frequencies, with shallower slopes at higher frequencies. CONCLUSIONS: In general, TEOAE amplitude is highly reliable. For those individuals in this study who were more variable, the variability was at low frequencies or across the entire frequency spectrum. For clinical applications, the choice of analyzing bandwidth should be based on consideration of both frequency specificity (where narrow analyzing bandwidths are optimal) and reliability (where wide analyzing bandwidths are optimal).

Acoustic Impedance Tests

Teaching DSM-III to clinicians. Some problems of the DSM-III system reducing reliability, using the diagnosis and classification of depressive disorders as an example.

Experiences from teaching DSM-III to more than three hundred Norwegian psychiatrists and clinical psychologists suggest that reliable DSM-III diagnoses can be achieved within a few hours training with reference to the decision trees and the diagnostic criteria only. The diagnoses provided are more reliable than the corresponding ICD diagnoses which the participants were more familiar with. The three main sources of reduced reliability of the DSM-III diagnoses are related to: poor knowledge of the criteria which often is connected with failure of obtaining diagnostic key information during the clinical interview; unfamiliar concepts and vague or ambiguous criteria. The two first issues are related to the quality of the teaching of DSM-III. The third source of reduced reliability reflects unsolved validity issues. By using the classification of five affective case stories as examples, these sources of diagnostic pitfalls, reducing reliability and ways to overcome these problems when teaching the DSM-III system, are discussed. It is concluded that the DSM-III system of classification is easy to teach and that the system is superior to other classification systems available from a reliability point of view. The current version of the DSM-III system, however, partly owes a high degree of reliability to broad and heterogeneous diagnostic categories like the concept major depression, which may have questionable validity. Thus, the future revisions of the DSM-III system should, above all, address the issue of validity.

Depressive Disorder

Action potential propagation through embryonic dorsal root ganglion cells in culture. II. Decrease of conduction reliability during repetitive stimulation.

1. The reliability of the propagation of action potentials (AP) through dorsal root ganglion (DRG) cells in embryonic slice cultures was investigated during repetitive stimulation at 1-20 Hz. Membrane potentials of DRG cells were recorded intracellularly while the axons were stimulated by an extracellular electrode. 2. In analogy to the double-pulse experiments reported previously, either one or two types of propagation failures were recorded during repetitive stimulation, depending on the cell morphology. In contrast to the double-pulse experiments, the failures appeared at longer interpulse intervals and usually only after several tens of stimuli with reliable propagation. 3. In the period with reliable propagation before the failures, a decrease in the conduction velocity and in the amplitude of the afterhyperpolarization (AHP), an increase in the total membrane conductance, and the disappearance of the action potential "shoulder" were observed. 4. The reliability of conduction during repetitive stimulation was improved by lowering the extracellular calcium concentration or by replacing the extracellular calcium by strontium. The reliability of conduction decreased by the application of cadmium, a calcium channel blocker, 4-amino pyridine, a fast potassium channel blocker, or apamin or muscarine, the blockers of calcium-dependent potassium channels. The reliability of conduction was not effected by blocking the sodium potassium pump with ouabain or by replacing extracellular sodium with lithium. 5. In the period with reliable propagation cadmium, apamin, and muscarine reduced the amplitude of the AHP. The shoulder of the action potential was more pronounced and not sensitive to repetitive stimulation when extracellular calcium was replaced by strontium. It disappeared when cadmium was applied. 6. In DRG somata changes of the intracellular Ca2+ concentration were monitored by measuring the fluorescence of the Ca2+ indicator Fluo-3 with a laser-scanning confocal microscope. During repetitive stimulation, an accumulation of intracellular calcium occurred that recovered very slowly (tens of seconds) after the AP trains. 7. Computer model simulations performed in analogy to the experimental protocols produced conduction failures during repetitive stimulation only when the calcium currents during the APs were reduced. 8. From these findings it is concluded that conduction failures during repetitive stimulation are dependent on an accumulation of intracellular calcium leading to an inactivation of calcium currents, combined with small contributions of an accumulation of extracellular potassium and a summation of slow potassium conductances.

Acetylcholine

An intervention to improve the reliability of manuscript reviews for the Journal of the American Academy of Child and Adolescent Psychiatry.

OBJECTIVE: The effects of methods used to improve the interrater reliability of reviewers' ratings of manuscripts submitted to the Journal of the American Academy of Child and Adolescent Psychiatry were studied. METHOD: Reviewers' ratings of consecutive manuscripts submitted over approximately 1 year were first analyzed; 296 pairs of ratings were studied. Intraclass correlations and confidence intervals for the correlations were computed for the two main ratings by which reviewers quantified the quality of the article: a 1-10 overall quality rating and a recommendation for acceptance or rejection with four possibilities along that continuum. Modifications were then introduced, including a multi-item rating scale and two training manuals to accompany it. Over the next year, 272 more articles were rated, and reliabilities were computed for the new scale and for the scales previously used. RESULTS: The intraclass correlation of the most reliable rating before the intervention was 0.27; the reliability of the new rating procedure was 0.43. The difference between these two was significant. The reliability for the new rating scale was in the fair to good range, and it became even better when the ratings of the two reviewers were averaged and the reliability stepped up by the Spearman-Brown formula. The new rating scale had excellent internal consistency and correlated highly with other quality ratings. CONCLUSIONS: The data confirm that the reliability of ratings of scientific articles may be improved by increasing the number of rating scale points, eliciting ratings of separate, concrete items rather than a global judgment, using training manuals, and averaging the scores of multiple reviewers.

Algorithms

A reliability study of the universal goniometer, fluid goniometer, and electrogoniometer for the measurement of ankle dorsiflexion.

This study investigated the reliability of three goniometers, the universal, fluid, and electro-goniometers, in the measurement of ankle dorsiflexion. Intra- and interobserver reliability were assessed using 10 healthy volunteers and five observers. A standardized ankle position was used to measure full range of active dorsiflexion. Intraobserver reliability was assessed using one observer over two successive occasions. Interobserver reliability was assessed among five observers over five separate occasions. A one-factor analysis of variance to examine intraobserver reliability demonstrated no significant difference between each of the devices on the two occasions. A multifactorial analysis of variance demonstrated significant differences among observers and again among devices (P < 0.001). Secondary analysis for interdevice reliability demonstrated significant differences among the three devices (p < 0.1). The study suggests that each device cannot be used reliably among observers or be used interchangeably, and clinical judgment based on angular changes of less than 10 degrees are invalid if rigid protocols are not followed.

Adult

Assessing reliability of a measure of self-rated health.

The test-retest reliability of self-rated health is analysed and compared with the reliability of health questions phrased more as well as less precisely. Differences in reliability between men and women and between age groups are also assessed. The study is based on 204 and 409 re-interviews from the 1991 Swedish Level of Living Survey and the 1989 Survey of Living Conditions respectively. The results show that the reliability of self-rated health is as good as or even better than that of most of the more specific questions. Only an indicator of high blood pressure showed significantly higher reliability. The reliability of self-rated health is good in all subgroups studied, and is even excellent among older men. It is concluded that the good overall reliability of self-rated health found in this study is in line with previous results concerning the validity of people's assessments of their general health as well as results concerning the basis upon which they make these judgements.

Adult

Interrater reliability of Alzheimer's disease diagnosis.

To determine interrater reliability of dementia diagnosis, 4 physicians experienced in the evaluation of dementia patients applied 3 sets of diagnostic criteria to each of 62 patients, based on a standardized set of medical record information. All patients had undergone similar examinations and follow-up to establish the initial clinical diagnosis (76% had autopsy). Raters were blind to the diagnosis and to follow-up information after the initial evaluation period. This paper presents interrater agreement (kappa values) for a diagnosis of Alzheimer's disease using the American Psychiatric Association diagnostic criteria from the Diagnostic and Statistical Manual (DSM-III), the National Institute of Neurological and Communicative Disorders and Stroke (NINCDS) criteria for the clinical diagnosis of Alzheimer's disease, and the Eisdorfer and Cohen Research Diagnostic Criteria (ECRDC) for primary neuronal degeneration. The NINCDS showed somewhat higher average interrater reliability (kappa = 0.64) than the DSM-III (kappa = 0.55) and considerably higher interrater reliability than the ECRDC (kappa = 0.37). One rater displayed conspicuously lower levels of interrater reliability than the other 3, especially in DSM-III and ECRDC. This study indicates that interrater reliability of DSM-III and NINCDS criteria are comparable. Documentation of interrater reliability and, if necessary, training to improve reliability is an important consideration in research where different observers are diagnosing dementing illnesses.

Aged

Reliability of the NINDS Myotatic Reflex Scale.

The assessment of deep tendon reflexes is useful for localization and diagnosis of neurologic disorders, but only a few studies have evaluated their reliability. We assessed the reliability of four neurologists, instructed in two different countries, in using the National Institute of Neurological Disorders and Stroke (NINDS) Myotatic Reflex Scale. To evaluate the role of training in using the scale, the neurologists randomly and blindly evaluated a total of 80 patients, 40 before and 40 after a training session. Inter- and intraobserver reliability were measured with kappa statistics. Our results showed substantial to near-perfect intraobserver reliability, and moderate-to-substantial interobserver reliability of the NINDS Myotatic Reflex Scale. The reproducibility was better for reflexes in the lower than in the upper extremities. Neither educational background nor the training session influenced the reliability of our results. The NINDS Myotatic Reflex Scale has sufficient reliability to be adopted as a universal scale.

Adult

Validity and reliability of self-reported drinking behavior: dealing with the problem of response bias.

This work assesses the validity and reliability of self-reported survey data on drinking behavior. There is evidence to suggest that data are adversely affected by bias from underreporting. This bias affects the validity of measures of consumption of alcohol and can have deleterious effects on the results of some forms of statistical estimation. Data for this study were collected at an isolated military base. The remoteness of this site and the fact that it is a military station made it possible to estimate the actual level of consumption of alcohol for the population by assessing apparent consumption through officially recorded sales of alcohol. The results of eight measures of consumption of alcohol were compared with apparent consumption, as established by documented sales, and the validity and reliability of the various measures were determined using the classical correlational approach. The validity and reliability of the data generated by the self-report survey were also analyzed using LISREL, the measurement model in particular. The results indicate that various instruments used to assess the consumption of alcohol produce very different outcomes in terms of their validity and reliability, some questions being considerably more valid and reliable than others. Two of the more salient characteristics of questions that affect validity and reliability were isolated, namely a question's ability to aid recall and its ability to mitigate the effects of persons providing socially desirable responses. The LISREL results show that these are two underlying factors for the measurement of the consumption of alcohol. It is concluded that questions that produce valid and reliable responses do so for identifiable reasons, and measurement instruments can be improved by incorporating particular features.

Alcohol Drinking

Two measurement techniques for assessing subtalar joint position: a reliability study.

Proper assessment of the subtalar joint is critical for foot and ankle evaluation. Yet, reliability of open kinetic chain goniometric measurements of the subtalar joint has been poor. Two alternative techniques, navicular height and calcaneal position with an inclinometer, have been reported in the literature but lack reliability assessment. The purpose of this study was to determine the intertester and intratester reliability of navicular height and calcaneal position using an inclinometer. Thirty healthy, volunteer subjects (22 females, age 24 +/- 3.6 years; eight males, age 25 +/- 5.1 years) participated in this study. Two testers performed repeated measures on both feet of each subject (N = 60) during two testing sessions. Testers determined the 1) subtalar neutral position, 2) resting position, and 3) difference between these two measurements using an inclinometer for calcaneal position and navicular height. Intratester and intertester reliabilities (ICC 2, 1), standard errors of measurement, and 95% confidence intervals were determined. Intertester and intratester reliability for calcaneal position ranged from .68 to .91 for all measurements. Intertester and intratester reliability for navicular height ranged from .73 to .96 for all measurements. We conclude that these weight-bearing measurement techniques are reliable and acceptable for clinical and research purposes as measured. In addition, we hypothesize that these measurement techniques are simpler than previously described open kinetic chain methods.

Adult

Reliability of the probability effect on event-related potentials during repeated testing.

The reliability of event-related potentials (ERPs) was studied in 10 healthy adults who were tested 8 times over 7-10 day intervals using a standard auditory oddball paradigm. The difference waveforms, obtained by subtracting the averaged waveforms for frequent trials from those obtained in rare trials, were designed to analyze the components of the ERPs, such as the P300, and to focus on the reliability of the probability effect on the ERPs. The between-session reliability (8 sessions) and the within-session reliability (order of blocks or of different visual procedures) were computed for the obtained difference waveforms. The between-session reliabilities, expressed as the intraclass correlations (r') for the P300 amplitude, area and latency, were 0.70, 0.61 and 0.65, respectively. The within-session reliability, presented as the Pearson correlation coefficients (r) for the three P300 measures were 0.43, 0.35 and 0.25 for different eyes. The values were 0.45, 0.39, 0.42 between the first and the second blocks (eyes-open) and 0.58, 0.47 and 0.29 (eyes-closed). These findings indicate that the P300 amplitude calculated from the difference waveforms may be the most stable marker for the between-session reliabilities. There were no significant differences in the P300 measures over the 8 sessions, suggesting that habituation may not occur with the difference waveform reflecting the probability effect on ERPs. The difference waveform may be useful in research on repetitive group ERPs.

Adult

Clinical reliability of shoulder function assessment in patients with rheumatoid arthritis.

A model for functional assessment and a dynamic test of the shoulder joint were designed and tested for normal variation and clinical inter- and intra-rater reliability. The functional assessments, which covered four common shoulder functions, were compared with assessments of pain, recordings of active motion range and the results of a Health Assessment Questionnaire, in eight patients with rheumatoid arthritis according to the ARA criteria. Intra-rater reliability was satisfactory for all four functions and inter-rater reliability was satisfactory for the hand-raising and hand-to-opposite-shoulder functions but less so for hand-behind-back and hand-to-neck. A second test-retest study in 15 patients, with a slight modification of one of the functional tests, confirmed the results and improved the reliability of the modified test. The reliability of the dynamic test and of the active motion range measurement was less satisfactory or not satisfactory. No significant correlation was found between shoulder functional assessment and the Fries index, but there were positive significant correlations between active motion range and shoulder functions. It is concluded that the method presented for evaluating shoulder functions has satisfactory reliability and in the first test-retest study was more reliable than conventional motion range measurement of the shoulder joint.

Adult

Validity and reliability of lupus activity measures in the routine clinic setting.

As part of a cohort study of 150 patients with systemic lupus erythematosus (SLE), we investigated the validity and reliability of several indices of lupus activity, including the UCSF/JHU Lupus Activity Index (LAI), the SLE Disease Activity Index (SLEDAI), and a simple Core Index combining common elements. Validity was assessed by measuring correlations of these indices at the first cohort visit with the physician's global assessment (PGA) of SLE activity. The correlation of M-LAI (LAI modified so as not to contain PGA) and SLEDAI with PGA was 0.64 (95% CI 0.50, 0.70) and 0.55 (95% CI 0.42, 0.64), respectively. Reliability was assessed in a study of 6 patients seen twice, one week apart, by 9 physicians. The interrater reliability and test-retest reliability was greater for LAI (or M-LAI) than for SLEDAI. The Core Index performed better in its correlation with PGA (R = 0.78), although it contained no treatment data or serologic tests. Its interrater reliability and test-retest reliability were comparable with LAI. We conclude that (1) all indices have high validity; (2) LAI and the Core Index have higher reliability; and (3) these indices can be readily assimilated into routine clinic practice.

Adult

Measurement precision and reliability in craniofacial anthropometry: implications and suggestions for clinical applications.

Craniofacial anthropometry has become an important tool used by both clinical geneticists and reconstructive surgeons. Yet little attention has been paid to the potentially serious problem of measurement error. This paper examines intra-observer measurement error and precision (also called repeatability or reliability) for 52 commonly used anthropometric variables of the head and face. Two factors proved critical to reliability: magnitude of the measurement in question and the degree to which its constituant landmarks could be readily identified. Thus, all of the measurement variables with means above 10 cm proved to have good or excellent reliability. In contrast measurement variables with means below 10 cm were more likely to have poor reliability. This trend was especially evident in variables with means of 6 cm or less where 18 of the 20 variables in this range had poor reliability. The least reliable variables were those like philtrum breadth, columella breadth, and nasal root breadth that combine small magnitude with difficult to define landmarks. While these results suggest that it may be prudent to avoid using craniofacial variables with small dimensions this may be neither practical nor desirable. In such cases repeat measurements may be the best means for optimizing reliability.

Adult

Interexaminer reliability of the electromagnetic radiation receiver for determining lumbar spinal joint dysfunction in subjects with low back pain.

Twenty subjects (6 male, 14 female) with low back pain were examined by two experienced and licensed chiropractic doctors (E1 and E2). Both examiners examined the patients using a Toftness Electromagnetic Radiation Receiver (EMRR) and by manual palpation (MP) of the spinous processes. Interexaminer reliability was calculated at three sites (L3, L4, L5) for the following combinations: a) E1,MP--E2,MP; b) E1,EMRR--E2,EMRR; c) E1,MP--E2,EMRR; and) d) E2,MP--E1,EMRR, and intraexaminer reliability was calculated for the following variables: e) E1,MP--E1,EMRR; and f) E2,MP--E2,EMRR. Results of a Kappa coefficient analysis for interexaminer reliability of the stated combinations and at the specific sites were: a) -0.071, 0.400, 0.200; b) -0.013, 0.100, -0.120; c) 0.286, 0.300, 0.200; d) -0.081, 0.000, 0.048. These results predominantly indicate a poor to fair interexaminer reliability. The results of a Kappa coefficient analysis for intraexaminer reliability of the stated combinations were: e) 0.111, 0.400, 0.737; f) 0.000, 0.100, 0.368. These results indicate a poor to fair reliability. It was concluded that in subjects with low back pain the EMRR may not be a reliable indicator of spinal joint dysfunction.

Adult

Effects of specific criteria and calibration on examiner reliability.

The purpose of this pilot study was to investigate the use of specific criteria and examiner calibration on the reliability of inexperienced examiners on dental sealant evaluations. Dental (N = 8) and dental hygiene (N = 8) students participated as examiners. The study objectives were to identify differences in calibrated and non-calibrated examiners, examiners calibrated by an expert or non-expert, and reliability between dental and dental hygiene student examiners. A criterion-referenced evaluation form was used to evaluate dental sealant end product on 20 teeth, twice by each examiner. Eight of 16 examiners participated in a one-hour calibration session between evaluations. The session consisted of a discussion of operational definitions, the evaluation procedure for dental sealants, and use of the criterion-referenced form. Intra- and interexaminer reliabilities were measured. There were no statistically significant differences (p less than .05) in intraexaminer reliability. Although calibration produced no significant increase in interexaminer reliability, the post-training reliability scores for the group calibrated by an expert decreased, and scores for the group calibrated by a non-expert increased. No significant difference was found in reliability between dental and dental hygiene student examiners.

Humans

Reliability in perimetry.

As perimetric instrumentation becomes more sophisticated, patient reliability emerges as an important limiting factor in testing. Modern instrumentation for threshold and suprathreshold perimetry incorporate up to five separate indicators of patient reliability. For these perimetric methods, patient reliability is enhanced with specific techniques such as refractive correction, control of pupil size, and actively monitoring patient responses. With the manual (Goldmann) perimeter and the tangent screen, special statokinetic techniques help in both assessment and enhancement of patient reliability. In screening perimetry, reliability is assessed by analyzing the relative number, relative location, and repeatability of misses. Reliability in confrontation perimetry is both assessed and enhanced by using finger-counting and color-naming techniques. Review of the ophthalmic literature on perimetry shows how the various methods of patient reliability assessment and enhancement can be applied in the clinic.

Humans

A study of the reliability of carcinoembryonic antigen blood levels in following the course of colorectal cancer.

Twenty-three patients were studied to assess the reliability of carcinoembryonic antigen (CEA) levels in following the course of colorectal cancer. CEA estimations were made prior to surgery and again postoperatively. The resected specimens were allocated a Dukes' Stage and histological grading (well, moderate or poorly differentiated). In addition, sections were stained for the presence of CEA by an immunoperoxidase method. Of the 23 patients, twelve had either disseminated disease at initial surgery or subsequently developed metastasis/recurrence. Eleven remain disease-free at a minimum follow-up of one year. In all of these the reliability of plasma CEA values in reflecting the disease status has been assessed. No false positive elevations of CEA were found. Three factors emerge as positive predictors of CEA estimation reliability: pre-operative CEA elevation; tumour grading as well differentiated; dark staining for the presence of CEA. These factors identified 15 of the 18 patients (83%) in whom CEA appeared reliable and were not present in any of the five patients where CEA was not reliable. This reliability achieves statistical significance (X2 = 8.5, p less than 0.02). Histological demonstration of CEA may contribute to the reliability placed on plasma CEA estimations and should be considered if serial estimations are to be performed.

Carcinoembryonic Antigen