PubMed HealthSearch

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Validity and reliability of lupus activity measures in the routine clinic setting.

As part of a cohort study of 150 patients with systemic lupus erythematosus (SLE), we investigated the validity and reliability of several indices of lupus activity, including the UCSF/JHU Lupus Activity Index (LAI), the SLE Disease Activity Index (SLEDAI), and a simple Core Index combining common elements. Validity was assessed by measuring correlations of these indices at the first cohort visit with the physician's global assessment (PGA) of SLE activity. The correlation of M-LAI (LAI modified so as not to contain PGA) and SLEDAI with PGA was 0.64 (95% CI 0.50, 0.70) and 0.55 (95% CI 0.42, 0.64), respectively. Reliability was assessed in a study of 6 patients seen twice, one week apart, by 9 physicians. The interrater reliability and test-retest reliability was greater for LAI (or M-LAI) than for SLEDAI. The Core Index performed better in its correlation with PGA (R = 0.78), although it contained no treatment data or serologic tests. Its interrater reliability and test-retest reliability were comparable with LAI. We conclude that (1) all indices have high validity; (2) LAI and the Core Index have higher reliability; and (3) these indices can be readily assimilated into routine clinic practice.

Adult

Measurement precision and reliability in craniofacial anthropometry: implications and suggestions for clinical applications.

Craniofacial anthropometry has become an important tool used by both clinical geneticists and reconstructive surgeons. Yet little attention has been paid to the potentially serious problem of measurement error. This paper examines intra-observer measurement error and precision (also called repeatability or reliability) for 52 commonly used anthropometric variables of the head and face. Two factors proved critical to reliability: magnitude of the measurement in question and the degree to which its constituant landmarks could be readily identified. Thus, all of the measurement variables with means above 10 cm proved to have good or excellent reliability. In contrast measurement variables with means below 10 cm were more likely to have poor reliability. This trend was especially evident in variables with means of 6 cm or less where 18 of the 20 variables in this range had poor reliability. The least reliable variables were those like philtrum breadth, columella breadth, and nasal root breadth that combine small magnitude with difficult to define landmarks. While these results suggest that it may be prudent to avoid using craniofacial variables with small dimensions this may be neither practical nor desirable. In such cases repeat measurements may be the best means for optimizing reliability.

Adult

Interexaminer reliability of the electromagnetic radiation receiver for determining lumbar spinal joint dysfunction in subjects with low back pain.

Twenty subjects (6 male, 14 female) with low back pain were examined by two experienced and licensed chiropractic doctors (E1 and E2). Both examiners examined the patients using a Toftness Electromagnetic Radiation Receiver (EMRR) and by manual palpation (MP) of the spinous processes. Interexaminer reliability was calculated at three sites (L3, L4, L5) for the following combinations: a) E1,MP--E2,MP; b) E1,EMRR--E2,EMRR; c) E1,MP--E2,EMRR; and) d) E2,MP--E1,EMRR, and intraexaminer reliability was calculated for the following variables: e) E1,MP--E1,EMRR; and f) E2,MP--E2,EMRR. Results of a Kappa coefficient analysis for interexaminer reliability of the stated combinations and at the specific sites were: a) -0.071, 0.400, 0.200; b) -0.013, 0.100, -0.120; c) 0.286, 0.300, 0.200; d) -0.081, 0.000, 0.048. These results predominantly indicate a poor to fair interexaminer reliability. The results of a Kappa coefficient analysis for intraexaminer reliability of the stated combinations were: e) 0.111, 0.400, 0.737; f) 0.000, 0.100, 0.368. These results indicate a poor to fair reliability. It was concluded that in subjects with low back pain the EMRR may not be a reliable indicator of spinal joint dysfunction.

Adult

Effects of specific criteria and calibration on examiner reliability.

The purpose of this pilot study was to investigate the use of specific criteria and examiner calibration on the reliability of inexperienced examiners on dental sealant evaluations. Dental (N = 8) and dental hygiene (N = 8) students participated as examiners. The study objectives were to identify differences in calibrated and non-calibrated examiners, examiners calibrated by an expert or non-expert, and reliability between dental and dental hygiene student examiners. A criterion-referenced evaluation form was used to evaluate dental sealant end product on 20 teeth, twice by each examiner. Eight of 16 examiners participated in a one-hour calibration session between evaluations. The session consisted of a discussion of operational definitions, the evaluation procedure for dental sealants, and use of the criterion-referenced form. Intra- and interexaminer reliabilities were measured. There were no statistically significant differences (p less than .05) in intraexaminer reliability. Although calibration produced no significant increase in interexaminer reliability, the post-training reliability scores for the group calibrated by an expert decreased, and scores for the group calibrated by a non-expert increased. No significant difference was found in reliability between dental and dental hygiene student examiners.

Humans

AIDS knowledge and attitudes among injection drug users: the issue of reliability.

Among injection drug users (IDUs), AIDS-related knowledge and attitudes have not consistently predicted AIDS risk behavior. This may be due in part to the limited reliability of indexes used to measure drug users' AIDS knowledge and attitudes. In addition, the substantive interpretation of findings is confounded if index reliability is lower for particular demographic groups (e.g., ethnic populations and women). This report is based on 8 measures of AIDS-related knowledge and attitudes in a sample of 332 injection drug users in Los Angeles. The reliability of knowledge and attitude indexes for the overall sample is generally acceptable for the purpose of group comparison (average alpha = .60). But reliability is consistently lower for respondents who are Hispanic (average alpha = .49) and respondents with less formal education (alpha = .56). The reliability of 2 measures of sex-related attitudes is lower for female respondents. It is therefore important that the reliability of knowledge and attitude indexes be assessed not just for drug-user samples as a whole, but also within demographic groups of substantive interest.

Acquired Immunodeficiency Syndrome

Reliability and validity in binary ratings: areas of common misunderstanding in diagnosis and symptom ratings.

Confusion may exist between the reliability of a binary rating (for example, schizophrenia versus not-schizophrenia) and its implications for validity. High reliability does not guarantee validity, but paradoxically, low reliability does not imply poor validity in all contexts. Changes in the base rate or in experimental design may indicate high validity even when the reliability was thought to be low. Attempts to improve the psychiatric nomenclature by increasing only reliability run the risk of the "attenuation paradox" where further increases in reliability will make the ratings less valid. Finally, the assumption of random error in making diagnoses does not always hold, so that statistical analyses must be adjusted accordingly. New statistical methods are needed to index only false-positive or false-negative rates in order to quantify the error that will reduce some validity coefficients.

Bipolar Disorder

Reliability of topographic quantitative EEG amplitude in healthy late-middle-aged and elderly subjects.

Reliabilities of quantitative measures of absolute and relative EEG amplitudes were assessed in healthy older adults under the eyes closed (n = 46) and eyes opened (n = 45) conditions. For the theta, alpha, beta 1, and beta 2 bands, reliabilities of 28 scalp derivations were stable over the 4.5 month test interval. Reliabilities of delta were lower. When appropriate transformations were applied, the reliabilities of absolute EEG amplitude measures tended to exceed those of relative measures. There were not, however, striking differences in reliabilities under the eyes closed, as compared to eyes opened condition. We concluded that when coupled with the criterion of interpretability, the generally higher reliabilities of absolute, as opposed to relative, amplitude measures render them preferable in clinical research.

Aged

Test-retest reliability of the P50 mid-latency auditory evoked response.

Attenuation in mid-latency auditory evoked responses (MLAERs) can be used to study sensory gating. If paired-click stimuli (S1 and S2) are used, lower amplitude in response to S2 vs. S1 (attenuation) is considered evidence for intact sensory gating. However, the need for reliable measurements of MLAER amplitude and attenuation is a recognized problem. Ten normal volunteers were studied six times each. An S1 amplitude test-retest reliability coefficient (r) of 0.585 was obtained when means of two recordings were used vs. reliability coefficients as high as 0.809 for means of six recordings. Averaging a higher number of runs (120 vs. 60) resulted in a reliability coefficient of 0.677/recording. Similar values were obtained for S1 and S2 latencies. Reliability coefficients for S2 attenuation (S2/S1) were not nearly as high (a value of 0.138 when means of all six recordings were used). The S1 amplitude as measured in this study (with 120 averages) appears to be a reliable psychophysiologic measurement, but the S2/S1 attenuation measure is more variable, perhaps reflecting a greater sensitivity of the S2/S1 to uncontrolled variables in this study. Further research to identify such variables is necessary.

Adult

The reliability and stability of a quantity-frequency method and a diary method of measuring alcohol consumption.

The study aimed to assess the test-retest reliability of two commonly used measures of alcohol consumption, the quantity-frequency (QF) method and the diary method, as well as the stability of scores on the two measures over time. Two methods of assessing reliability and stability were employed. The first was a traditional method based on calculation of correlation coefficients for agreement between scores on repeated measures over a short retest interval to yield test-retest reliability coefficients, and over a long retest interval to yield stability coefficients. The second method was that devised by Wiley and Wiley (1970) to differentiate the effects of reliability and stability on repeated measures over time. The two methods were applied to a sample of heavy drinkers and to a sample of light drinkers. The results indicated that both the QF and diary measures are reliable in measuring alcohol consumption of light drinkers. Both measures are less reliable for heavy drinkers. The results indicate, in addition, that drinking consumption levels of light drinkers demonstrate a high degree of stability. However, the consumption levels of heavy drinkers demonstrate less stability, especially over a long time period. Heavy drinkers significantly reduced reported levels of alcohol consumption on both measures after the first test, suggesting a regression to the mean effect or the possibility of unintended intervention effects due to repeated measurement of drinking behaviour.

Adult

Reliability of seven measures of social intelligence in a sample of adolescents with mental retardation.

This study evaluates the reliability of seven measures, selected to assess the social-cognitive variables hypothesized by Greenspan to define social intelligence. Responses from 75, 30 and 20 adolescents with mental retardation were used to assess each test's internal, interrater, and test-retest reliabilities, respectively. Interrater reliability coefficients were high to very high (.76 to .98), internal reliabilities were moderate to very high (.66 to .90), and test-retest reliabilities were moderate to high (.54 to .74). Internal and test-retest reliability coefficients compared favourably with those reported for the subtests of the Revised Wechsler Intelligence Scale for Children.

Activities of Daily Living

Reliability of clinical measurements of lumbar lordosis taken with a flexible rule.

The purpose of this study was to examine the intratester and intertester reliability of lumbar lordosis measurements taken with a flexible rule. Two physical therapists (Tester 1 and Tester 2) took measurements on 40 subjects without low back pain (LBP) and on 40 subjects with LBP. Intraclass correlation coefficients (ICCs) were used to determine the degree of agreement between repeated measurements taken by the same therapist and between measurements taken by the two therapists. The ICC values for intratester reliability of Tester 1 were .84 for subjects without LBP and .94 for subjects with LBP. The ICC values of Tester 2 were .73 for subjects without LBP and .83 for subjects with LBP. Intertester reliability generally was poor, with ICC values of .41 for subjects without LBP and .50 for subjects with LBP. The results suggest that measurements of lumbar lordosis with a flexible rule may be reliable if taken by the same physical therapist. The degree of reliability, however, may vary from therapist to therapist. The intertester reliability of these measurements appears to be poor, but these conclusions must be interpreted carefully because of the limited number of therapists participating in this study.

Adult

Intrarater reliability of manual muscle test (Medical Research Council scale) grades in Duchenne's muscular dystrophy.

The purpose of this study was to document the intrarater reliability of manual muscle test (MMT) grades in assessing muscle strength in patients with Duchenne's muscular dystrophy (DMD). Subjects were 102 boys, aged 5 to 15 years, who were participating in a double-blind, multicenter trial to document the effects of prednisone on muscle strength in patients with DMD. Four physical therapists participated in the study. Two identical (duplicate) evaluations were performed within 5 days of each other by the same examiner initially and after 6 and 12 months of treatment. A total of 18 muscle groups were tested on each patient, 16 of them bilaterally, using a modification of the Medical Research Council scale. Reliability of muscle strength grades obtained for individual muscle groups and of individual muscle strength grades was analyzed using Cohen's weighted Kappa. The reliability of grades for individual muscle groups ranged from .65 to .93, with the proximal muscles having the higher reliability values. The reliability of individual muscle strength grades ranged from .80 to .99, with those in the gravity-eliminated range scoring the highest. We conclude the MMT grades are reliable for assessing muscle strength in boys with DMD when consecutive evaluations are performed by the same physical therapist.

Adolescent

Reliability of lumbar isometric torque in patients with chronic low back pain.

In this study, the test-retest reliability of lumbar isometric strength testing in patients with chronic low back pain (CLBP) was assessed. Isometric torque measurements were obtained from 89 patients with CLBP at seven different angles of lumbar flexion. Because previous studies have demonstrated significant strength differences between male and female subjects, separate data analyses were performed for each gender. Results indicated moderate to high reliability for patients with CLBP when tested at individually determined angles of flexion within their idiosyncratic range of motion (ROM) (female subjects: r = .59-.96, P less than .05, SEE = 12.0-24.2 N.m; male subjects: r = .71-.93, P less than .05, SEE = 25.1-62.1 N.m). For comparison with previously published data on asymptomatic controls, an additional set of analyses was conducted for subjects with full lumbar ROM. Similar reliability was demonstrated for this subsample (female subjects: r = .57-.93, P less than .05, SEE = 12.4-27.9 N.m; male subjects: r = .63-.93, P less than .05, SEE = 34.2-44.2 N.m). The authors concluded that isometric lumbar extension torque could be reliably measured in patients with CLBP at multiple positions within the full ROM, although reliability decreased at the most extended positions. The demonstrated reliability will allow researchers to assess treatment effects and group differences without undue concern for artifact attributable to measurement error.

Back Pain

On the methods and theory of reliability.

This paper reviews the most frequently used and misused reliability measures appearing in the mental health literature. We illustrate the various types of data sets on which reliability is assessed (i.e., two raters, more than two raters, and varying numbers of raters with dichotomous, polychotomous, and quantitative data). Reliability statistics appropriate for each data format are presented, and their pros and cons illustrated. Inadequancies of some methods are highlighted. The meaning of different levels of reliability obtained with various statistics is discussed. This critique is intended for the reading professional and the investigator who has an occasional need for reliability assessment. Statistical expertise is not required and theoretical material is referenced for the interested reader. Necessary formulas for computations are presented in the appendices. A summary table of some suitable reliability measures is presented.

Humans

The diagnosis of hypersensitivity to ingested foods. Reliability of skin prick testing and the radioallergosorbent test with different materials.

The diagnostic reliability in food allergy of skin prick tests (SPT) and the radio-allergosorbent test (RAST) was investigated in paediatric patients with respiratory and skin allergies. SPT and RAST were found to be reliable for the diagnosis of allergy to codfish, peas, nuts, peanuts and egg white. Positive SPT and RAST to cereals were common, but were most often without clinical significance or were correlated with respiratory allergy to the inhalation of flour dust. SPT and RAST were only partly reliable with regard to allergy to cow's milk, and were mostly reliable when used together and showing corresponding results. Experimental allergosorbents for RAST with soy beans and white beans were not reliable. The study shows the need to improve the diagnostic materials and to establish the diagnostic reliability of the material and tests used for each food item in question.

Adolescent

Variability and reliability of joint measurements.

The purpose of this study was to determine the variability and reliability of joint measurements as carried out by three physician observers. The intratester variation and reliability of nine different joint measurements was determined in eight healthy subjects. The measurements were taken in eight sessions by each tester. In this population also the intertester variation and reliability was determined by the three observers. This was also done in a population of middle-aged athletes over a period of 2.5 years. The results indicate that it is difficult to show either an improvement or worsening of a joint motion of less than 5 degrees to 10 degrees for most joints measured by the same tester. The intertester variation is not consistent over a longer period of time, so differences between observers during long-term studies cannot be corrected on the basis of a single study at a single point in time. The reliability of all nine joint measurements is not very high, but is probably sufficient if the results are used to compare groups within a single population and for large studies with experienced observers. Because the reliability strongly depends on the interindividual variation, it is preferable to determine the reliability for each study population.

Adult

Percent of agreement among raters and rater reliability of the copying subtest of the Stanford-Binet Intelligence Scale: Fourth Edition.

The purpose of this study was to investigate the interrater reliability of the visual-motor portion of the Copying subtest of the Stanford-Binet Intelligence Scale: Fourth Edition. Eight raters independently scored 11 protocols completed by children aged 5 through 10 years, using the scoring criteria and guidelines in the manual. The raters marked each of 10 items pass or fail and computed a total raw score for each protocol. Interrater reliability coefficients were obtained for each child's protocol, and the Kappa coefficient was computed for each item. Significant raters' reliability coefficients ranged from .82 to .91, which were low in comparison to test-retest reliability and Kuder-Richardson-20 coefficients for this and other subtests of the Stanford-Binet in the technical manual. Percent agreement among 8 raters also indicated weak reliability. Although the obtained results suggested some interrater reliability coefficients within acceptable levels, questions were raised about the scoring criteria for individual items. Caution is warranted in the use of cognitive measures which include subjective judgement of the examiner in applying scoring criteria.

Child

Number of stimuli as a reliability parameter in perimetry.

Catch trials test patient performance during automated, static perimetry, but their adequacy to estimate reliability is uncertain even though up to 10% of the test time is reserved for catch trials. The 308 visual fields (program G1, all 3 phases, Octopus 201) of 308 eyes of 308 glaucoma, suspected glaucoma, and normal subjects were studied. The 108 visual fields (mean sensitivity > 10 dB; corrected loss variance < 50 dB2) without false responses to catch trials were considered reliable. A multiple linear regression analysis of these 108 fields was performed and revealed the following result (r2 = 0.751): Number of stimuli = 480 + (40.short-term fluctuation) + (8.8.the square root of the index corrected loss variance) - (2.2.mean sensitivity). This equation was used to estimate the number of stimuli required of a reliable subject to complete an examination. Excess stimuli would thus be a sign of reduced reliability. The difference between the estimated and the actual number of stimuli was called the 'stimulus discrepancy'. In 169 fields with false-positive and 58 fields with false-negative responses, the false-positive and false-negative responses correlated with the 'stimulus discrepancy' (r = 0.19, P = 0.014; r = 0.29, P < 0.026, respectively). The number of stimuli depends not only on reliability but also on the software and hardware of the perimeter. 'Stimulus discrepancy' may be an additional useful perimetric reliability parameter which does not require extra testing time.

Adult