PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Testing reliability of plaque and gingival indices. Two methods.

This investigation was undertaken to compare two methods of interexaminer and intraexaminer reliability in the evaluation of Plaque and Gingival Indices prior to a study of toothbrushing. Inter-/intraexaminer reliabilities were compared using a projected slide series consisting of 40 slides of clinical examples of gingival inflammation and plaque accumulation. Time between assessments was three weeks. Using the slide technique, intraexaminer reliability was established for: (1) Gingival Indices and (2) Plaque Indices. Interexaminer reliability was also established for Gingival Indices. Interexaminer reliability could not be established for Plaque Indices on the first assessment but was established on the post-assessment. Intraexaminer reliability was also determined through clinical examinations of patients. A third clinician was used to manipulate the tissue while investigators evaluated bleeding on provocation and plaque accumulation. Significant results were established for the Gingival Indices and Plaque Indices. Results of this investigation suggest that significant inter-/intraexaminer reliabilities may be obtained for gingival indices using the slide technique. In addition, the clinic technique appeared useful for assessing interexaminer reliability for Gingival Indices. Plaque Indices using the slide technique required more practice than those using the clinic technique.

Dental Health Surveys↗

Measures of reliability in sports medicine and science.

Reliability refers to the reproducibility of values of a test, assay or other measurement in repeated trials on the same individuals. Better reliability implies better precision of single measurements and better tracking of changes in measurements in research or practical settings. The main measures of reliability are within-subject random variation, systematic change in the mean, and retest correlation. A simple, adaptable form of within-subject variation is the typical (standard) error of measurement: the standard deviation of an individual's repeated measurements. For many measurements in sports medicine and science, the typical error is best expressed as a coefficient of variation (percentage of the mean). A biased, more limited form of within-subject variation is the limits of agreement: the 95% likely range of change of an individual's measurements between 2 trials. Systematic changes in the mean of a measure between consecutive trials represent such effects as learning, motivation or fatigue; these changes need to be eliminated from estimates of within-subject variation. Retest correlation is difficult to interpret, mainly because its value is sensitive to the heterogeneity of the sample of participants. Uses of reliability include decision-making when monitoring individuals, comparison of tests or equipment, estimation of sample size in experiments and estimation of the magnitude of individual differences in the response to a treatment. Reasonable precision for estimates of reliability requires approximately 50 study participants and at least 3 trials. Studies aimed at assessing variation in reliability between tests or equipment require complex designs and analyses that researchers seldom perform correctly. A wider understanding of reliability and adoption of the typical error as the standard measure of reliability would improve the assessment of tests and equipment in our disciplines.

Algorithms↗

Short- and long-term reliability of information on previous illness and family history as compared with that on smoking and drinking habits in questionnaire surveys.

To assess the reliability of responses to questionnaires regarding previous illness, family history of cancer, and smoking and drinking habits, we repeated questionnaire surveys four times at intervals of 2 weeks and 1 year (short-term), and 4.5 years (long-term) among 440 subjects aged 40-69. The reliability was assessed using kappa statistic. Kappa was calculated both for complete data and data including missing values. The changes of mode of pre-after paired responses were also investigated. Our results from complete data showed both short- and long-term reliabilities of replies regarding smoking or drinking were excellent (mean kappa 0.85-0.99). The reliability of previous illness was excellent except for stroke for short-term intervals (mean kappa 0.85-1.00), but varied depending on the kinds of illness with long-term intervals (mean kappa -0.01-0.75). Responses to family history had fair to excellent short-term reliability (mean kappa 0.54-0.85). Inclusion of missing value as an independent category reduced reliability remarkably. Subjects stating absence of medical history were more likely to have missing values for this item than subjects with some history. In conclusion, the reliability for information given on previous illness was as good as that on smoking and drinking for a short interval, but was lower for a long-interval probably due to the development of new cases. The reliability of a family history on cancer was slightly poorer than that of individual's previous cancer or other illnesses and that of smoking and drinking even for a short interval.

Adult↗

Reliability of health information on the Internet: an examination of experts' ratings.

BACKGROUND: The use of medical experts in rating the content of health-related sites on the Internet has flourished in recent years. In this research, it has been common practice to use a single medical expert to rate the content of the Web sites. In many cases, the expert has rated the Internet health information as poor, and even potentially dangerous. However, one problem with this approach is that there is no guarantee that other medical experts will rate the sites in a similar manner. OBJECTIVES: The aim was to assess the reliability of medical experts' judgments of threads in an Internet newsgroup related to a common disease. A secondary aim was to show the limitations of commonly-used statistics for measuring reliability (eg, kappa). METHODS: The participants in this study were 5 medical doctors, who worked in a specialist unit dedicated to the treatment of the disease. They each rated the information contained in newsgroup threads using a 6-point scale designed by the experts themselves. Their ratings were analyzed for reliability using a number of statistics: Cohen's kappa, gamma, Kendall's W, and Cronbach's alpha. RESULTS: Reliability was absent for ratings of questions, and low for ratings of responses. The various measures of reliability used gave conflicting results. No measure produced high reliability. CONCLUSIONS: The medical experts showed a low agreement when rating the postings from the newsgroup. Hence, it is important to test inter-rater reliability in research assessing the accuracy and quality of health-related information on the Internet. A discussion of the different measures of agreement that could be used reveals that the choice of statistic can be problematic. It is therefore important to consider the assumptions underlying a measure of reliability before using it. Often, more than one measure will be needed for "triangulation" purposes.

Delivery of Health Care↗

Intramachine and intermachine reliability for selected dynamic muscle performance tests.

The Cybex 6000 isokinetic dynamometer is a new isokinetic device for which no published reports of reliability have been presented in the literature. In addition, the manufacturer not only claims that the new Cybex 6000 is reliable but that torque data obtained from the Cybex 6000 are consistent with data obtained from past Cybex systems, such as the Cybex II. The purpose of this study was to investigate the intramachine reliability of the Cybex 6000 to itself and the intermachine reliability of the Cybex 6000 and the Cybex II. Data on peak torque, work, and power were collected using the Cybex 6000, and data on peak torque were obtained using the Cybex II for knee flexion and extension in 20 volunteers (10 males, 10 females). Subjects were tested three times, twice on the Cybex 6000 and once on the Cybex II, approximately 1 week apart across a 3-week period of time at angular velocities of 60, 180, and 300 degrees/sec. Data were analyzed using intraclass correlations. Results indicated that the majority of test-retest correlation coefficients for all parameters for intramachine reliability of the Cybex 6000 were above .90. Comparing peak torque obtained with the Cybex 6000 to that obtained with the Cybex II (intermachine reliability), correlation coefficients ranged from .72 to .89. In conclusion, information obtained on the Cybex 6000 appears to be quite reliable in a test-retest situation using the same equipment and moderately reliable when compared to the Cybex II. Clinical implications for these results are discussed.

Adult↗

Relationship of the pelvic angle to the sacral angle: measurement of clinical reliability and validity.

There is a need to better document the reliability and validity of assessment measures used in physical therapy. Studies documenting the reliability of measurement of the pelvic angle and its relationship to sacral motion are presently inconclusive. The purpose of this study was twofold. First, we wanted to determine the reliability and validity of a goniometric measurement of the pelvic angle. We also wanted to test the hypothesis that there is a relationship between the pelvic angle and the sacral angle. Intertester and intratester reliability of goniometric pelvic angle measurements of 23 healthy young adults were examined using three different raters. Radiographic measurements of the pelvic and sacral angle using two raters and goniometric measurement of the pelvic angle using a single rater were taken from 15 patients with low back pain who had been referred for X-rays. Intraclass correlation coefficients (ICCs) of intratester reliability for goniometric measurements of the pelvic angle were .93, .96, and .96. The intertester reliability was .95. The ICCs for intratester reliability for radiological measurements were .92 and .95 for the sacral angle and .98 for both measurements of the pelvic angle. Intertester reliability coefficients were .86 and .88, respectively. The Pearson correlation coefficients for the goniometric and radiological measurements of the pelvic angle were .85 and .68. A comparison of the radiological and goniometric measurements of the pelvic angle with the sacral angle demonstrated low average correlations of .43 and .58, respectively. The results indicate a high level of correlation between and within testers for goniometric measurements of the pelvic angle but only a fair correlation between goniometric and radiological measurements of the pelvic angle.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult↗

Reliability of McConnell's classification of patellar orientation in symptomatic and asymptomatic subjects.

STUDY DESIGN: Test-retest reliability study with blinded testers. OBJECTIVES: To determine the intratester reliability of the McConnell classification system and to determine whether the intertester reliability of this system would be improved by one-on-one training of the testers, increasing the variability and numbers of subjects, blinding the testers to the absence or presence of patellofemoral pain syndrome, and adhering to the McConnell classification system as it is taught in the "McConnell Patellofemoral Treatment Plan" continuing education course. BACKGROUND: The McConnell classification system is currently used by physical therapy clinicians to quantify static patellar orientation. The measurements generated from this system purportedly guide the therapist in the application of patellofemoral tape and in assessment of the efficacy of treatment interventions on changing patellar orientation. METHODS AND MEASURES: Fifty-six subjects (age range, 21-65 years) provided a total of 101 knees for assessment. Seventy-six knees did not produce symptoms. A researcher who did not participate in the measuring process determined that 17 subjects had patellofemoral pain syndrome in 25 knees. Two testers concurrently measured static patellar orientation (anterior/posterior and medial/lateral tilt, medial/lateral glide, and patellar rotation) on subjects, using the McConnell classification system. Repeat measures were performed 3-7 days later. A kappa (kappa) statistic was used to assess the degree of agreement within each tester and between testers. RESULTS: The kappa coefficients for intratester reliability varied from -0.06 to 0.35. Intertester reliability ranged from -0.03 to 0.19. CONCLUSION: The McConnell classification system, in its current form, does not appear to be very reliable. Intratester reliability ranged from poor to fair, and intertester reliability was poor to slight. This system should not be used as a measurement tool or as a basis for treatment decisions.

Adult↗

Reliability of some tremor measurement outcome variables in field testing situations.

OBJECTIVE: Many summary measures of data obtained from tremor measurement procedures are commonly reported. The reliability of many of these summary measures of tremor measurements made in field testing situations is unknown. The purpose of the present investigation was to assess the reliability of a number of summary measures produced by the software of a widely used, commercially available tremor measurement instrument using data collected in three field epidemiologic studies. METHODS: Tremor data were obtained from 689 participants in 3 previously conducted studies of groups exposed to elemental mercury or arsenic. A widely used, commercially available tremor measurement instrument was used. Two-axis accelerometer measurements were obtained on 2 or more trials from each hand for each participant. Estimates of trial-to-trial and internal consistency reliability were calculated for 5 summary measures calculated by instrument manufacturer's software and 5 additional summary measures calculated from data output by the software. RESULTS: An RMS acceleration measure had the highest reliability in all 3 studies. The average over 4 trials of RMS acceleration and its logarithm had high reliability (>0.9). Recalculation of a tremor summary index and a harmonicity index as suggested by Edwards and Beuter (1999) resulted in measures with higher reliability and better distributional shape than the corresponding measures provided by the instrument manufacturer's software. The results in all three studies were similar. CONCLUSIONS: For the tremor measurement instrument and testing procedure that we employed, we recommend using the common logarithm of the RMS accelerations and recalculated tremor index as summary measures. We also recommend employing multiple trials of each type (e.g., with each hand) and averaging summary measures from those trials to derive outcome measures of tremor for use in epidemiologic studies. We recommend at least 2 trials for RMS acceleration measures and more for less reliable measures, particularly for designs employing repeated measurements of individuals. Summary measures averaged over at least 4 trials for mean frequency, dispersion of frequency, and power in the 3-6.5 and 6.6-10 Hz frequency ranges have sufficiently high reliability for use in epidemiologic studies.

Aged↗

The reliability of H-reflex recordings in standing subjects.

OBJECTIVE: Several studies have used the H-reflex to investigate the effect of upright stances and locomotion on spinal reflex excitability. The reliability of eliciting this reflex during weight-bearing has however yet to be addressed. This study was undertaken to determine the reliability of individual differences in the H-reflex recorded from healthy subjects during quiet standing. Secondary aims of the study were to evaluate individual reliability during prolonged standing, and to establish the minimum number of trials required to provide reliable measurements of the H-reflex. DESIGN AND PARTICIPANTS: Twenty neurologically healthy volunteers participated in a repeated measures design consisting of 8 blocks of 20 trials evenly distributed over two testing sessions. H-reflex recordings were elicited from the subject's dominant side soleus muscle by percutaneously stimulating the posterior tibial nerve. Stimuli were presented every 10 s, with 2 min seated rest provided between blocks of trials. Peak-to-peak amplitude of H-reflexes and m-responses determined for individual trials were used for subsequent analysis. RESULTS: It was found that the reliability of measuring the H-reflex and m-response during quiet standing was extremely robust (r = .97 for both measures). This pattern of individual differences remained consistent over 80 trials confirming the stability of the measures. High reliabilities (r = .96 and .87 for the H-reflex and m-responses respectively) were also observed when as few as four trials were analysed. When measures obtained during the first session of testing were compared with those obtained for session two, the correlation coefficients were generally of a lower order (r = .54 to .90). DISCUSSION: The results demonstrate that the H-reflex in quiet standing provides high intra-individual reliability, suggestive of a stable reflex resistant to potentially confounding postural influences, or other sources of biological variation. The between-session reliability underscores the difficulty in reproducing conditions between sessions, and emphasises the need for within-session comparisons of H-reflex amplitudes. Given the functional challenge of maintaining an upright posture, the H-reflex appears to be a well maintained and stable phenomenon.

Adolescent↗

Test-retest reliability of the soleus H-reflex in three different positions.

PURPOSE: H-reflex has been clinically useful in the diagnosis of radiculopathies, developmental disorders, and measurement of motoneuron excitability. However, variability of the H-reflex precluded its routine application. The purpose of this study is to evaluate the test-retest and within-subject reliability of the soleus H-reflex tested in three different positions. SUBJECTS: Seven men and eight women healthy volunteered (20-50 y) with no history of significant low back pain or radiculopathy consented to the study. METHODS: The soleus H-reflexes for both lower extremities were elicited and recorded using Cadwell 5200-A EMG unit and surface recording. The tibial nerve was electrically stimulated at the popliteal fossa using 0-5 ms., 0.2 pps pulses at intensity equivalent to H-max. Each subject was tested randomly in three different positions: pronelying, free standing, and standing while lifting 20% of his/her body weight. Signal were amplified (1-5 K) using surface electrodes applied on the soleus muscle at midline and 3 cm below the gastrocnemius musculotendinous junction. The peak-to-peak amplitude and onset latencies of four separate traces were averaged for each trial. Subjects were re-tested within 10 days by the same tester following the same protocol. RESULTS: Test-retest reliability of the H-reflex amplitude ranged from r = .29 in prone position to r = .56 in the loading position. Within day reliability of the H-amplitude was high between the three different positions and ranged from r = .56 to r = .97. The test-retest reliability of the H-latency were extremely high and robust, with the coefficients ranged from r = .92 to r = .94. Also the within day reliability of the H-latency ranged from r = .96 to r = .99. CONCLUSIONS: Results indicated that, when the H-amplitude is the measure of choice, testing the H-reflex in standing and loading positions is more reliable than testing in pronelying. Also testing the subject during various procedures in the same session is more reliable than testing subject in different days/sessions. The H-latency is highly reliable in all three testing positions.

Adult↗

Reliability of the craniomandibular index.

AIMS: To examine various dimensions of reliability of the Craniomandibular Index, a commonly used instrument for quantifying the severity of signs and symptoms of temporomandibular disorders. METHODS: Classical psychometric theory and generalizability theory were used to assess the reliability of data obtained from a calibration study of examiners participating in a multi-site clinical trial and from a random community sample. RESULTS: The reliability of aggregate scores formed by summing individual binary scored items was high, with intraclass correlations ranging from 0.81 to 0.88. When it was required that examiners recognize and agree upon a specific pattern of signs and symptoms exhibited by a patient, however, reliability dropped dramatically (multivariate kappas ranged from 0.26 to 0.32). A group of practicing examiners also showed limited ability to agree with the pattern of signs and symptoms identified by a "gold standard" examiner (multivariate kappas ranging from 0.25 to 0.32). Generalizability analysis failed to identify the specific sources of measurement error that played a major role in limiting reliability but demonstrated that generalizability of aggregate scores was very high. CONCLUSION: Methods of classical psychometric theory and generalizability theory support the conclusion that the reliability of aggregate scores is acceptably high. Individual items assessing certain aspects of jaw mobility and joint sounds are measured with poor reliability. Reliability declines when it is defined as the ability of examiners to agree among themselves upon a specific constellation of signs and symptoms or their ability to identify correctly a "correct" constellation identified by an expert examiner.

Adult↗

Intrarater Reliability of Functional Performance Tests for Subjects With Patellofemoral Pain Syndrome.

OBJECTIVE: Patellofemoral pain syndrome (PFPS) is a common clinical entity seen by the sports medicine specialist. The ultimate goal of rehabilitation is to return the patient to the highest functional level in the most efficient manner. Therefore, it is necessary to assess the progress of patients with PFPS using reliable functional performance tests. Our purpose was to evaluate the intrarater reliability of 5 functional performance tests in patients with PFPS. DESIGN AND SETTING: We used a test-retest reliability design in a clinic setting. SUBJECTS: Two groups of subjects were studied: those with PFPS (n = 29) and those with no known knee condition (n = 11). The PFPS group included 19 women and 10 men with a mean age of 27.6 +/- 5.3 years, height of 169.80 +/- 10.5 cm, and weight of 69.59 +/- 15.8 kg. The normal group included 7 women and 4 men with a mean age of 30.3 +/- 5.2 years, height of 169.55 +/- 9.9 cm, and weight 69.42 +/- 14.6 kg. MEASUREMENTS: The reliability of 5 functional performance tests (anteromedial lunge, step-down, single-leg press, bilateral squat, balance and reach) was assessed in 15 subjects with PFPS. Secondly, the relationship of the 5 functional tests to pain was assessed in 29 PFPS subjects using Pearson product moment correlations. The limb symmetry index (LSI) was calculated in the 29 PFPS subjects and compared with the group of 11 normal subjects. RESULTS: The 5 functional tests proved to have fair to high intrarater reliability. Intrarater reliability coefficients (ICC 3,1) ranged from.79 to.94. For the PFPS subjects, a statistical difference existed between limbs for the anteromedial lunge, step-down, single-leg press, and balance and reach. All functional tests correlated significantly with pain except for the bilateral squat; values ranged from.39 to.73. The average LSI for the PFPS group was 85%, while the average LSI for the normal subjects was 97%. CONCLUSIONS: The 5 functional tests proved to have good intrarater reliability and were related to changes in pain. Future research is needed to examine interrater reliability, validity, and sensitivity of these clinical tests.

Journal Article↗

Consultation competence in general practice: testing the reliability of the Leicester assessment package.

BACKGROUND: An acceptable assessment must be both valid and reliable; the face validity of the Leicester assessment package has already been established. AIM: This study set out to test the reliability of the Leicester assessment package, and the factors influencing it, when used by multiple assessors to assess performance in general practice consultations. METHOD: Six randomly selected course organizer assessors simultaneously used the package to conduct independent assessments of the performance of five doctors of widely varying abilities in consultation with six simulated patients. The scores allocated were subjected to generalizability analysis. RESULTS: The mean scores allocated for consultation performance of individual doctors ranged from 51% to 70%, with the lower scores being allocated to the less experienced doctors. Scores of each assessor across the cases were examined for internal consistency and five of the six assessors consistently scored the doctors with an alpha coefficient of the minimum accepted level of 0.80 or greater. The other assessor had a consistency of only 0.22. Measurements of consistency within cases between markers indicated that the first case produced unreliable results (alpha coefficient 0.25) but all other cases were scored consistently. Two independent assessors scoring eight consultations are the requisite numbers to achieve acceptable levels of reliability in a formal assessment process; seven consultations produce the minimum acceptable generalizability coefficient of 0.80 plus the first 'non-counting' consultation. CONCLUSION: Required levels of reliability can be achieved when the package is used by multiple markers assessing the same consultations over a wide range of consultation performance. To achieve reliability only two hours of assessment time are required using the Leicester package compared with the previously regarded minimum of 32 hours. Although assessors can produce reliable scores with minimal training, intra-assessor reliability cannot be taken for granted and all assessors should be trained and calibrated before being sanctioned to conduct assessments, particularly for regulatory purposes. The Leicester assessment package has now been shown to be valid, reliable, feasible and easy to use in practice. It can, therefore, be recommended for use in both formative and summative assessment of consultation competence in general practice.

Communication↗

Chiropractic biophysics lateral cervical film analysis reliability.

OBJECTIVE: To determine the degree to which the geometric line drawings used in Chiropractic Biophysics Technique (CBP) on lateral cervical radiographs are reliable. DESIGN: A blind, delayed repeated measures design was used. Three examiners were presented radiographs in random order. All identifying marks were removed prior to each examiner's individual marking and measurement. Each examiner was blinded as to how the previous examiners marked and measured the radiographs. SETTING: Primary care private chiropractic clinic. PATIENTS PARTICIPANTS: Sixty-five subject films were provided from the patient records of a primary care private chiropractic clinic. The 65 radiographs qualified for inclusion in the study based on two criteria: C1 through C7 had to be clearly visible, and there had to be no identifying artifacts. MAIN OUTCOME MEASURES: Anterior head translation in millimeters, atlas plane to horizontal, Ruth Jackson's cervical stress lines, and five relative rotation angles for C2-C3, C3-C4, C4-C5, C5-C6, C6-C7. Inter- and intrareliability of the three examiners were statistically analyzed. RESULTS: Intraexaminer for a) C1 to horizontal reliability was .98-.99 with confidence intervals of .96-.99, b) absolute rotation angle from C2 to C7 reliability was .82-.95 with confidence intervals of .80-.99, c) anterior head translation [+Sz] reliability was .86-.99, with confidence intervals of .74-.99, d) relative rotation angle reliability ranges were (C2-C3) .99, and (C3-C4) .98-.99, (C4-C5) .88-.99, (C5-C6) .80-.99, and (C6-C7) .94-.98. Interexaminer reliabilities across examiners ranged from a) Winer:.89-.99 and b) Bartko: .72-.96. CONCLUSIONS: The reliabilities for intra- and interexaminer were all greater than .70, indicating that these measurements in CBP technique would be considered accurate enough to provide measurements for future clinical studies. The data indicated that the C6-C7 relative rotation angle was the least reliable measurement. This might be due to the very small angles found at this level.

Analysis of Variance↗

Intra- and interexaminer reliability of the chiropractic biophysics lateral lumbar radiographic mensuration procedure.

OBJECTIVE: To determine the intra- and interexaminer reliability of a specific method of mensuration commonly used to evaluate the positional configuration of the lumbopelvic spine viewed on lateral lumbar radiographs. DESIGN: A blind, repeated-measures design was used. Lateral lumbopelvic radiographs were presented to each of three examiners in random order. Each film was marked and measurements were recorded. The films were cleaned of all markings and randomized again for a second run by each examiner. Each examiner's measurements were unavailable to the other examiners. SETTING: Private, primary-care chiropractic clinic. MAIN OUTCOME MEASURES: Anterior/posterior thoracic translation in millimeters, Ferguson's sacral-plane angle to horizontal, arcuate line angle to horizontal, L1 to L5 absolute rotation angle and four relative rotation angles for L1-L2, L2-L3, L3-L4 and L4-L5. Intra- and interrelibility of the three radiographic examiners were analyzed. RESULTS: Intraexaminer reliability for (a) L1-L5 absolute rotation angle was .98, with confidence intervals included in the range of 0.95-0.99, (b) anterior/posterior thorax translation [+/- Sz] was .97-.99, with confidence intervals included in the range of 0.94-1.00, (c) arcuate angle (AA) .40-.81, with confidence intervals included in the range of 0.07-0.90, (d) Ferguson's angle (FA) was .91-.97, with confidence intervals included in the range of 0.82-0.98, (e) relative rotation angle reliability ranges were L1-L2, .84-.94; L2-L3, .80-.85; L3-L4, .78-.89; L4-L5, .87-.92. Interexaminer reliabilities for the three examiners ranged from .66-.98. CONCLUSION: With the exception of the arcuate angle measurement, the reliabilities for all other measurements were at least .78. Those measurements with reliabilities approaching .80 or better would be considered accurate enough for use in future clinical studies. The arcuate angle measurement may have been least reliable because of the subjective nature of the method of affixing a best-fit line to a radiographic landmark that often takes on the appearance of a mild curvature. Establishing reliability is an important first step toward evaluating these and other similar radiographic measurements that have yet to be examined for their validity.

Analysis of Variance↗

Reliability of nerve conduction studies among active workers.

Nerve conduction studies play an important role in clinical practice and research. Given their widespread use, reliability of tests merits careful attention. We assessed interexaminer and intraexaminer reliability of median and ulnar sensory nerve measures of amplitude, onset latency, and peak latency. In a two-phase cross-sectional study, two examiners tested 158 workers. Reliability was assessed with intraclass correlations (ICC) and kappa statistics. Median nerve measures were more reliable (ICC range, 0.76 to 0.92) than ulnar measures (ICC range, 0.22 to 0.85). Ulnar-onset latencies had the worst reliability. The median-ulnar peak latency difference was a particularly stable measure (ICC range, 0.79 to 0.92). The median-ulnar peak latency difference had high interexaminer reliability (kappa range, 0.71 to 0.79) for normal tests defined by cut points of 0.8 ms and 0.5 ms. Intraexaminer reliability was higher with the 0.8-ms cut point (kappa = 0.90 and kappa = 0.85 for examiners 1 and 2, respectively). Rather than absolute cut points to describe normality, a more rational interpretation of results can be made with ordered categories or continuous measures.

Adult↗

Validity and test-retest reliability of a disability questionnaire for essential tremor.

BACKGROUND: One important outcome in clinical trials is patients' own opinions about whether the medication alleviates their symptoms and improves their ability to function. A valid and reliable method with which to assess this subjective information is important. OBJECTIVE: To determine the validity and test-retest reliability of the Columbia University Disability Questionnaire for Essential Tremor (ET). METHODS: Patients with ET underwent a 2.5-hour evaluation, including a 36-item tremor disability questionnaire, to assess the functional impact of tremor, a 26-item videotaped tremor examination rated by a neurologist, a 15-item performance-based test, and quantitative computerized tremor analysis. We determined the validity and test-retest reliability of the tremor disability questionnaire. Correlations between variables were assessed using Pearson's correlation coefficients and test-retest reliability with the weighted kappa statistic. RESULTS: Ninety-five patients with ET participated. The score on tremor disability questionnaire correlated with the neurologist's clinical ratings (r = 0.57, p <0.001) and the total score on the performance-based test (r = 0.69, p < 0.001). Correlations with quantitative computerized tremor analysis results were less robust, but each remained significant, including mean amplitude of dominant arm tremor while arms were extended (r = 0.56, p <0.001), while drawing a spiral (r = 0.42, p = 0.01), and while pouring (r = 0.34, p = 0.04). The questionnaire was readministered to 32 subjects, and the test-retest reliability was substantial (weighted kappa = 0.67). CONCLUSIONS: This Tremor Disability Questionnaire demonstrated substantial reliability, and it correlated with multiple measures of tremor severity, including a neurologist's clinical ratings, a performance-based test of function, and quantitative computerized tremor analysis results. The questionnaire would be useful in clinical trials in which it could be used as a reliable and valid tool to assess disability in ET.

Activities of Daily Living↗

Test-retest reliability of the Upper Extremity Questionnaire among keyboard operators.

BACKGROUND: Questionnaires are often used in research among workers although few have been tested in the working population. The Upper Extremity Questionnaire is a self-administered questionnaire designed for epidemiological studies and tested among workers. This study assessed reliability of the instrument. METHODS: A two-part assessment was conducted among 138 keyboard operators as part of a large medical survey. Test-retest reliability was analyzed using the kappa statistic, paired t-test, and intraclass correlation coefficient (ICC). Logistic regression models were used to test the effect of demographic and work-related factors on reliability. RESULTS: The average respondent was a white woman, age 35 years, with some college education, in permanent employment with tenure of 1.4 years. Overall, reports of symptoms were stable from Round 1 to 2. Most kappa values for symptom reports were between 0.60 and 0.89. Kappa values for right and left hand diagrams were 0.57 and 0.28, respectively. Among psychosocial items, Perceived Stress and Job Dissatisfaction Scales were most reliable (ICC = 0.88); co-worker support was least reliable (ICC = 0.44). CONCLUSION: Reliability of items on the Upper Extremity Questionnaire were generally good to excellent. Reports of symptom severity and interference with work were less stable. Demographic and work-related factors were not statistically significant in modeling the variation in reliability. Repeated use of the questionnaire with similar results suggests findings are applicable to a larger working population.

Adult↗