PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Reliability of isokinetic dynamometry in assessing plantarflexion torque following Achilles tendon rupture.

BACKGROUND: Research investigating the most effective management of Achilles tendon injury has been limited by a lack of reliable outcome measurements. Calf strength may be a valid outcome measure, not only in terms of identifying possible risk factors for reoccurrence of rupture, but also as an indicator of recovery. Isokinetic dynamometry has been suggested as an effective tool for measuring the torque of the calf muscles. Such measurements have demonstrated high reliability for the assessment of calf muscle torque in healthy subjects. However, there are no published data to demonstrate the reliability of isokinetic dynamometry in subjects with pathology in the Achilles tendon. The purpose of this study was to assess the inter- and intraobserver reliability of isokinetic dynamometry for assessing plantarflexion torque following Achilles tendon rupture. METHODS: Two independent observers used the Kin-com Dynamometer to measure the torque of the plantarflexors in 22 subjects 6 months after unilateral rupture of the Achilles tendon. Twelve subjects had been managed operatively and 10 nonoperatively. Subjects were placed in the prone position with the knee extended. Measurements of peak torque, average torque, and total work were made for both concentric and eccentric plantarflexion movements at 60 degrees per second. RESULTS: Intraclass correlation coefficients were used to calculate reliability of measurements both within and between observers. Reliability was slightly greater on the healthy side (0.74-0.92 ICC) in comparison with the injured side (0.74-0.89 ICC). CONCLUSION: The results of this study suggest that isokinetic dynamometry provides a reliable method of measuring the torque of the plantarflexors following rupture of the Achilles tendon, with levels of reliability comparable with those from healthy subjects. The study concludes that this would be a valuable and reliable outcome measure for use in clinical trials.

Achilles Tendon↗

Reliability of treadmill testing in peripheral arterial disease: a comparison of a constant load with a graded load treadmill protocol.

This study aims to evaluate the reliability of repeated graded workload treadmill testing (G-test; 2 mph; 0% grade, increasing 2% every 2 min) and to compare the reliability of a constant workload treadmill protocol (C-test; 2 mph; 12% grade) versus the graded workload treadmill protocol in patients with intermittent claudication, studied longitudinally. A clinical trial investigating an orally stable prostacycline derivative that included 330 patients with intermittent claudication was performed. The trial employed three active treatment groups and one placebo group. Because there were no significant inter-group differences at baseline or after treatment, data from all groups were pooled for the evaluation of treadmill test reliability. Treadmill data were obtained from a 2-week run-in phase where three G-tests were performed, as well as from the beginning and the end of a 3-month double-blind phase where a G-test and a C-test were performed in random order. Treadmill test reliability was described through test process-related and between-subject variances and also using variance-derived parameters such as the reliability coefficient (RC) and the relative precision (RP). A higher value for the RC and a lower value for the RP indicate that the test variability is predominantly due to between-subject variance and not to test process-related variance. Estimates of variance were described for both the maximal or absolute claudication distance (ACD) and the initial claudication distance (ICD) with each treadmill test. Reliability estimates are reported for the total study sample and for patients with baseline claudication distances < or =300 feet and >300 feet (approximately < or =100 m; >100 m), as measured with the C-test. The cut-off value was empirically chosen to separate severely diseased from mild to moderately diseased claudicants. Theoretical considerations suggest that reliability measures may differ in these subgroups. With repeated testing during the run-in phase for the measure of ACD, the G-test had an RC of 0.952 and an RP of 21.9%. With the comparison of both test protocols in the entire study population for the measurement of ACD, the G-test had an RC of 0.902 and an RP of 31.3%, while the C-test had an RC of 0.876 and an RP of 35.2%. The results for ICD on the G-test were an RC of 0.809 and an RP of 43.7%, while the C-test had an RC of 0.737 and an RP of 51.3%. The reliability of the ACD measurement for RC and RP was numerically superior to those for the ICD for both protocols. In patients with a baseline ACD < or =300 feet, the RC for ACD on the G-test was 0.827 and the RP was 41.4%. In contrast, on the C-test the RC decreased to 0.250 and the RP increased to 86.6%. These changes in RC and RP were due to a marked decrease in the between-subject variance, demonstrating the inability of the C-test to separate appropriately the different claudication distances in populations with highly limited baseline claudication distances. During a run-in phase, the G-test has excellent test characteristics. During the longitudinal phase of a trial, the reliability of G-tests and C-tests are comparable in the entire study population. However, in patients with low claudication distances, the G-test should be given preference over the C-test.

Adult↗

The role of measurement reliability in clinical trials.

One of the principal characteristics of an outcome measure in a clinical trial, and any measurement in general, is its reliability. Reliability refers to the reproducibility of the measurement when repeated at random in the same subject or specimen. Reliability is often confused with validity, which refers to the extent to which the variable properly measures the underlying trait of interest. The coefficient of reliability is an estimate of the proportion of all variation that is not due to measurement error and is readily estimated from replicate measurements. The reliability of a measurement determines its maximal correlation or R2 and slope (or effect size) in regression models, its sensitivity and specificity when used for classifications or predictions, and the power of a statistical test employing the measurement. All decline as the reliability of the measure declines. The reliability of a measurement is an important consideration in the choice of the primary outcome measure for a clinical trial and in the choice of measures used for assessment of eligibility and exclusion. Reliability of measures should be assessed and assured by a quality control program based on randomly selected duplicate assessments. Just as the power of a study is reported in a final publication, so also should the reliability of the outcome and eligibility measurements so as to allow the authors to better describe, and readers to better understand, the sources of imprecision in study results, and those who follow to improve the design of future trials.

Clinical Trials as Topic↗

Reliability of the Camberwell Assessment of Need--European Version. EPSILON Study 6. European Psychiatric Services: Inputs Linked to Outcome Domains and Needs.

BACKGROUND: The five-country European Psychiatric Services: Inputs Linked to Outcome Domains and Needs (EPSILON) Study aimed to develop standardised and reliable outcome instruments for people with schizophrenia. This paper reports reliability findings for the Camberwell Assessment of Need--European Version (CAN-EU). METHOD: The CAN-EU was administered in each country, at two points in time to assess test-retest reliability, and was rated by two interviewers at the first administration. Cronbach's alpha, test-retest reliability and interrater reliability were compared between the five sites. Reliability coefficients and standard errors of measurement for summary scores were estimated. RESULTS: Sites varied in levels and spread of needs. Alphas were 0.48, 0.58 and 0.64 for total, met and unmet needs respectively. Test-retest reliability estimates, pooled over sites, were 0.85 for the total needs, 0.69 for met needs and 0.78 for unmet needs. Pooled estimates for interrater reliability were higher, at 0.94, 0.85 and 0.79 for total, met and unmet needs respectively. There were statistically significant differences in interrater reliability between sites. CONCLUSION: The results confirm the feasibility of using CAN-EU across sites in Europe and its psychometric adequacy.

Adolescent↗

The test-retest reliability of standardized instruments among homeless persons with substance use disorders.

OBJECTIVE: Standardized instruments are widely used to assess homeless persons, but basic data on their reliability and validity in these populations have not been available. The purpose of this study was to examine the reliability of standardized instruments used in a cooperative agreement on homeless persons with substance use disorder. METHOD: This study examined the 1-week test-retest reliability of the Alcohol Dependence Scale, the Addiction Severity Index and the Personal History Form, using 189 randomly selected subjects participating in a multisite study of services for homeless persons with alcohol and other drug abuse problems. In addition to scales and items, factors hypothesized to influence reliability related so subject, interviewer and setting were examined. RESULTS: Results showed substantial reliability for scale scores (> .60) but mixed reliability for individual items. Reliability was greater when items were factual and based on a recent time interval, and when subjects were interviewed in a protected setting. Higher reliability was also related to younger age, female gender, a first episode of homelessness and lower severity of psychiatric problems. CONCLUSIONS: Reliability should be examined in individual studies of homeless persons, and efforts should be made to minimize controllable sources of unreliability.

Adult↗

The reliability of the Brazilian version of the Composite International Diagnostic Interview (CIDI 2.1).

The objective of the present study was to determine the reliability of the Brazilian version of the Composite International Diagnostic Interview 2.1 (CIDI 2.1) in clinical psychiatry. The CIDI 2.1 was translated into Portuguese using WHO guidelines and reliability was studied using the inter-rater reliability method. The study sample consisted of 186 subjects from psychiatric hospitals and clinics, primary care centers and community services. The interviewers consisted of a group of 13 lay and three non-lay interviewers submitted to the CIDI training. The average interview time was 2 h and 30 min. General reliability ranged from kappa 0.50 to 1. For lifetime diagnoses the reliability ranged from kappa 0.77 (Bipolar Affective Disorder) to 1 (Substance-Related Disorder, Alcohol-Related Disorder, Eating Disorders). Previous year reliability ranged from kappa 0.66 (Obsessive-Compulsive Disorder) to 1 (Dissociative Disorders, Maniac Disorders, Eating Disorders). The poorest reliability rate was found for Mild Depressive Episode (kappa = 0.50) during the previous year. Training proved to be a fundamental factor for maintaining good reliability. Technical knowledge of the questionnaire compensated for the lack of psychiatric knowledge of the lay personnel. Inter-rater reliability was good to excellent for persons in psychiatric practice.

Adolescent↗

Reliability of radiographic grading of osteoarthritis of the hip and knee.

We review studies on the reliability of radiographic assessment of osteoarthritis of the hip and knee. Reliability studies were reported for 10 among 24 identified scores. In general, moderate to good agreement was found for overall scores and for separate grading of joint space narrowing of the hip and osteophytes of the knee in the majority of studies, while reliability tended to be lower for other radiographic features. Overall scores of the knee were more reliable than overall scores of the hip, and intra-rater-reliability was considerably higher than inter-rater-reliability in most instances. Comparison of reliability between scores can only be made with caution, given the difference in the design of reliability studies, particularly the different qualification of involved observers. The limits of existing knowledge on reliability of commonly used radiologic scores are outlined, and, proposals are made to overcome those limits in future studies.

Adult↗

The reliability of goniometric measurements of active and passive wrist motions.

A reliability study was conducted to determine (a) the intrarater and interrater reliability of goniometric measurement of active and passive wrist motions under clinical conditions and (b) the effect of a therapist's specialization on the reliability of measurement. Randomly paired therapists performed repeated measurements of active and passive wrist motions in 48 subjects who had been referred to one of four occupational therapy or hand management clinics for evaluation and treatment. The data were analyzed with an intraclass correlation coefficient. A posteriori data analyses were performed to determine the effects of identified sources of error on the reliability of measurement. The results indicated that measurement of wrist motion by individual therapists is highly reliable and that intrarater reliability is higher than interrater reliability for all active and passive motions. Interrater reliability was generally higher among specialized therapists for reasons not immediately apparent from this study. With the exception of pain, identified sources of error were found to have surprisingly little effect on the reliability of measurement.

Adult↗

Test-retest reliability of the ulnar F-wave minimum latency versus ulnar distal motor latency in healthy adults.

BACKGROUND AND PURPOSE: The purposes of this study were to explore reliability of the ulnar F-wave minimum latency (Fmin) and the ulnar distal motor latency (DML) and to contrast those levels of reliability in order to reveal whether physiologic lability is the primary contributor to unwanted variability in Fmin measurements. SUBJECTS AND METHODS: Fmin and DML in the Abductor Digiti Minimi muscle were measured bilaterally by two raters in 50 healthy adults (n = 100 hands, 70 male, 30 female) with 3-14 days between testing sessions. RESULTS: Intrarater reliability (ICC 3,1) for the Fmin was 0.89 with a standard error of the measurement (SEM) of 0.77 msec. Interrater reliability (ICC 2,1) for the Fmin was 0.80 with a SEM of 1.04 msec. Intrarater reliability (ICC 3,1) for the DML was 0.71 with a SEM of 0.18 msec. Interrater reliability (ICC 2,1) for the DML was 0.76 with a SEM of 0.19 msec. DISCUSSION AND CONCLUSIONS: Contrary to our hypothesis, the Fmin had a higher reliability than the DML. The DML did not display the high reliability other investigators have reported. We conclude the Fmin is a reliable measurement when 10 supramaximal stimulations are administered to healthy, young to middle-aged adult subjects. However, no inferences were made regarding relative levels of psychologic lability for the two latencies.

Adult↗

Reliability and validity of measures from the Behavioral Risk Factor Surveillance System (BRFSS).

OBJECTIVES: To assess the reliability and validity of measures on the BRFSS, to assist users in evaluating the quality of BRFSS data, and to identify areas for further research. METHODS: Review and summary of reliability and validity studies of measures on the BRFSS and studies of measures that were the same or similar to those on the BRFSS from other surveys. RESULTS: Measures determined to be of high reliability and high validity were current smoker, blood pressure screening, height, weight, and BMI, and several demographic characteristics. Measures of both moderate reliability and validity included when last mammography was received, clinical breast exam, sedentary lifestyle, intense leisure-time physical activity, and fruit and vegetable consumption. Few measures were of low validity and only one measure was determined to be of low reliability. Several other measures were of high or moderate reliability or validity, but not both. The reliability or validity could not be determined for some measures, primarily due to lack of research. CONCLUSIONS: Most questions on the core BRFSS instrument were at least moderately reliable and valid, and many were highly reliable and valid. Additional research is needed for some measures.

Adult↗

Reliability of work-related assessments.

Insufficient evidence of the reliability of work-related assessments is a major concern in this area of practice. Despite this concern there has been ongoing development of new assessments, while existing assessments have been revised, modified and updated and others are no longer used or available. OBJECTIVES: The purpose of this study was to determine the extent and quality of evidence for the reliability of work-related assessments. STUDY DESIGN: This study examined available literature and sources in order to review the extent which reliability has been established for 28 work-related assessments. RESULTS: The levels of evidence and reliability are presented for each assessment. This indicates that a number of commercially available work-related assessments have insufficient evidence of reliability. For the limited number of work-related assessments with an adequate level of evidence on which to judge their reliability, most demonstrate a moderate to good level. Few assessments, however, have demonstrated levels of reliability sufficient for clinical (and legal) purposes. CONCLUSION: With this study clinicians will be able to examine their options with regard to the reliability of the assessments they choose to use. Interpretation of changes in test results can be considered in the light of the evidence for the reliability of the instrument used.

Journal Article↗

OMERACT Rheumatoid Arthritis Magnetic Resonance Imaging Studies. Exercise 5: an international multicenter reliability study using computerized MRI erosion volume measurements.

Scoring erosions on magnetic resonance imaging (MRI) is one method of estimating damage in patients with rheumatoid arthritis (RA), but it has limitations. The aim of this pilot study was to assess the feasibility and inter-reader reliability of computer assisted erosion volume estimation in patients with RA. Intra-reader and inter-occasion reliability was also assessed, and different slice thicknesses were compared in terms of erosion volume estimation. A 3 mm slice thickness 3D gradient-echo sequence followed by a 1 mm sequence was performed at baseline and repeated within 24 h with metacarpophalangeal (MCP) joints 2 to 5 of the dominant hand included in the field of view. Three readers were instructed to grade MCP 2 and 3 using the OMERACT grading system and then to measure the erosion volume of the same joints using OSIRIS software. The inter-reader reliability of the grading method and the volume method was calculated, as well as the inter-occasion reliability, by comparing results from each reader from baseline to the followup scan. One reader performed repeat volume measurements on 5 patients to assess the intra-reader reliability. Five patients were included in the study. Expressed in terms of intraclass correlation coefficients (ICC), the inter-reader and inter-occasion reliability of the volume method were comparable to the existing OMERACT scoring system, but large systematic differences in volume estimations were found between readers. The intra-reader reliability was excellent. Good correlation was demonstrated between the total erosion scores and the total erosion volumes. For both erosion volumes and erosion scores, 1 mm and 3 mm acquisitions produced variable results between readers, with no clear pattern of underestimation or overestimation for either slice thickness. The volume estimation method was more time consuming, taking roughly 5 times as long as the scoring method. Computerized MRI erosion volume measurements are feasible, with high intra-observer and inter-occasion reliabilities. Despite high ICC, the inter-observer reliability is not sufficient for multicenter use without prior reader training and calibration. The optimal slice thickness was not determined.

Arthritis, Rheumatoid↗

[The appraisal of reliability and validity of subjective workload assessment technique and NASA-task load index].

OBJECTIVE: To test the reliability and validity of two mental workload assessment scales, i.e. subjective workload assessment technique (SWAT) and NASA task load index (NASA-TLX). METHODS: One thousand two hundred and sixty-eight mental workers were sampled from various kinds of occupations, such as scientific research, education, administration and medicine, etc, with randomized cluster sampling. The re-test reliability, split-half reliability, Cronbach's alpha coefficient and correlation coefficients between item score and total score were adopted to test the reliability. The test of validity included structure validity. RESULTS: The re-test reliability coefficients of these two scales and their items were ranged from 0.516 to 0.753 (P < 0.01), indicating the two scales had good re-test reliability; the split-half reliability of SWAT was 0.645, and its Cronbach's alpha coefficient was more than 0.80, all the correlation coefficients between its items score and total score were more than 0.70; as for NASA-TLX, both the split-half reliability and Cronbach's alpha coefficient were more than 0.80, the correlation coefficients between its items score and total score were all more than 0.60 (P < 0.01) except the item of performance. Both scales had good inner consistency. The Pearson correlation coefficient between the two scales was 0.492 (P < 0.01), implying the results of the two scales had good consistency. Factor analysis showed that the two scales had good structure validity. CONCLUSION: Both SWAT and NASA-TLX have good reliability and validity and may be used as a valid tool to assess mental workload in China after being revised properly.

Adolescent↗

Reliability of first ray position and mobility measurements in experienced and inexperienced examiners.

CONTEXT: Neither reliability nor validity data exist for the Root method of clinically assessing first ray position or mobility by experienced and inexperienced examiners. OBJECTIVE: To determine intrarater and interrater reliability for first ray position and mobility measurements in experienced and inexperienced examiners. DESIGN: Single-blind prospective reliability study. SETTING: Physical therapy clinic. PATIENTS OR OTHER PARTICIPANTS: Four examiners, 2 experienced and 2 inexperienced, obtained first ray position and mobility measurements. Both feet of 36 subjects (14 males, 22 females) were measured. INTERVENTION(S): Each examiner evaluated first ray position and mobility for each of the subjects' feet on 2 separate occasions using the manual assessment techniques described by Root. MAIN OUTCOME MEASURE(S): First ray position (normal, plantar flexed, dorsiflexed) and mobility (normal, hypermobile, hypomobile) decisions were made. RESULTS: We calculated kappa correlation coefficients for intrarater and interrater reliability. For position, intrarater and interrater reliability ranged from .03 to .27 for all examiners, experienced and inexperienced. For mobility, intrarater and interrater reliability ranged from .02 to .26 for experienced, inexperienced, and experienced/inexperienced. The percentage agreement (P(O)) values for all examiners were less than 58%. For individual values for position, intrarater and interrater reliability ranged from .00 to .26. For individual values for mobility, intrarater and interrater reliability ranged from .00 to .26. The P(O) values for all examiners were less than 50%. CONCLUSIONS: Clinical experience was not associated with higher kappa coefficients or P(O) values when examiners assessed first ray position or mobility. Clinicians should acknowledge the poor reliability of first ray measurements, especially when making treatment decisions. Finally, a validity study to compare the Root techniques with a gold standard is warranted.

Journal Article↗

Reliability of clinical findings in temporomandibular disorders.

The aim of the present investigation was to study the interexaminer reliability of orthopedic tests and palpation techniques routinely used in the clinical diagnosis of disorders of the masticatory system. The tests were performed by a dentist and a physiotherapist, who both used the tests routinely when examining patients with temporomandibular disorders. Seventy-nine patients participated in this study. In the analysis, percentage agreement, intraclass correlation, and Cohen's kappa were used. The interexaminer reliability of the tests measuring maximal active mouth opening and registration of clicking during active mouth opening was high. The interexaminer reliability was fair for the tests measuring the intensity of pain during active movements and moderate for tests recording joint sounds (kappa = 0.47 to 0.59). There was high interobserver agreement on several items of the traction and translation tests, although the kappa values were low. The interexaminer reliability of the multitest scores for compression was substantial for joint sounds (kappa = 0.66) and fair for pain (kappa = 0.40). The interexaminer reliability of the multitest scores for muscle palpation and joint palpation was moderate (kappa = 0.51) and fair (kappa = 0.33), respectively. It can be concluded that most variables determined during active movements can be measured with satisfactory reliability, whereas variables for other tests are not measured with the same reliability on the basis of the kappa scores. The main symptoms of temporomandibular disorders can be evaluated reliably with multitest scores. It is recommended that clinicians calibrate their techniques regularly to improve the reliability of results in daily practice.

Adolescent↗

Interexaminer reliability in physical examination of the neck.

BACKGROUND: There are numerous clinical tests used in the evaluation of patients with symptoms arising from the cervical spine. It is necessary to use clinical tests with high validity and reliability. Previous studies of reliability of clinical tests used in the evaluation of the cervical spine have come to various conclusions, most of which suggest low reliability. This might be explained by differences between examiners in performance and where the limit of normality is placed. OBJECTIVE: To evaluate the interexaminer reliability of clinical tests used in everyday clinical work, where the examiners base their evaluations on a comparison between left and right sides. STUDY DESIGN: A total of 50 volunteers were examined by two physiotherapists. The interexaminer reliability of clinical tests included in the physical examination of patients with symptoms from the cervical spine was evaluated. METHODS: Two physiotherapists independently examined volunteers. RESULTS: An acceptable reliability was found for two of 10 clinical tests. CONCLUSION: When it is possible to compare left and right sides, it is possible to show acceptable reliability for some clinical tests. Reliability studies most often find low reliability, perhaps because of bias; clinical tests are not standardized. In future studies, greater efforts should be taken to reduce bias.

Adolescent↗

Reliability, Device Agreement and Validity of Load-Velocity Profiles: A Systematic Review with Meta-analysis.

BACKGROUND: For a valid one-repetition maximum (1RM) prediction via load-velocity (LV) relationships, high reliability and accuracy must be assumed. OBJECTIVE: Since individual study results indicate ambivalent prediction, this systematic review and meta-analysis was designed to provide a updated and comprehensive overview, extending knowledge about the validity and reliability of commercially available velocity sensors in Part I and the validity and reliability of velocity-based 1RM prediction models in Part II. METHODS: A systematic literature search was conducted in PubMed/MEDLINE, Web of Science, and Scopus. Validity and/or reliability studies or velocity-based 1RM prediction evaluations were included. Methodological quality was assessed using adapted COSMIN. The analysis was performed for intraclass correlation coefficient (ICC), Lin's concordance correlation coefficient (CCC), and Pearson's correlation coefficient (r). The review was preregistered in PROSPERO (CRD42025634595). RESULTS: Sixty-three studies were included for sensor validity and reliability and 38 for 1RM prediction models. Part I: Velocity sensors demonstrated good-to-excellent pooled validity and device agreement (ICC&#x2009;=&#x2009;0.91-0.92 [0.83-0.97]; k&#x2009;=&#x2009;55 and 439, respectively); intra- and inter-day reliability were classified as good to excellent with ICC&#x2009;=&#x2009;0.90-0.91 [0.85-0.95] (k&#x2009;=&#x2009;228 and 608, respectively), with sensor technology moderating the results. However, substantial heterogeneity and wide ranges of study-level estimates indicated considerable variability across moderators, linear position transducer (LPT) generally showing more consistent performance than inertial measurement units (IMU). Part II: Velocity-based 1RM prediction showed ICCs&#x2009;=&#x2009;0.90 [0.83-0.94] (k&#x2009;=&#x2009;124) and ICC&#x2009;=&#x2009;0.91 [0.72-0.98] (k&#x2009;=&#x2009;9); for reliability and validity, respectively. DISCUSSION: Commercial velocity sensors generally provide high relative validity and reliability. Results varied depending on exercise complexity, intensity, sensor technology, and modeling approach. While velocity-based 1RM prediction demonstrated high average validity, large heterogeneity in lower body exercises significantly biased the results. Furthermore, the dearth of measurement error and agreement analyses prohibits final conclusions. CONCLUSION: Therefore, velocity-based monitoring and 1RM prediction require cautious interpretation, as sensor- and exercise-specific evidence remains limited.

Load&#x2013;velocity relationship↗

Reliability of physical examination of the upper extremity among keyboard operators.

BACKGROUND: Physical examination is a traditional outcome measure in epidemiological research. Its value as a reliable measure depends, in part, on the prevalence of positive findings. The purpose of this paper is to determine the empirical reliability of physical examination and anthropometry in a field study of upper extremity disorders among keyboard operators. METHODS: Two experienced examiners independently performed common provocative tests and procedures in physical examinations of the neck and upper extremity among 160 keyboard operators. Two additional examiners conducted anthropometric surveys among 137 workers. Inter-examiner reliability was assessed with observed agreement, kappa statistics, and intra-class correlations (ICC). RESULTS: Observed agreement was between 96% and 100% for neck and upper extremity signs, muscle stretch reflexes, and muscle strength, however, with the exception of provocative tests, reliability statistics were unstable. Among the provocative tests, Phalen and Tinel tests had modest agreement after adjusting for chance (kappa range: 0.20-0.43). The carpal compression test had the best reliability (kappa=0.60 and kappa=0.67, left and right side, respectively). The ICCs for anthropometry ranged from 0.36-0.91. CONCLUSIONS: Results from the study showed that statistically, except for the carpal compression test, physical examination contributed minimal reliable information. This was attributed mainly to the low prevalence of positive findings, and generally mild nature of upper extremity disorders in this population. The results are the best estimate of what would be found in a field study with experienced examiners. While it may reduce bias, separating physical examination from medical history may contribute to the poor reliability of findings. With a shift toward reliable measures, resources can be allocated to more effective tools, like questionnaires, in epidemiological research of upper extremity disorders among keyboard operators.

Adult↗