PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Reliability of performance-based clinical skill assessment of emergency medicine residents.

OBJECTIVE: To test the overall reliability of a performance-based clinical skill assessment for entering emergency medicine (EM) residents. Also, to investigate the reliability of separate reporting of diagnostic and management scores for a standardized patient case, subjective scoring of patient notes, and interstation exercise scores. METHODS: Thirty-four first-year EM residents were tested using a 10-station standardized patient (SP) examination. Following each 10-minute encounter, the residents completed a patient note that included differential diagnosis and management. The residents also were asked to read an ECG or chest x-ray (CXR) associated with each case. History, physical examination, and interpersonal skills were scored by the SPs. The patient note, CXR, and ECG readings were scored by faculty emergency physicians. Intercase reliability was determined for the residents. RESULTS: Global score reliability, Cronbach's alpha = 0.85. Reliabilities for the other components were: history, 0.77; physical examination, 0.83; and interpersonal skills, 0.80. Differential diagnosis and management reliabilities were 0.61 and 0.66, respectively. Subjective scoring of the patient note resulted in acceptable reliability for legibility (0.80), history completeness (0.80), and history organization (0.81). Physical examination completeness and organization reliabilities were 0.74 and 0.73. For ECG and CXR readings, alpha = 0.74 and 0.34, respectively. CONCLUSIONS: SPs can be used to reliably assess bedside clinical skills of EM residents. While component reliability levels are slightly lower than the global clinical skill reliability coefficient, they are still high enough to use for identification of individual strengths and weaknesses.

Clinical Competence↗

Reliability of two goniometric methods of measuring active inversion and eversion range of motion at the ankle.

BACKGROUND: Active inversion and eversion ankle range of motion (ROM) is widely used to evaluate treatment effect, however the error associated with the available measurement protocols is unknown. This study aimed to establish the reliability of goniometry as used in clinical practice. METHODS: 30 subjects (60 ankles) with a wide variety of ankle conditions participated in this study. Three observers, with different skill levels, measured active inversion and eversion ankle ROM three times on each of two days. Measurements were performed with subjects positioned (a) sitting and (b) prone. Intra-class correlation coefficients (ICC[2,1]) were calculated to determine intra- and inter-observer reliability. RESULTS: Within session intra-observer reliability ranged from ICC[2,1] 0.82 to 0.96 and between session intra-observer reliability ranged from ICC[2,1] 0.42 to 0.80. Reliability was similar for the sitting and the prone positions, however, between sessions, inversion measurements were more reliable than eversion measurements. Within session inter-observer measurements in sitting were more reliable than in prone and inversion measurements were more reliable than eversion measurements. CONCLUSION: Our findings show that ankle inversion and eversion ROM can be measured with high to very high reliability by the same observer within sessions and with low to moderate reliability by different observers within a session. The reliability of measures made by the same observer between sessions varies depending on the direction, being low to moderate for eversion measurements and moderate to high for inversion measurements in both positions.

Adult↗

Between-days reliability of electromyographic measures of paraspinal muscle fatigue at 40, 50 and 60% levels of maximal voluntary contractile force.

OBJECTIVE: To ascertain which percentage of maximal voluntary contractile force of the paraspinal muscles, when tested in a functional position, is most reliable for assessing electromyographic (EMG) fatigue changes. SUBJECTS: Ten healthy volunteers with no history of low back pain (six males). MAIN OUTCOME MEASURES: The surface EMG signal during 60-second isometric contractions of the paraspinal muscles at 40, 50 and 60% levels of maximal voluntary contractile force was captured and analysed. Each contraction level was assessed on two occasions, at least three days apart. The initial median frequency, the decline in median frequency slope and the increase in root mean square values were assessed for between-days reliability, using intraclass correlation coefficients (ICCs) and standard errors of measurements (SEM). Normalized median frequency and root mean square values were also assessed. RESULTS: At 40% of maximal voluntary contraction, little or no EMG fatigue changes occurred in any of the observed parameters. At 50% maximal voluntary contraction the initial mean frequency and root mean square changes proved highly reliable, with ICCs ranging from 0.74 to 0.86 and 0.75 to 1.00 respectively. Normalizing the root mean square data reduced the reliability, but this was still acceptable with ICCs 0.70-0.83. The median frequency decline slope proved less reliable with ICCs 0.24-0.74 for raw and 0.26-0.77 for normalized data. At 60% maximal voluntary contraction the initial mean frequency proved as reliable as initial median frequency at 50% with ICCs 0.70-0.89. The raw and normalized root mean squares (ICCs 0.43-0.89 and 0.30-0.87 respectively) and raw and normalized median frequency (ICCs 0.27-0.51 and 0.24-0.53 respectively) changes were less reliable than at 50% MVC. Overall, the reliability is better at the L4/5 than at the L2/3 level. CONCLUSION: Outcome measures taken at 50% maximal voluntary contraction are the most reliable in functional testing the paraspinal muscles of healthy volunteers. With initial median frequency and root mean square values being more reliable parameters than median frequency decline. At the L4/5 level, however, all parameters were acceptably reliable at 50% of maximum effort. However the between-subject variability of the median frequency decline and root mean square incline slopes suggest that these parameters are not yet fully suitable for monitoring fatigue changes during prolonged isometric contraction.

Adult↗

Intraobserver and interobserver reliability of the classification of thoracic adolescent idiopathic scoliosis.

The system described by King et al. is the standard method for the classification of thoracic adolescent idiopathic scoliosis. Although it is widely used and referenced, its reliability and reproducibility among scoliosis surgeons are unknown. We used a scoliosis case-presentation format to examine the interobserver and intraobserver reliability of the classification of thoracic adolescent idiopathic scoliosis with the system of King et al. Eight active, current members of the Scoliosis Research Society reviewed twenty-seven full-length radiographs that had been made before operative correction of the scoliotic deformity. On the basis of these images, which included posteroanterior and lateral radiographs made with the patient standing as well as right and left forced-side-bending radiographs made with the patient supine, the reviewers assigned a type to each curve according to the classification system of King et al. Kappa coefficients were used to test statistical reliability. The mean interobserver reliability of the classification was only 64 per cent (range, 54 to 77 per cent) when the responses of seven of the reviewers were compared with those of one of the originators of the classification. The mean kappa coefficient was 0.49 (range, 0.27 to 0.73), which indicates poor reliability. When each reviewer's responses were compared with those of the other reviewers, the reliability was similarly poor (interobserver reliability, 55 per cent [range, 33 to 81 per cent] and mean kappa coefficient, 0.40 [range, 0.21 to 0.63]). Intraobserver reliability was evaluated in a trial in which five reviewers in a group setting were shown the same radiographs in a different order at two different viewings. Comparison of the results at the two viewings revealed a mean intraobserver reliability of 69 per cent (range, 56 to 85 per cent) and a mean kappa coefficient of 0.62 (range, 0.34 to 0.95), which indicates fair reliability. The current method of classification of adolescent idiopathic scoliosis does not appear to have sufficient intraobserver or interobserver reliability among scoliosis surgeons to portray curve types accurately. Thus, it may not help to guide treatment with use of modern spinal fixation methods.

Adolescent↗

Intrasession and intersession reliability of the soleus H-reflex in supine and standing positions.

The Hoffmann reflex (H-reflex) is a measure of motoneuron pool excitability, which is valuable in determining muscle inhibition caused by joint damage (arthrogenic muscle inhibition). In order to detect changes in H-reflex due to injury, the reliability of such a measurement must be established. The purpose of this study was to establish the intrasession and intersession reliability of soleus H-reflex in a supine and standing position. Thirteen healthy volunteers (age 10 +/- 2.63 yr, height 171.35 +/- 10.19 cm, mass 69.62 +/- 13.03 Kg) with no lower extremity orthopedic or neurological disorders within the past year participated in this study. To determine the intrasession and intersession reliability of this measure in a supine resting position and a one-leg standing position, EMG data were collected from the soleus while the tibial nerve was stimulated in the popliteal space. A high voltage (120-200 V), short duration (1.0 msec) stimulus was automatically triggered, eliciting a reflex twitch detected by surface EMG. Several of these measurements were performed with 20 second rest intervals to find the maximum H-reflex. The maximum H-reflex was located by adjusting the intensity of the stimulus. Once a maximum H-reflex was found, 12 measurements were taken in that position with 20 second rest intervals. These steps were repeated for each position (supine and standing) at the same time for 5 consecutive days. Intrasession reliability was computed using 12 measurement trials (12), 12 measurement trials dropping the high and low score (12x), the first 7 measurement trials dropping the high and low score (7x), and the first 5 measurement trials (5). Intrasession and intersession reliability over five consecutive days was estimated using intraclass correlation coefficients (ICC (3, 1)). The supine intrasession reliability measurements were as follows: 0.932 (12), 0.932 (12x), 0.935 (7x), and 0.932 (5). The standing intrasession reliability was 0.853 (12), 0.852 (12x), 0.865 (7x), and 0.862 (5). The intersession reliability was 0.938 in the supine position and 0.803 in the standing position. These results indicate that the H-reflex measured using our protocol in a supine and standing position is a reliable assessment within sessions and between sessions. Five measurements are sufficient to observe reliable measurements within a single session. Most importantly, this data shows that the H-reflex is a reliable assessment that may be used to measure small changes in motoneuron pool excitability over time.

Adult↗

An evaluation of ventilator reliability: a multivariate, failure time analysis of 5 common ventilator brands.

INTRODUCTION: Mechanical ventilator failures expose patients to unacceptable risks and are expensive. By identifying factors that correlate with the amount of time between consecutive ventilator failures, we might reduce patient risk, save money, and shed light on a number of important questions concerning whether reliability changes as a function of time. OBJECTIVE: Investigate the correlation between several explanatory variables and the time between consecutive ventilator failures and address the following questions: (1) Are ventilators as safe and reliable following repairs as they were before failing? (2) Does reliability change significantly as a ventilator is used or ages? (3) Does a hospital's particular operating environment play a role in ventilator reliability? (4) Are ventilator service contracts worth the money? METHODS: A retrospective review was conducted using repair and maintenance records from 2 hospitals: a 570-bed teaching hospital and a 410-bed local community hospital. Records were examined from a total of 66 individual ventilators, of 5 different brands, used between July 1, 1991, and January 3, 2001. The ventilators included 13 Tyco-Mallinckrodt Infant Star, 10 Bird VIP, 11 Bird 6400ST, 16 Bird 8400STi, and 16 Tyco-Mallinckrodt 7200ae. The dependent variable was the operating time between or before unexpected mechanical failures; this was determined by the difference between hours logged on the ventilator hour meter at the time of failure and that recorded when the study began, or when the ventilator was new. Thereafter (when applicable), the time before failure was the difference in hours at consecutive failures. Seven independent explanatory covariates were selected and analyzed as potential correlates with time between failures. Another independent variable, the site of ventilator use (community or teaching hospital), was also tested for significance. Data were analyzed using the Cox proportional hazard model, the multiple-groups survival statistic, and the Cox-Mantel test. RESULTS: In 2,567,365 hours of ventilator operation, 290 observations were recorded (226 failures and 64 censored observations). Two of the 7 covariates were judged time-dependent, excluded from the Cox model, and evaluated using other techniques. Of the 5 remaining covariates, 2 were significantly related to reliability, both indirectly. There was no difference in reliability, regardless of how many times a ventilator had been previously repaired, but hospital environment did significantly affect reliability. CONCLUSIONS: Ventilator reliability depends on a number of factors. This study indicates that, on average, ventilator reliability improves the more a ventilator is used and the longer the brand has been commercially available. The number of previous ventilator repairs did not affect reliability, but the hospital environment did. These data, if validated, should help to enhance our understanding of ventilator reliability and could eventually have profound economic and safety implications as well.

Contract Services↗

Intrasession and intersession reliability of the quadriceps Hoffmann reflex.

The Hoffmann reflex (H-reflex) is a resting electromyographic (EMG) measurement of motoneuron pool recruitment. While the soleus H-reflex has been studied extensively, the study of the quadriceps H-reflex has been limited for various reasons. To date no data exist regarding the reliability of this measurement within and between sessions over an extended period of time. The purpose of this study was to establish quadriceps H-reflex reliability over a four-week period consisting of 6 testing sessions. Eleven neurologically sound volunteers (age: 20 +/- 2 yr, height: 181.9 +/- 9.9 cm; mass: 84.2 +/- 17.8 Kg) participated in this study. Subjects were prepared for EMG surface electrodes over the vastus medialis and medial malleolus (ground). A stimulating bar electrode was placed over the femoral nerve, and a stimulus was delivered at 20 sec intervals with increasing amplitude to obtain a peak quadriceps H-reflex. Ten peak H-reflex measurements were recorded for each session. The stimulator amplitude was increased further to obtain the peak efferent motor response (M-response) for normalization of peak H-reflex measurements. The procedure was repeated at 1 hr, 24 hr, 1, 2, and 3 weeks following the initial session. Intraclass correlation coefficients (ICC (2.1) and ICC (3.1)) were calculated using normalized peak H-reflex measurements. Strong reliability was detected within a session (10 trials ICC (2.1) = 0.957, ICC (3.1) = 0.970; 5 trials ICC (2.1) = 0.961, (ICC (3.1) = 0.970). ICC (2.1) calculations yielded strong reliability between the first and second (1 hr) session (0.956) but moderate reliability between days 1 and 2 (0.787) and between all session over 4 weeks (0.756). ICC (3.1) calculations were also computed to determine the reliability for the fixed selection of sessions. This calculation yielded strong reliabilities between days 1 and 2 (0.969) and between all sessions over 4 weeks (0.911). Within a measurement session quadriceps H-reflex measurements are very reliable, and we recommend 5 trials to establish a measurement. Between measurement sessions these data provide evidence of strong reliability for these fixed sessions (ICC (3.1)), and generalized evidence for moderate reliability over a 4-week period has also been provided (ICC (2.1)). These data provide valuable information for the researcher and practitioner as to the reliability of this measurement in detecting changes in the neuromuscular system.

Adult↗

Reliability of refraction--a literature review.

BACKGROUND: The repeatability (reliability) of methods of measuring refractive error is an important consideration in patient management decisions and in research design and interpretation. Several papers on reliability of refraction have appeared in the literature. METHODS: Studies on the reliability of conventional clinical refraction, including repeated refractions by the same examiner (intraexaminer reliability) and comparisons of the results obtained by different examiners (interexaminer reliability), were reviewed. Studies comparing the results obtained by autorefractors and conventional subjective refraction were also reviewed. RESULTS: The intraexaminer reliability and interexaminer reliability of subjective refraction in most studies were close to 80 percent agreement within +/- 0.25D and 95 percent agreement within +/- 0.50D for spherical equivalent, sphere power, and cylinder power. The reliability of most autorefractors is similar to the reliability of conventional subjective refraction. CONCLUSIONS: The results of the reliability studies support Bannon's 1977 conclusion that conventional subjective refraction is reliable within 0.25 to 0.50D. Studies comparing objective autorefractors and conventional subjective refraction indicate that these autorefractors are satisfactory for a preliminary refraction but are not satisfactory as substitutes for conventional subjective refraction.

Humans↗

Infant polysomnography: reliability. Collaborative Home Infant Monitoring Evaluation (CHIME) Steering Committee.

Infant polysomnography (IPSG) is an increasingly important procedure for studying infants with sleep and breathing disorders. Since analyses of these IPSG data are subjective, an equally important issue is the reliability or strength of agreement among scorers (especially among experienced clinicians) of sleep parameters (SP) and sleep states (SS). One basic issue of this problem was examined by proposing and testing the hypothesis that infant SP and SS ratings can be reliably scored at substantial levels of agreement, that is, kappa (kappa) > or = 0.61. In light of the importance of IPSG reliability in the collaborative home infant monitoring evaluation (CHIME) study, a reliability training and evaluation process was developed and implemented. The bases for training on SP and SS scoring were CHIME criteria that were modifications and supplements to Anders, Emde, and Parmelee (10). The kappa statistic was adopted as the method for evaluating reliability between and among scorers. Scorers were three experienced investigators and four trainees. Inter- and intrarater reliabilities for SP codes and SSs were calculated for 408 randomly selected 30-second epochs of nocturnal IPSG recorded at five CHIME clinical sites from healthy full term (n = 5), preterm (n = 4), apnea of infancy (n = 2), and siblings of the sudden infant death syndrome (SIDS) (n = 4) enrolled subjects. Infant PSG data set 1 was scored by both experienced investigators and trained scorers and was used to assess initial interrater reliability. Infant PSG data set 2 was scored twice by the trained scorers and was used to reassess inter-rater reliability and to assess intrarater reliability. The kappa s for SS ranged from 0.45 to 0.58 for data set 1 and represented a moderate level of agreement. Therefore, rater disagreements were reviewed, and the scoring criteria were modified to clarify ambiguities. The kappa s and confidence intervals (CIs) computed for data set 2 yielded substantial inter-rater and intrarater agreements for the four trained scorers; for SS, the kappa = 0.68 and for SP the kappa s ranged from 0.62 to 0.76. Acceptance of the hypothesis supports the conclusion that the IPSG is a reliable source of clinical and research data when supported by significant kappa s and CIs. Reliability can be maximized with strictly detailed scoring guidelines and training.

Humans↗

Reliability of the Eating Disorder Examination in patients with binge eating disorder.

OBJECTIVE: This study examined the interrater and test-retest reliabilities of the Eating Disorder Examination (EDE) in patients with binge eating disorder (BED). METHOD: Interrater reliability and short-term (6-14 days) test-retest reliability of the EDE were examined in two study groups of 18 patients with BED. RESULTS: Interrater reliability was excellent for objective bulimic episodes and days (correlations above .98) and very good for the EDE scales, albeit somewhat variable (correlations range from .65 to .96). Test-retest reliabilities were very good for objective bulimic episodes (.70) and days (.71) and were good (significant) for the EDE scales, albeit somewhat variable (correlations range from .50 to .88). Interrater reliability was excellent for subjective bulimic episodes and days but test-retest reliabilities were unacceptable. CONCLUSIONS: These findings support the reliability of the EDE for patients with BED. The EDE has utility for assessing the number of large binge episodes (objective bulimic episodes), as well as the number of days during which large binge episodes occurred. The EDE also demonstrates very good interrater and test-retest reliabilities for assessing the associated features of eating disorders in patients with BED. The results for subjective bulimic episodes are consistent with previous studies, suggesting that these eating behaviors may not be reliable indicators of eating disorders.

Adult↗

Can the stages of change for smoking acquisition be measured reliably in adolescents?

BACKGROUND: was to examine the reliability of the algorithm. METHODS: As part of a randomized controlled trial, 3,930 adolescents completed a paper version of the algorithm questions and a differently worded computerized version on the same day: a parallel form reliability assessment. In a separate assessment, another group of 118 adolescents completed 2 identical paper versions of the same questionnaire 2 weeks apart: a test-retest reliability assessment. Kappa (kappa) for agreement for stage and the individual questions were calculated. Logistic regression was used to examine whether demographic characteristics, smoking status, and stage predicted agreement for stage. RESULTS: Kappa (95% confidence intervals) for stage was 0.57 (0.55-0.60) in the first assessment and 0.46 (0.28-0.63) in the second assessment, indicating moderate reliability. The question concerning trying smoking in the next 6 months was moderately reliable, but that concerning trying within the next thirty days was poorly reliable. Acquisition precontemplation was significantly more reliably coded than all other stages. Demographic characteristics did not predict reliability. CONCLUSIONS: The algorithm reliably allocates individuals into acquisition precontemplation, but for all other stages, its reliability is fair.

Adolescent↗

Plagiocephalometry: a non-invasive method to quantify asymmetry of the skull; a reliability study.

UNLABELLED: Deformational plagiocephaly (DP) in newborns and very young children is a common problem in daily practice. The intrarater and interrater reliability of plagiocephalometry (PCM), a new, non-invasive, inexpensive instrument to assess and quantify the asymmetry of the skull, is evaluated at the outpatient Department of Physical Therapy of the Bernhoven Hospital at Veghel, The Netherlands. Using a thermoplastic material to mould the outline of the infant's skull, a reproduction of the skull shape is performed on paper, allowing for accurate cephalometric measurements. Fifty children (aged 0-24 months), with or without positional preference of the head, and with or without DP, were measured three times by two separate, experienced pediatric physical therapists. Intraclass correlation coefficients (ICC) regarding the measurements of the drawn lines were all above 0.92 (intrarater reliability) and 0.90 (interrater reliability). The ICCs of the plagiocephaly indicators ear deviation (ED), antero-sinistra-antero-dextra (ASAD), postero-dextra-postero-sinistra (PDPS) and oblique diameter difference (ODD) were 0.88, 0.57, 0.92 and 0.96, respectively, for the intrarater reliability and 0.90, 0.65, 0.94 and 0.96, respectively, for the interrater reliability. The ICCs of the two indices oblique diameter difference index (ODDI) and cranial proportional index (CPI) were 0.97 and 0.96, respectively, for the intrarater reliability and 0.95 and 0.92, respectively, for the interrater reliability. The limits of agreement according to Bland Altman, comprising 95% of the differences between two measurements (2 sd), were 4.3 mm (ED), 5.9 mm (ASAD), 3.0 mm (PDPS), 3.4 mm (ODD), 2.7% (ODDI) and 4.5% (CPI) for the intrarater reliability, and 3.7 mm (ED), 5.2 mm (ASAD), 2.4 mm (PDPS), 3.3 mm (ODD), 2.9% (ODDI) and 5.8% (CPI) for the interrater reliability. CONCLUSION: We conclude that PCM is an easy-to-apply, non-invasive and reliable measurement instrument to assess skull asymmetry with good clinical accuracy and low application costs. PCM might serve as an instrument to be used in all levels of care for children with DP, and might provide information concerning the natural course of DP, as well as the assessment of the effects of conservative treatment strategies on DP.

Cephalometry↗

Inter-examiner reliability in the assessment of low back pain (LBP) using the Kirkaldy-Willis classification (KWC).

Reliable classification systems and clinical tests are sought for the care of patients with low back pain (LBP). The objectives of this clinical study were to evaluate inter-examiner reliability in the classification of patients with LBP, the influence of radiological findings on the classification and the reliability of some clinical tests. Two examiners independently assessed 50 outpatients with LBP. Inter-examiner reliability in classification of patients with LBP using Kirkkaldy-Willis classification (KWC) system and in 30 clinical tests was calculated as percentage agreement and kappa coefficients (kappa). Inter-examiner reliability was excellent (kappa>0.8) for classification according to KWC. Radiological findings did not influence the reliability. Age of the patient, movement range, and pain and neurological signs seemed to guide the decision on classification. The reliability of clinical tests was good (kappa>0.6) in 6 tests and moderate (kappa>0.4) in 12 tests. Good inter-examiner reliability was found for the SLR test, movement range and sensibility testing with spurs in dermatome areas. We conclude that the KWC for classifying patients with LBP seems to be a reliable classification system depending on a few key observations and that moderate and good inter-examiner reliability can be achieved in several clinical tests in the assessment of LBP.

Adolescent↗

Analysis of reliability indices from Humphrey visual field tests in an urban glaucoma population.

PURPOSE: Visual field assessment is extremely important in glaucoma management, but interpretation is affected by the quality of the patient's performance. The authors have investigated the reliability of visual field performance by a randomly selected sample of the chronic glaucoma population at an urban tertiary care practice. METHODS: Patient reliability in Humphrey automated visual field testing was studied in 106 randomly selected chronic open-angle glaucoma patient charts, which provided 768 tests (mean, 7.2 +/- 4.8 fields; range, 2-18 fields). Reliability criteria were established as less than 20% fixation losses, less than 33% false-negative error, and less than 33% false-positive error, as recommended by Humphrey Instruments, Inc (San Leandro, CA). RESULTS: Patients performed reliably in 61% of right eye fields, 58% of left eye fields, and 59.5% overall. Of the 106 patients, only 35 (33%) were always reliable in both eyes, whereas 8 (7.5%) were always unreliable in both eyes. The most common cause of unreliability was fixation loss (39%), whereas false-positive error (5%) and false-negative error (9%) were less frequent. A more severely depressed mean deviation correlated significantly with poorer performance on the three reliability indices, with false-negative error having the greatest correlation, followed by fixation loss and false-positive error. Corrected pattern standard deviation correlated closely only with false-negative error. Prolonged test time also correlated with all three reliability indices. Age was a significant factor for fixation loss but not for false-negative or false-positive error. CONCLUSIONS: The authors conclude that fewer than two thirds of the Humphrey visual fields were reliable with the authors' urban tertiary care population of patients with glaucoma. Relaxing the fixation loss criterion to less than 33% improved the rate of reliability to approximately 75%. The severity of glaucomatous visual field defects, test time, and age were identified as factors influencing the reliability of the Humphrey visual fields.

Aged↗

Reliability of paramedic ratings of laryngoscopic views during endotracheal intubation.

BACKGROUND: Prior studies have related prehospital endotracheal intubation (ETI) difficulty to paramedic visualization of the vocal cords using the Cormack-Lehane (C-L) scale. However, the reliability of paramedic C-L ratings has not been formally studied. OBJECTIVE: To evaluate the reliability of C-L and a more recently described scale, percentage of glottic opening (POGO), when used by paramedics to rate laryngoscopic views during ETI. METHODS: Twenty-five standard slide images of laryngoscopic views were obtained during ETI. The 25 images were duplicated to facilitate evaluation of intrarater agreement (total 50 slides). Seven paramedics rated the degree of vocal cord visualization in each image using C-L (I-IV, ordinal scale; I = full visualization of vocal cords, IV = only epiglottis seen) and POGO (0-100 continuous scale; 0 = no vocal cords seen, 100 = full visualization of vocal cords). We assessed intra- and interrater reliabilities using Cohen's multirater kappa for C-L and intraclass correlation coefficients (ICCs) for POGO. RESULTS: C-L showed variable intrarater reliability (kappa range = 0.37-0.90) and poor interrater reliability (Cohen's multirater kappa = 0.22). POGO demonstrated good to excellent intrarater reliability (one-way random-effects ICC range = 0.57-0.87) and fair to good interrater reliability (two-way random-effects ICC = 0.59, 95% Confidence interval: 0.48-0.71). CONCLUSIONS: Paramedic C-L ratings exhibit poor intra- and interrater reliabilities. Paramedic POGO ratings exhibit fair to good intra- and interrater reliabilities. POGO may be more appropriate than C-L for prehospital clinical and scientific application. Reliability must be formally evaluated for any proposed laryngoscopic exposure classification system.

Allied Health Personnel↗

Goniometric reliability in a clinical setting. Elbow and knee measurements.

Reliability of goniometric measurements has been examined only under standardized conditions and usually with healthy subjects. The purpose of this study was to assess goniometric reliability in a clinical setting. The reliability of goniometric measurements of passive elbow and knee positions was assessed using patients as subjects. The effect of using the means of repeated measurements and the interdevice reliability of three common goniometers were also examined. Results showed that intratester reliability for flexion and extension of the knee and the elbow joints was high (r = .91 to .99). Intertester reliability was also high (r = .88 to .97) for these measurements except for measurements of knee extension (r = .63 to .70). Although previous investigators have suggested that using the means of multiple measurements improves reliability, our data indicate that this procedure never improves the correlation coefficient more than .12. The reliability was similar for all three devices. The results of this study indicate that for the knee and elbow joints, goniometric measurements performed in a clinical setting can be highly reliable. The method described in this study provides a simple protocol that can be used clinically to investigate goniometric reliability.

Elbow Joint↗

A critical assessment of factors influencing reliability in the classification of fractures, using fractures of the tibial plafond as a model.

OBJECTIVE: To investigate three factors that may influence the reliability of a fracture classification system: (a) the quality of the radiographs; (b) the ability of observers to identify the fracture fragments; and (c) the use of binary decision making. DESIGN: Assessment of interobserver reliability of blinded observers. SETTING: Medical school department of orthopaedics. PARTICIPANTS: Two attending orthopaedists, two PGY-5 orthopaedic residents, and two PGY-3 orthopaedic residents served as observers. INTERVENTION: Observers classified radiographs of twenty-five tibial plafond fractures according to the Rüedi-Allgöwer and binary classification systems, and also rated the quality of each radiograph as adequate or inadequate for accurately classifying the fracture. At a second session, observers classified the same radiographs after marking the fragments of the tibial articular surface, as well as radiographs that had the articular fragments premarked by the senior author. MAIN OUTCOME MEASURES: Pairwise interobserver reliability was analyzed by kappa statistics, and mean kappa values were compared for each method of fracture classification. RESULTS: No difference in interobserver reliability was detected between the Rüedi-Allgöwer and binary classification systems. Interobserver agreement on the adequacy of the radiographs was poorer than agreement on the classification of the fractures themselves. Having observers mark the fragments of the tibial articular surface had no effect on interobserver reliability; having the articular fragments premarked, however, significantly improved interobserver reliability in classifying the fractures. CONCLUSIONS: The results of this study underscore the complexity of tibial plafond fractures and the difficulty observers have in reliably interpreting fracture radiographs. Fracture classification systems, such as the Rüedi-Allgöwer, predicated on identification of the number and displacement of articular fragments, may inherently perform poorly on reliability analyses because of observer difficulty in reliably identifying the fragments. Because binary decision making did not improve the reliability of fracture classification in this study, further investigation of the effectiveness of binary decision making may be advisable before such strategies are put into widespread use.

Adult↗

Comparison of reliability between the Lenke and King classification systems for adolescent idiopathic scoliosis using radiographs that were not premeasured.

STUDY DESIGN: Multisurgeon comparison of two radiographic scoliosis curve classification systems was performed. OBJECTIVE: To determine the reliability of the King and Lenke classifications systems for adolescent idiopathic scoliosis using radiographs that had not been premeasured. SUMMARY OF BACKGROUND DATA: Recent studies introducing the new Lenke classification system for idiopathic scoliosis have reported reliability improved over that of the King classification system. This newer classification system evaluates three different parameters (curve type, lumbar modifier, and sagittal thoracic modifier) and then combines them. The reliability of both classification systems had been determined using radiographs in which all of the curves had been premeasured (recorded on the radiographs) before review by examiners. However, in a normal clinical situation, spine surgeons need to determine the Cobb angles independently, thus introducing another variable. METHODS: On two separate occasions, four orthopedic surgeons independently evaluated preoperative radiographs (standing posteroanterior, lateral, and two supine side-bending views) of 50 patients with adolescent idiopathic scoliosis. All measurements had been removed on every radiograph before each evaluation. The results were determined by calculating the average percentage of intraobserver and interobserver agreement. Reliability was quantified using kappa statistics. RESULTS: The King classification demonstrated good intraobserver and fair interobserver reliability. Intraobserver percentage of agreement averaged 83.5% (kappa coefficient, 0.81). Interobserver percentage of agreement averaged 68.0% (kappa coefficient, 0.61). All three parameters of the overall Lenke curve classification demonstrated fair reliability. Intraobserver percentage of agreement averaged 65.0% (kappa coefficient, 0.60). Interobserver percentage of agreement averaged 55.5% (kappa coefficient, 0.50). When the Lenke curve type was examined separately, intraobserver percentage of agreement averaged 81.5% (kappa coefficient, 0.76) and interobserver percentage of agreement averaged 71.5% (kappa coefficient, 0.64). The results for this variable (curve type) were similar to those for the King classification. For the Lenke lumbar modifier, the percentage of agreement and reliability were excellent. For the sagittal thoracic modifier, the percentage of agreement was good, but the kappa values were low because of an extreme imbalance in the grouping of hypokyphotic, normal, and hyperkyphotic spines. CONCLUSIONS: In this study, with each investigator performing the radiographic measurements, the King classification was found to be better than had been reported recently. The Lenke classification system for adolescent idiopathic scoliosis was found to be less reliable than previously reported when the radiographs were premeasured. This was particularly true when all three parameters of this new classification system were combined. This difference in reliability of the Lenke classification between studies can be attributed to the additional variable of determining the Cobb measurements on each of the unmarked radiographs. Although this new classification system has limitations with respect to interobserver and intraobserver reliability, for planning operative treatment, it offers a more comprehensive radiographic evaluation of patients with adolescent idiopathic scoliosis than previous systems.

Adolescent↗