PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Reliability between two observers using a protocol for diagnosing essential tremor.

Protocols with demonstrated reliability have been established for the diagnosis of numerous movement disorders. whereas in the essential tremor (ET) literature, there is no discussion about the reliability of diagnostic protocols. Lack of knowledge of the reliability of diagnostic protocols in ET limits the use of these protocols because reliability is an essential requirement for scientific quality in data management. The objective of this study was to determine the reliability of a protocol for diagnosing ET. The protocol consists of a Tremor Interview, a videotaped Tremor Examination, and a diagnostic algorithm. Eighty-three subjects with ET, identified in a community-based health study in Washington Heights-Inwood, New York, were matched with 83 control subjects from the same community. These subjects and their relatives are being recruited to participate in the Washington Heights-Inwood Genetic Study of ET. Two hundred twenty-six subjects have been evaluated to date (35 ET cases, 40 controls, 151 relatives). All 226 underwent an 84-item Tremor Interview and 26-item videotaped Tremor Examination. Diagnoses (normal, possible ET, probable ET, definite ET) were independently assigned by two blinded neurologists specializing in movement disorders. The kappa statistic, k, was used to determine diagnostic agreement between these two neurologists. The concordance rate between two raters using diagnostic categories definite ET, probable ET. possible ET, and normal was 80%; kw = 0.84 (near perfect to perfect agreement). The concordance rate between two raters using two diagnostic categories (definite ET and normal) was 100%; k = 1.00 (perfect agreement). There was high correlation between the two raters' total tremor scores (r = 0.89, p < 0.00001). This diagnostic protocol is highly reliable. Research in ET would greatly benefit from diagnostic protocols with demonstrated reliability.

Adolescent↗

Intratester and intertester reliability and criterion validity of the parallelogram and universal goniometers for active knee flexion in healthy subjects.

BACKGROUND AND PURPOSE: A new parallelogram goniometer was designed by the Rehabilitation Centre of the Royal Ottawa Health Care Group in 1983. The advantage of using such a goniometer is that the clinician is not required to estimate the joint axis of rotation when taking a measurement. The parallelogram goniometer has obtained a good intratester and intertester reliability when measuring active range of motion of hip abduction on eight individuals with hip pathologies. However, the validity of the parallelogram goniometer has not been examined. The purposes of this study were to examine the intratester and intertester reliability and the criterion validity of the parallelogram and universal goniometers for active knee flexion on healthy individuals. SUBJECTS: Sixty healthy university students (44 females and 16 males; mean age of 20.6 yrs.) participated to this study. METHODS: Measurements with the universal and parallelogram goniometers were taken in two different positions, the smaller and larger angles of active knee flexion. All measurements were taken by two trained testers. A radiograph was taken in both positions to serve as the 'gold standard'. The sequence of the measurements and radiographs were randomly selected. The intra and intertester reliability of both goniometers were established by calculating the intraclass correlation coefficients (ICCs) using the repeated-measures ANOVA. The criterion validity was examined by calculating Pearson product-moment correlation coefficients (tau) between each goniometric and radiologic measurements. A 0.05 level of significance was chosen for each statistical test. RESULTS: Intratester reliability ranged from good to excellent for the small angles (ICC = 0.85 and 0.87) and the large angles (ICC = 0.91 and 0.96) when using the parallelogram goniometer. Intertester reliability was fair for the small angles of flexion (ICC = 0.43 to 0.52) and good to excellent for the large angles of flexion (ICC = 0.82 to 0.88). The parallelogram goniometer was found to have greater validity when measuring the large angles of knee flexion (r = 0.73 and 0.77) compared to the small angles of knee flexion (r = 0.33 and 0.41). Similar results of reliability and validity were obtained with the universal goniometer. CONCLUSION: The results of this study have clinical importance. The use of the parallelogram goniometer was found to be as reliable and valid as the universal goniometer when measuring active knee flexion. However, the parallelogram goniometer offered clinicians the advantages of obtaining precise angular measurements with fewer adjustments, and a faster application technique. Further studies on the parallelogram goniometer are necessary among individuals presenting with altered range of motion at different joints.

Adult↗

The reliability of virtual organ computer-aided analysis (VOCAL) for the semiquantification of ovarian, endometrial and subendometrial perfusion.

OBJECTIVES: Three-dimensional power Doppler angiography (3D-PDA) has been largely used for the subjective assessment of vascular patterns but semiquantification of the power Doppler signal is now possible. We examined the intraobserver and interobserver reliability of the semiquantification of ovarian, endometrial and subendometrial blood flow using 3D-PDA, virtual organ computer-aided analysis (VOCAL) and shell-imaging. METHODS: 3D-PDA was used to acquire 20 ovarian and 20 endometrial volumes from 40 different patients at various stages of in vitro fertilization treatment. VOCAL was then used to delineate the 3D areas of interest and the 'histogram facility' employed to generate three indices of vascularity: the vascular index, the flow index and the vascularization flow index. Intraobserver and interobserver reliability was assessed by two-way, mixed, intraclass correlation coefficients (ICCs) and general linear modeling was used to examine for differences in the mean values between each observer. RESULTS: The intraobserver reliability for both observers was extremely high and there were no differences in reliability between the observers for measurements of both volume and vascularity within the ovary or endometrium and its shells. With the exception of the outside subendometrial shell volumes, there were no significant differences between the two observers in the mean values obtained for either endometrial or ovarian volume and vascularity measurements. The interobserver reliability of measurements was equally high throughout with all measurements obtaining a mean ICC of above 0.985. CONCLUSIONS: 3D-PDA and shell-imaging offer a reliable, practical and non-invasive method for the assessment of ovarian, endometrial and subendometrial blood flow. Future work should concentrate upon confirming the reliability of data acquisition and the validity of the technique before its predictive value can be truly tested in prospective clinical studies.

Endometrium↗

The reliability theory of aging and longevity.

Reliability theory is a general theory about systems failure. It allows researchers to predict the age-related failure kinetics for a system of given architecture (reliability structure) and given reliability of its components. Reliability theory predicts that even those systems that are entirely composed of non-aging elements (with a constant failure rate) will nevertheless deteriorate (fail more often) with age, if these systems are redundant in irreplaceable elements. Aging, therefore, is a direct consequence of systems redundancy. Reliability theory also predicts the late-life mortality deceleration with subsequent leveling-off, as well as the late-life mortality plateaus, as an inevitable consequence of redundancy exhaustion at extreme old ages. The theory explains why mortality rates increase exponentially with age (the Gompertz law) in many species, by taking into account the initial flaws (defects) in newly formed systems. It also explains why organisms "prefer" to die according to the Gompertz law, while technical devices usually fail according to the Weibull (power) law. Theoretical conditions are specified when organisms die according to the Weibull law: organisms should be relatively free of initial flaws and defects. The theory makes it possible to find a general failure law applicable to all adult and extreme old ages, where the Gompertz and the Weibull laws are just special cases of this more general failure law. The theory explains why relative differences in mortality rates of compared populations (within a given species) vanish with age, and mortality convergence is observed due to the exhaustion of initial differences in redundancy levels. Overall, reliability theory has an amazing predictive and explanatory power with a few, very general and realistic assumptions. Therefore, reliability theory seems to be a promising approach for developing a comprehensive theory of aging and longevity integrating mathematical methods with specific biological knowledge.

Adult↗

Reliability of self-reported breast screening information in a survey of lower income women.

BACKGROUND: Self-reported behavior is widely used to estimate the prevalence of breast cancer screening and to evaluate programs for promoting screening, but detailed studies of reliability have not previously been performed. METHODS: Reliability was assessed by comparing responses to questions about screening behavior from repeat personal interviews of 382 women age 40 and older living in low-income census tracts of two Florida communities. Reliability was assessed using Pearson's correlation (r) and kappa (kappa) coefficients. RESULTS: Estimated reliabilities were kappa = 0.38 for "ever had clinical breast examination," kappa = 0.82 for "ever had mammogram," kappa = 0.65 for "mammogram in past year," r = 0.54 for "date of last mammogram," and r = 0.72 for "number of mammograms." The dates of last mammogram reported at the two interviews agreed within 1 month for 64% of the women, while the dates of last clinical breast examination agreed within 1 month for 50% of the women. Reliability of "ever had mammogram" was significantly related to demographic variables. CONCLUSIONS: Women reliably report ever having mammography, but information about timing and frequency has lower reliability. The results have implications for breast screening research because measurement error affects the precision of estimates and the sample sizes needed to detect program effects.

Adult↗

Reliability of psychophysiological responding as a function of trait anxiety.

This study examined the temporal stability of three psychophysiological responses (frontal electromyographic activity, hand surface temperature, and heart rate) recorded over four sessions (days 1, 2, 8, and 28) on 34 subjects, 17 with high Spielberger Trait Anxiety Inventory scores and 17 with low scores. Each session consisted of a 20-minute adaptation period, a baseline condition, and two stressors (one cognitive, the other physical). Two forms of reliability coefficients were employed, intraclass correlations and Pearson Product Moment; the two types of reliability coefficients arrived at the same conclusions. Results indicated that reliability coefficients for the two anxiety groups did not differ on frontal EMG or heart rate responses; however, hand surface temperature responding was considerably less reliable for high anxious individuals than low anxious individuals. Reliability coefficients on absolute scores were, for the most part, reliable. Treating the responses as relative measures (percent change from baseline or simple change scores from baseline) produced smaller and less reliable coefficients. Magnitudes of the three physiological responses did not significantly differ as a function of high or low trait anxiety. Findings are discussed in terms of their clinical, as well as basic psychophysiological, importance.

Adult↗

Reliability and validity of clinical outcome measurements of osteoarthritis of the hip and knee--a review of the literature.

High reliability and validity of clinical rating schemes is crucial for their use as outcome measurements of treatment of hip and knee osteoarthritis. In this paper, we review the empirical evidence on the reliability and validity of commonly used clinical scores. Clinical scores and related reliability and validity studies were identified by systematic literature search. Scores were classified according to the type and joint. Reliability and validity studies were characterized according to design, population, number and qualification of observers, number of measurements, time interval between repeat measurements and results. Reliability and validity studies were reported for only 6 and 15 of the 45 identified clinical scores, respectively. Although comparisons are difficult due to differences in study design, relatively high reliability was reported for most measurements of pain, stiffness, and physical function, while results are less conclusive for clinical signs. Most validity studies focused on the correlation between various scores. Correlation was generally found to be high for overall numerical ratings, but scores often differed with respect to the interpretation of these ratings. Validity has been more comprehensively studied for Lequesne's scores, WOMAC, and ILAS, and these scores have shown satisfactory responsiveness to different treatment effects. Overall, knowledge on reliability and validity of clinical scores of hip and knee osteoarthritis is limited, underlining the need for further properly designed and conducted studies.

Hip↗

Reliability of the Health Utilities Index--Mark III used in the 1991 cycle 6 Canadian General Social Survey Health Questionnaire.

This study presents information on the test-retest reliability of the Health Utility Index--Mark III (HUI) system used in cycle 6 of the Canadian General Social Survey (GSS). The HUI system used in this reliability study consists of an eight-attribute health status classification system (HSCS) and a function for generating a summary score of health-related quality of life. To estimate test-retest reliability, a stratified random sample of individuals (n = 506) completing GSS telephone interviews during August and September, 1991 were interviewed again 1 month later. Weighting adjustments based on the probability of selection were invoked during the analyses to provide unbiased estimates of test-retest reliability for all GSS respondents in the August-September period. The results indicate that the individual questions, attributes and provisional index scores generally provided reliable information on health status in the GSS. The exceptions to this were limitations in speech and dexterity which were reported very infrequently. Kappa estimates of test-retest reliability for individual questions varied from 0.184 to 0.766. For the eight attributes, kappa estimates varied from 0.137 to 0.728. Using the provisional index scores to quantify health overall, a test-retest reliability of 0.767 was obtained (intra-class correlation coefficient).

Activities of Daily Living↗

Factors affecting the reliability of ratings of students' clinical skills in a medicine clerkship.

OBJECTIVE: To determine the overall reliability and factors that might affect the reliability of ratings of students' clinical skills in a medicine clerkship. DESIGN: A nine-item instrument was used to evaluate students' clinical skills. Raters were also asked to provide a grade of each student's overall clinical performance. Generalizability studies were performed to estimate the reliability of the ratings. The effects of rater experience and clerkship setting were investigated by regression analysis. SETTING: Teaching hospitals and community-based sites in three Northwestern states. PARTICIPANTS: All students (328) who had completed the 12-week clerkship in internal medicine at one medical school during the academic years 1987-1989. Raters included attending physicians, chief residents, and other residents. RESULTS: Seven observations were needed to provide a reliable rating of the overall clinical grade. More observations were needed to obtain reliable ratings for individual items, ranging from seven observations needed for the rating of data gathering skills to 27 observations needed for the rating of interpersonal relationships with patients. Rater experience and clerkship setting (i.e., teaching hospitals vs. community-based clinics) were found, in general, not to affect significantly the ratings received by students. CONCLUSIONS: Reliable ratings of students' overall clinical skills, including overall clinical grades, can be achieved by collecting a minimum of seven observations. More observations are needed to measure reliably the interpersonal aspects of clinical performance. These findings support the use of performance ratings to evaluate clinical skills and knowledge of students in clerkship settings.

Clinical Clerkship↗

Reliability of the timeline follow-back sexual behavior interview.

The reliability of self-reported sexual behavior is a question of utmost importance to human immunodeficiency virus (HIV) prevention research. The Timeline Follow-Back (TLFB) interview, which was developed to assess alcohol consumption on the event level, incorporates recall-enhancing techniques that result in reliable information. In this study, the TLFB interview was adapted to assess HIV-related sexual behaviors and their antecedents, and its reliability was assessed. The interview was administered to 110 participants (46% women, M age = 19.7; range = 18-41), and 58 participants who reported sexual behavior during the previous three months returned one week later for a second interview. Test-retest intraclass correlations (rho) from the TLFB protocol showed that all sexual behaviors were reported reliably (rho range = .86 to .97, median = .96). Bootstrapping, a nonparametric statistical technique, was used for significance testing in the reliability analyses. Reliability was equivalent across each of the three months assessed with the TLFB and was equivalent to conventional assessment methods (i.e. single-item questions). These findings show that the TLFB sexual behavior interview provides reliable reports of sexual behavior over three months and yields event-level data that are extremely valuable for sexual behavior and HIV-prevention research.

Adolescent↗

Reliability and validity of the Assessment of Daily Activity Performance (ADAP) in community-dwelling older women.

BACKGROUND AND AIMS: The Assessment of Daily Activity Performance (ADAP) test was developed, and modeled after the Continuous-scale Physical Functional Performance (CS-PFP) test, to provide a quantitative assessment of older adults' physical functional performance. The aim of this study was to determine the intra-examiner reliability and construct validity of the ADAP in a community-living older population, and to identify the importance of tester experience. METHODS: Forty-three community-dwelling, older women (mean age 75 yr +/-4.3) were randomized to the test-retest reliability study (n=19) or validation study (n=24). The intra-examiner reliability of an experienced (tester 1) and an inexperienced tester (tester 2) was assessed by comparing test and retest scores of 19 participants. Construct validity was assessed by comparing the ADAP scores of 24 participants with self-perceived function by the SF-36 Health Survey, muscle function tests, and the Timed Up and Go test (TUG). RESULTS: Tester 1 had good consistency and reliability scores (mean difference between test and retest scores (DIF), -1.05+/-1.99; 95% confidence interval (CI), -2.58 to 0.48; Cronbach's alpha (alpha) range, 0.83 to 0.98; intraclass correlation (ICC) range, 0.75 to 0.96; Limits of Agreement (LoA), -2.58 to 4.95). Tester 2 had lower reliability scores (DIF, -2.45+/-4.36; 95% CI, -5.56 to 0.67; alpha range, 0.53 to 0.94; ICC range, 0.36 to 0.90; LoA, -6.09 to 10.99), with a systematic difference between test and retest scores for the ADAP domain lower-body strength (-3.81; 95% CI, -6.09 to -1.54), ADAP correlated with SF-36 Physical Functioning scale (r=0.67), TUG test (r=-0.91) and with isometric knee extensor strength (r=0.80). CONCLUSIONS: The ADAP test is a reliable and valid instrument. Our results suggest that testers should practise using the test, to improve reliability, before applying it to clinical settings.

Activities of Daily Living↗

Reliability of a criterion-based test of athletes with knee injuries; where the physiotherapist and the patient independently and simultaneously assess the patient's performance.

UNLABELLED: A new criterion-based evaluation test method, has been developed in order to assess the functional ability of athletes with knee injuries, 'Tests for Athletes with Knee-injuries' (TAK). The physiotherapist and the patient assess independently and simultaneously the patient's performance. The TAK comprises eight demanding functional activities with emphasis on strength, stability, springiness and endurance. OBJECTIVES: To evaluate the inter-rater and intra-rater reliability of TAK between the physiotherapist's and the patient's assessments. Further, to evaluate the relation between the functional tests in TAK and the isokinetic quadriceps muscle strength. MATERIALS AND METHODS: Fifty-nine subjects were included in the study. Thirty-one were anterior cruciate ligament (ACL) reconstructed, fourteen were ACL-injured not reconstructed and fourteen were healthy athletes. The inter-rater-reliability was evaluated by assessments of 59 subjects carried out by two independent physiotherapists using visual observation. The assessment was rated on a 0-10-point scale according to five elaborate criteria drawn up for each test. Simultaneously, the subjects were asked to rate their own performance on each test using a 0-10-point scale. The intra-rater-reliability of TAK was evaluated by a test-retest of 31 patients. The relation between the physiotherapist's and the patients' ratings as well as of the patients' ratings at two different occasions were evaluated. Isokinetic quadriceps muscle strength was measured in a Biodex dynamometer on all 59 subjects in order to study the relation between quadriceps muscle strength and the results of the functional tests in TAK. RESULTS: Inter-rater-reliability showed good consistency between the assessments of the two physiotherapists in seven of eight tests (kappa = 0.62-0.78). The intra-rater-reliability was moderate to good (kappa = 0.43-0.65) in the test-retest study. The consistency of the physiotherapist and the patients' assessments differed, but showed good correlation. The consistency of the test-retest study of the patients' assessment was low. The correlation between the isokinetic quadriceps muscle strength measured in a Biodex dynamometer and the results of the functional tests was moderate in this study. CONCLUSIONS: This criterion-based test method for athletes with knee injuries showed good inter-rater reliability and acceptable intra-rater reliability for the physiotherapists' assessment. The consistency of the patients' ratings was low. The correlation between isokinetic quadriceps muscle strength and functional tests in TAK was moderate. The validity has not been evaluated in this study but will be done in the future.

Adolescent↗

Improving reliability in the classification of fractures of the acetabulum.

BACKGROUND: Plain radiographs of the pelvis are routinely used in the initial assessment of patients with suspected fractures of the acetabulum. It is necessary for orthopaedic resident trainees, emergency physicians as well as orthopaedic surgeons who infrequently treat trauma patients to be able to describe these fracture patterns reliably to traumatologist orthopaedic surgeons who ultimately take over the patient care. Our purpose was two-fold: (1) to determine the reliability of the component parts of the Letournel classification of acetabular fractures involving six anteroposterior (AP) radiographic lines, and (2) to examine whether the addition of oblique radiograph views (Judet views) would improve the reliability. METHODS: Thirty sets of AP and oblique radiographs (Judet views) of the pelvis were selected from a hospital database to represent various types of acetabular fractures. Six reviewers (three orthopaedic trainees and three community orthopaedic surgeons) independently reviewed the radiographs. For each radiograph, the reviewer classified the acetabular fracture according to the Letournel classification. In addition, each reviewer utilized a simplified classification scheme using six radiographic lines on the AP pelvic radiograph. Interobserver reliabilities among reviewers were reported along with the intraclass correlation coefficient (ICC) and kappa values. RESULTS: Agreement for the Letournel classification increased with increasing physician experience (trainees ICC=-0.14 and community surgeons ICC=0.56). Interobserver reliability between trainees and community surgeons improved when the six radiographic lines were used (range kappa=0.09-0.89). The oblique pelvic radiographs (Judet views) did not significantly improve reliability among physicians. CONCLUSIONS: In this study we report the following: (1) the reliability of the Letournel classification improves with level of training, (2) physicians with less experience with acetabular fractures have significantly better agreement in identifying fractures using the six radiographic lines on the AP film than the Letournel classification, and (3) agreement among the reviewers for the AP pelvic radiograph is not improved with additional oblique (Judet) views.

Acetabulum↗

Reliability and validity of Functional Capacity Evaluation methods: a systematic review with reference to Blankenship system, Ergos work simulator, Ergo-Kit and Isernhagen work system.

OBJECTIVES: Functional Capacity Evaluation methods (FCE) claim to measure the functional physical ability of a person to perform work-related tasks. The purpose of the present study was to systematically review the literature on the reliability and validity of four FCEs: the Blankenship system (BS), the ERGOS work simulator (EWS), the Ergo-Kit (EK) and the Isernhagen work system (IWS). METHODS: A systematic literature search was conducted in five databases (CINAHL, Medline, Embase, OSH-ROM and Picarta) using the following keywords and their synonyms: functional capacity evaluation, reliability and validity. The search strategy was performed for relevance in titles and abstracts, and the databases were limited to literature published between 1980 and April 2004. Two independent reviewers applied the inclusion criteria to select all relevant articles and evaluated the methodological quality of all included articles. RESULTS: The search resulted in 77 potential relevant references but only 12 papers were identified for inclusion and assessed for their methodological quality. The interrater reliability and predictive validity of the IWS were evaluated as good while the procedure used in the intrarater reliability (test-retest) studies was not rigorous enough to allow any conclusion. The concurrent validity of the EWS and EK was not demonstrated while no study was found on their reliability. No study was found on the reliability and validity of the BS. CONCLUSIONS: More rigorous studies are needed to demonstrate the reliability and the validity of FCE methods, especially the BS, EWS and EK.

Humans↗

Inter- and intrajudge reliability of a clinical examination of swallowing in adults.

This study investigates inter- and intrajudge reliability of a clinical examination of swallowing in adults. Several investigations have sought correlations between clinical indicators of dysphagia and the actual presence of dysphagia as determined by videofluoroscopy. Whereas some investigations have reported interjudge reliability for the videofluoroscopic measures employed, none have reported reliability for clinical measures. Without established reliability for rating clinical measures, conclusions drawn regarding the utility of a measure for detecting aspiration can be called into question. Results of the present study indicate that fewer than 50% of the measures clinicians typically employ are rated with sufficient inter- and intrajudge reliability. Measures of vocal quality and oral motor function were rated more reliably than were history measures or measures taken during trial swallows. There is a need to define more clearly the measures employed in clinical examinations and to be consistent in reporting reliability for clinical measures of swallowing function in future research.

Adult↗

Test-retest reliability of self-reported sexual behavior, sexual orientation, and psychosexual milestones among gay, lesbian, and bisexual youths.

Despite the importance of reliable self-reported sexual information for research on sexuality and sexual health, research has not examined reliability of information provided by gay, lesbian, and bisexual (GLB) youths. Test-retest reliability of self-reported sexual behaviors, sexual orientation, sexual identity, and psychosexual developmental milestones was examined among an ethnically diverse sample of 64 self-identified GLB youths. Two face-to-face interviews were conducted approximately 2 weeks apart using the Sexual Risk Behavior Assessment Schedule for Homosexual Youths (SERBAS-Y-HM). Overall, the mean of the test-retest reliability coefficients was substantial for 6 of the 7 domains: lifetime sexual behaviors (M=.89), sexual behavior in the past 3 months (M=.96), unprotected sexual behavior in the past 3 months (M=.93), sexual identity (kappa=.89), sexual orientation (M=.82), and ages of various psychosexual developmental milestones (M=.77). Inconsistent reliability was found for reports of sexual behaviors while using substances. A small number of gender differences emerged, with lower reliability among female youths in the lifetime number of same-sex partners. The overall findings suggest that a wide range of self-reported sexual information can be reliably assessed among GLB youths by means of interviewer-administered questionnaires, such as the SERBAS-Y-HM.

Adolescent↗

Cosmetic outcomes following breast conservation therapy: in search of a reliable scale.

INTRODUCTION: Multiple scales to evaluate breast cosmesis following breast conserving treatment (BCT) have been developed, however reliability is a problem. Panel scores, where scores from two or more individuals are combined, were assessed to examine their effect on reliability for two different cosmetic scales. METHODS: Women, two or more years following BCT, were recruited from a single breast centre. Photographs of each participant were evaluated independently by six health care professionals on two separate occasions. A simple four-point scale and more involved multi-item scale were used to assess cosmetic outcome. Reliability was assessed with the weighted kappa statistic for increasing panel sizes. RESULTS: Ninety-nine women were evaluated. Intra rater reliability increased from 0.73 to 0.83 for the four-point scale, for increasing panel sizes, however 95% confidence intervals generally overlapped. A smaller and more unpredictable effect was seen on the multi-item subscale, range 0.69 to 0.73. Inter rater reliability increased from 0.68 to 0.93 for the four-point scale, and 0.75 to 0.96 for the multi-item scale, for increasing panel sizes; 95% confidence intervals did not overlap. A panel of three for either scale provided almost perfect kappa values with only small improvements with larger panel sizes. CONCLUSIONS: Care should be used in interpreting results where cosmetic outcomes have been obtained from a single evaluator. Panel scores can be used to significantly improve inter-rater, but not intra rater reliability, for the scales studied. Comparable reliability, in combination with simplicity of use and interpretation, would favour the four-point scale for breast cosmetic evaluation over the multi-item scale.

Breast Neoplasms↗

Reliability of goniometric measurements and visual estimates of ankle joint active range of motion obtained in a clinical setting.

We examined intratester and intertester reliability for goniometric measurements of ankle dorsiflexion (ADF) and ankle plantar flexion (APF) active range of motion (AROM). Parallel-forms intratester reliability for ankle AROM measurements obtained by the universal goniometer (UG) and by visual estimation (VE) and intertester reliability for VE of ADF and APF were examined. Repeated measurements were obtained on 38 patients with orthopedic problems by 10 physical therapists in a clinical setting. For intratester reliability of measurements obtained with UG, intraclass correlation coefficients (ICC) for all physical therapists were 0.64 to 0.92 (median, 0.825) for ADF and 0.47 to 0.96 (median, 0.865) for APF. Intertester reliability was quantified with use of ICC. ICCs for measurements obtained by UG were 0.28 for ADF and 0.25 for APF; ICC of VE for ADF was 0.34 and was 0.48 for APF. ICC for parallel-forms intratester reliability obtained with UG and VE ranged from 0 to 0.94 (median, 0.58) for ADF and 0 to 0.86 (median, 0.625) for APF. Thus, a physical therapist should use a goniometer when making repeated measurements of ankle joint AROM. Considerable inconsistency exists when two or more physical therapists make repeated goniometric and visual measurements of ankle motion on the same subject. Physical therapists may erroneously conclude that a patient's AROM has changed because of treatment when the change could be attributed to a lack of intertester reliability.

Adolescent↗