PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

The rating reliability calculator.

BACKGROUND: Rating scales form an important means of gathering evaluation data. Since important decisions are often based on these evaluations, determining the reliability of rating data can be critical. Most commonly used methods of estimating reliability require a complete set of ratings i.e. every subject being rated must be rated by each judge. Over fifty years ago Ebel described an algorithm for estimating the reliability of ratings based on incomplete data. While his article has been widely cited over the years, software based on the algorithm is not readily available. This paper describes an easy-to-use Web-based utility for estimating the reliability of ratings based on incomplete data using Ebel's algorithm. METHODS: The program is available public use on our server and the source code is freely available under GNU General Public License. The utility is written in PHP, a common open source imbedded scripting language. The rating data can be entered in a convenient format on the user's personal computer that the program will upload to the server for calculating the reliability and other statistics describing the ratings. RESULTS: When the program is run it displays the reliability, number of subject rated, harmonic mean number of judges rating each subject, the mean and standard deviation of the averaged ratings per subject. The program also displays the mean, standard deviation and number of ratings for each subject rated. Additionally the program will estimate the reliability of an average of a number of ratings for each subject via the Spearman-Brown prophecy formula. CONCLUSION: This simple web-based program provides a convenient means of estimating the reliability of rating data without the need to conduct special studies in order to provide complete rating data. I would welcome other researchers revising and enhancing the program.

Humans↗

Reliability and cultural applicability of the Greek version of the International Personality Disorders Examination.

BACKGROUND: The International Personality Disorders Examination (IPDE) constitutes the proposal of the WHO for the reliable diagnosis of personality disorders (PD). The IPDE assesses pathological personality and is compatible both with DSM-IV and ICD-10 diagnosis. However it is important to test the reliability and cultural applicability of different IPDE translations. METHODS: Thirty-one patients (12 male and 19 female) aged 35.25 +/- 11.08 years, took part in the study. Three examiners applied the interview (23 interviews of two and 8 interviews of 3 examiners, that is 47 pairs of interviews and 70 single interviews). The phi coefficient was used to test categorical diagnosis agreement and the Pearson Product Moment correlation coefficient to test agreement concerning the number of criteria met. RESULTS: Translation and back-translation did not reveal specific problems. Results suggested that reliability of the Greek translation is good. However, socio-cultural factors (family coherence, work environment etc) could affect the application of some of the IPDE items in Greece. The diagnosis of any PD was highly reliable with phi >0.92. However, diagnosis of non-specific PD was not reliable at all (phi close to 0) suggesting that this is a true residual category. Diagnosis of specific PDs were highly reliable with the exception of schizoid PD. Diagnosis of antisocial and Borderline PDs were perfectly reliable with phi equal to 1.00. CONCLUSIONS: The Greek translation of the IPDE is a reliable instrument for the assessment of personality disorder but cultural variation may limit its applicability in international comparisons.

Adult↗

Health measurement using the ICF: test-retest reliability study of ICF codes and qualifiers in geriatric care.

BACKGROUND: The International Classification of Functioning, Disability and Health (ICF) was published by the World Health Organization (WHO) to standardize descriptions of health and disability. Little is known about the reliability and clinical relevance of measurements using the ICF and its qualifiers. This study examines the test-retest reliability of ICF codes, and the rate of immeasurability in long-term care settings of the elderly to evaluate the clinical applicability of the ICF and its qualifiers, and the ICF checklist. METHODS: Reliability of 85 body function (BF) items and 152 activity and participation (AP) items of the ICF was studied using a test-retest procedure with a sample of 742 elderly persons from 59 institutional and at home care service centers. Test-retest reliability was estimated using the weighted kappa statistic. The clinical relevance of the ICF was estimated by calculating immeasurability rate. The effect of the measurement settings and evaluators' experience was analyzed by stratification of these variables. The properties of each item were evaluated using both the kappa statistic and immeasurability rate to assess the clinical applicability of WHO's ICF checklist in the elderly care setting. RESULTS: The median of the weighted kappa statistics of 85 BF and 152 AP items were 0.46 and 0.55 respectively. The reproducibility statistics improved when the measurements were performed by experienced evaluators. Some chapters such as genitourinary and reproductive functions in the BF domain and major life area in the AP domain contained more items with lower test-retest reliability measures and rated as immeasurable than in the other chapters. Some items in the ICF checklist were rated as unreliable and immeasurable. CONCLUSION: The reliability of the ICF codes when measured with the current ICF qualifiers is relatively low. The result in increase in reliability according to evaluators' experience suggests proper education will have positive effects to raise the reliability. The ICF checklist contains some items that are difficult to be applied in the geriatric care settings. The improvements should be achieved by selecting the most relevant items for each measurement and by developing appropriate qualifiers for each code according to the interest of the users.

Activities of Daily Living↗

Reliability of a measure of post-stroke shoulder pain in patients with and without aphasia and/or unilateral spatial neglect.

OBJECTIVE: To determine the inter/intra-rater reliability of expert physiotherapists (PTs) measuring post-stroke shoulder pain with 100 mm vertical visual analogue scales (VAS; intensity, frequency and affective response) and a categorical site-of-pain scale. DESIGN: Three PTs independently rated subjects (normal clinical procedure but with a standardized starting position) on three days, at the same time of day, during one week in a randomized order determined by a nested latin square. Reliability for VAS scores was determined with the intraclass correlation coefficient (ICC) and for site-of-pain with the kappa statistic (kappa). Acceptable reliability was set at 0.75. The limits of agreement were also calculated. SETTING: Community. SUBJECTS: Thirty-three patients, mean time post stroke 42 months (range 7-360). RESULTS: Mean inter-rater reliability was 0.79 for intensity, 0.75 for frequency and 0.62 for affective response (ICC). The limits of agreement were wide and rater bias was significant for 6/27 ratings. Mean intra-rater reliability was 0.70 for intensity, 0.77 for frequency and 0.69 for affective response (ICC). For site-of-pain inter-rater reliability ranged from 0.156 (kappa) to 0.385 (kappa) and intrarater reliability ranged from 0.300 (kappa) to 0.559 (kappa). CONCLUSIONS: Although inter-rater reliability was acceptable for intensity and frequency there was a consistently large systematic bias between pairs of raters. Agreement might be improved if a standardized assessment procedure was used and/or if training in pain behaviour interpretation was provided.

Aged↗

Pressure mapping systems: reliability of pressure map interpretation.

BACKGROUND: Pressure mapping systems offer a new technology to assist with pressure care assessment. Data output from such systems can be presented in three forms: numerical data, a three-dimensional grid and a colour-coded pressure map. OBJECTIVES: To (1) investigate whether sole use of the pressure map was a reliable method of interpreting interface pressures when compared with use of the numerical data; (2) establish the inter- and intra-rater reliability of using pressure maps to assess pressure and determine whether reliability depended upon system operator experience; and (3) examine whether reliability extended to the range of seating surfaces being tested. DESIGN: A reliability study assessing the ranking of pressure maps recorded by the Force Sensing Array pressure mapping system. SETTING: A university occupational therapy department and a community NHS trust. SUBJECTS: Fifteen occupational therapists with experience in pressure mapping and 50 occupational therapy students with no practical experience of pressure mapping. INTERVENTIONS: Two sets of pressure maps were pre-recorded with an able-bodied adult seated on a variety of surfaces, with maps on each individual surface recorded over a 20-minute period at 2-minute intervals. Subjects ranked both sets of maps in terms of 'best to poorest' distribution of pressure. MAIN OUTCOME MEASURES: Rank orders of (1) pressure maps; (2) average interface pressures (mmHg); (3) maximum interface pressures (mmHg). RESULTS: The use of pressure maps to interpret interface pressures was a reliable method. Significant agreement existed within (p < 0.001) and between groups of operators and reliability extended over the range of seating surfaces tested. CONCLUSIONS: The practice of using pressure maps to interpret interface pressures in seating as opposed to using the associated numerical data can be supported. This was shown to be a reliable method of assessment by both experienced and less experienced operators across a range of seating surfaces.

Adult↗

Reliability and validity of functional balance tests post stroke.

OBJECTIVE: To contribute to the reliability and validity of a series of functional balance tests for use post stroke. DESIGN: Within-session, test-retest and intertester reliability was tested using the kappa coefficient and intraclass correlations. The tests were performed three times and the first and third attempts compared to test the within-session reliability. The tests were repeated a few days later to assess test-retest reliability and were scored simultaneously by two physiotherapists to assess the intertester reliability. To test criterion-related validity the tests were compared with the sitting section of the Motor Assessment Scale, Berg Balance Scale and Rivermead Mobility Index using Spearman's rho. SETTING: Stroke physiotherapy services of six National Health Service hospitals. PARTICIPANTS: People with a post stroke hemiplegia attending physiotherapy who had no other pathology affecting their balance took part. Thirty-five people participated in the reliability testing and 48 people took part in the validity testing. MAIN OUTCOME MEASURES: The following functional balance tests were used: supported sitting balance, sitting arm raise, sitting forward reach, supported standing balance, standing arm raise, standing forward reach, static tandem standing, weight shift, timed 5-m walk with and without an aid, tap and step-up tests. RESULTS: The ordinal level tests (supported sitting and standing balance and static tandem standing tests) showed 100% agreement in all aspects of reliability. Intraclass correlations for the other tests ranged from 0.93 to 0.99. All the tests showed significant correlations with the appropriate comparator tests (r = 0.32-0.74 p < 0.05), except the weight shift test and step-up tests which did not form significant relationship with Berg Balance Scale (r = 0.26 and 0.19 respectively). CONCLUSION: These functional balance tests are reliable and valid measures of balance disability post stroke.

Aged↗

The influence of contractures and variation in measurement stretching velocity on the reliability of the Modified Ashworth Scale in patients with severe brain injury.

OBJECTIVE: To determine the influence of contractures and different stretching velocities on the reliability of the Modified Ashworth Scale (MAS) in patients with severe brain injury and impaired consciousness. DESIGN: Cross-section observational study. SETTING: A rehabilitation centre for adult persons with neurological disorders. SUBJECTS: Fifty patients with impaired consciousness due to severe cerebral damage of various aetiologies. MEASUREMENT PROTOCOL: Three experienced and trained medical professionals rated each patient in a randomized order once daily for two consecutive days. Shoulder, elbow, wrist, knee and ankle spasticity were assessed by the use of the MAS with different stretching velocities. The presence of contractures was assessed by a goniometer. MAIN OUTCOME MEASURES: Retest and inter-rater reliability (k(w) = weighted kappa) of the MAS. RESULTS: The retest reliability of the MAS was good (shoulder joints (k(w) 0.74), elbow joints (k(w) 0.74), wrist joints (k(w) 0.72), knee joints (k(w) 0.72), ankle joints (k(w) 0.77)) and the inter-rater reliability was moderate (shoulder joints (k(w) 0.49), elbow joints (k(w) 0.52), wrist joints (k(w) 0.51), knee joints (k(w) 0.54) ankle joints (k(w) 0.49)). The presence of contractures significantly influenced the reliability of MAS in shoulder and wrist joints. No influence of stretching velocity on the reliability of the MAS was found. CONCLUSION: In patients with impaired consciousness due to severe brain injury the MAS has good retest, but only limited inter-rater, reliability. The presence of contractures may influence reliability of the MAS, but stretching velocity does not.

Brain Injuries↗

The Erasmus MC modifications to the (revised) Nottingham Sensory Assessment: a reliable somatosensory assessment measure for patients with intracranial disorders.

OBJECTIVE: To investigate the intra-rater and inter-rater reliability of the Erasmus MC modifications to the Nottingham Sensory Assessment (EmNSA). SUBJECTS: A consecutive sample of 18 inpatients, with a mean age of 57.7 years, diagnosed with an intracranial disorder and referred for physiotherapy. SETTING: The inpatient neurology and neurosurgery wards of a university hospital. DESIGN: Through discussions between four experienced neurophysiotherapists, the testing procedures of the revised Nottingham Sensory Assessment were further standardized. Subsequently, the intra-rater and inter-rater reliabilities of the EmNSA were investigated. RESULTS: The intra-rater reliability of the tactile sensations, sharp blunt discrimination and the proprioception items of the EmNSA were generally good to excellent for both raters with a range of weighted kappa coefficients between 0.58 and 1.00. Likewise the inter-rater reliabilities of these items were predominantly good to excellent with a range of weighted kappa coefficients between 0.46 and 1.00. An exception was the two-point discrimination that had a poor to good reliability, with the range for intra-rater reliability of 0.11-0.63 and for inter-rater reliability -0.10-0.66. CONCLUSION: The EmNSA is a reliable screening tool to evaluate primary somatosensory impairments in neurological and neurosurgical inpatients with intracranial disorders. Further research is necessary to consolidate these results and establish the validity and responsiveness of the Erasmus MC modifications to the NSA.

Adult↗

Reliability of assessment tools in rehabilitation: an illustration of appropriate statistical analyses.

OBJECTIVE: To provide a practical guide to appropriate statistical analysis of a reliability study using real-time ultrasound for measuring muscle size as an example. DESIGN: Inter-rater and intra-rater (between-scans and between-days) reliability. SUBJECTS: Ten normal subjects (five male) aged 22-58 years. METHOD: The cross-sectional area (CSA) of the anterior tibial muscle group was measured using real-time ultrasonography. MAIN OUTCOME MEASURES: Intraclass correlation coefficients (ICCs) and the 95% confidence interval (CI) for the ICCs, and Bland and Altman method for assessing agreement, which includes calculation of the mean difference between measures (d), the 95% CI for d, the standard deviation of the differences (SDdiff), the 95% limits of agreement and a reliability coefficient. RESULTS: Inter-rater reliability was high, ICC (3,1) was 0.92 with a 95% CI of 0.72 --> 0.98. There was reasonable agreement between measures on the Bland and Altman test, as d was -0.63 cm2, the 95% CI for d was -1.4 --> 0.14 cm2, the SDdiff was 1.08 cm2, the 95% limits of agreement -2.73 --> 1.53 cm2 and the reliability coefficient was 2.4. Between-scans repeatability was high, ICCs (1,1) were 0.94 and 0.93 with 95% CIs of 0.8 --> 0.99 and 0.75 --> 0.98, for days 1 and 2 respectively. Measures showed good agreement on the Bland and Altman test: d for day 1 was 0.15 cm2 and for day 2 it was -0.32 cm2, the 95% CIs for d were -0.51 --> 0.81 cm2 for day 1 and -0.98 --> 0.34 cm2 for day 2; SDdiff was 0.93 cm2 for both days, the 95% imits of agreement were -1.71 --> 2.01 cm2 for day 1 and -2.18 --> 1.54 cm2 for day 2; the reliability coefficient was 1.80 for day 1 and 1.88 for day 2. The between-days ICC (1,2) was 0.92 and the 95% CI 0.69 --> 0.98. The d was -0.98 cm2, the SDdiff was 1.25 cm2 with 95% limits of agreement of -3.48 --> 1.52 cm2 and the reliability coefficient 2.8. The 95% CI for d (-1.88 --> -0.08 cm2) and the distribution graph showed a bias towards a larger measurement on day 2. CONCLUSIONS: The ICC and Bland and Altman tests are appropriate for analysis of reliability studies of similar design to that described, but neither test alone provides sufficient information and it is recommended that both are used.

Adult↗

Reliability of an idiographic Q-sort measure of defense mechanisms.

Despite the important insights the concept of defense mechanisms may offer to our understanding of human behavior, no standardized definitions of defense mechanisms have been universally accepted. Inconsistencies in the definition and conceptualization of defense mechanisms has limited the practical utility of research involving these constructs. In addition, lack of interrater reliability, use of anecdotal evidence, and reliance on self-reports has retarded their investigation. Conducting methodologically rigorous investigations with a psychometrically sound instrument is the first step in addressing some of the issues concerning defense mechanisms and their theoretical postulates. This study was conducted to determine the reliability associated with the Defense-Q, an observer-based Q-sort measure of defense mechanisms. Thirty participants who had undergone an interpersonally stressful interview (the Type A Structured Interview; Rosenman, 1978) were rated by 11 trained coders, both for their use of the 25 defense mechanisms and for their ego strength. Reliability was assessed using Cronbach's alpha. Individual defense mechanisms demonstrated reliability ranging from .28 (undoing) to .92 (humor), with an average reliability of .73. Coder reliability ranged from .63 to .76, with an average of .69. These results indicate that defense mechanisms can be reliably assessed by the Defense-Q. Reliability of the Defense-Q is compared to existing observational measures of defense mechanisms.

Journal Article↗

Technical reliability assessment of three accelerometer models in a mechanical setup.

PURPOSE: To determine which of the three most commonly used accelerometer models has the best intra- and interinstrument reliability using a mechanical laboratory setup. Secondly, to determine the effects that acceleration and frequency have on these reliability measures. METHODS: Three experiments were performed. In the first, five each of the Actical, Actigraph, and RT3 accelerometers were placed on a hydraulic shaker plate and simultaneously accelerated in the vertical plane at varying accelerations and frequencies. Six different conditions of varying intensity were used to produce a range of accelerometer counts. Reliability was calculated using standard deviation, standard error of the measurement, coefficient of variation, and intraclass correlation coefficients. In the second and third experiments, 39 Actical and 50 Actigraph accelerometers were put through the same six conditions. RESULTS: Experiment 1 showed poor reliability in the RT3 (intra- and interinstrument CV > 40%). Experiments 2 and 3 clearly indicated that the Actical (CVintra = 0.5%, CVinter = 5.4%) was more reliable than the Actigraph (CVintra = 3.2%, CVinter = 8.6%). Variability in the Actical was negatively related to the acceleration of the condition, whereas no relationship was found between acceleration and reliability in the Actigraph. Variability in the Actigraph was negatively related to the frequency of the condition, whereas no relationship was found between frequency and reliability in the Actical. CONCLUSION: Of the three accelerometer models measured in this study, the Actical had the best intra- and interinstrument reliability. However, discrepant trends in the variability of Actical and Actigraph counts across accelerations and frequencies preclude the selection of a superior model. More work is needed to understand why accelerometers designed to measure the same thing behave so differently.

Acceleration↗

The effect of training on rater reliability on the scoring of the NART.

OBJECTIVES: This study investigates whether the accuracy of judging National Adult Reading Test (NART) words known to have lower inter-rater reliability can be improved by training and use of the pronunciation guide. DESIGN: Two groups (Experimental and Control), were compared with three repeated measures: Occasion (first and second i.e. 'post-training'), Word Reliability (high and low) and Pronunciation Guide (without and with guide). METHODS: Ten words were selected from the NART: five lower reliability and five high reliability words. These were presented aurally in correct and incorrect form to participants (N = 20) who judged correctness of pronunciation without or with a pronunciation guide. Each group repeated the task again, the Experimental group having received training. RESULTS: Accuracy was significantly worse for the low reliability words. The experimental group's accuracy was significantly better after training than the control group's and their own performance prior to training. The use of the guide enhanced accuracy, particularly for the low reliability words. CONCLUSION: Training in administration of the NART improves raters' accuracy and use of the pronunciation guide. This offers an alternative to the suggestion of improving the NART's reliability by replacing lower reliability words and therefore would avoid the need to re-standardize a modified test.

Adult↗

The reliability of Form 90: an instrument for assessing alcohol treatment outcome.

OBJECTIVE: Project MATCH is a randomized clinical trial consisting of five outpatient and five aftercare units at nine sites. Of importance in this multisite trial examining the efficacy of client-treatment matching was the cross- and within-site reliability of the structured interview used to assess alcohol treatment outcomes, the Form 90. Evaluation of the reliability of Form 90 is the subject of this article. METHOD: The reliability of Form 90 was evaluated in two test-retest studies. The cross-site reliability study consisted of 70 paired test-retest interviews conducted by different interviewers. Clients for this study were recruited from inpatient, outpatient and college settings. The within-site reliability study had a total of 108 paired test-retest interviews, with 54 of the retests conducted by different interviewers and 54 by the same interviewer. Clients for this study were most often presenting for alcohol treatment at the nine sites and were selected to be representative of the larger Project MATCH sample. RESULTS: Good-to-excellent reliability was found for all key summary measures of alcohol consumption and psychosocial functioning, and most frequently used illicit drugs had moderate reliability. No decay in consistency of self-reported drinking was found at more distal points from dates of test-retest interviews. Application of 68% confidence intervals for primary alcohol consumption measures suggests that trained researchers and clinicians can obtain consistent information regarding client drinking. CONCLUSIONS: Form 90 appears to be a reliable instrument for alcohol treatment assessment research when interviewers have received careful training and supervision in its use.

Adult↗

Testing reliability of plaque and gingival indices. Two methods.

This investigation was undertaken to compare two methods of interexaminer and intraexaminer reliability in the evaluation of Plaque and Gingival Indices prior to a study of toothbrushing. Inter-/intraexaminer reliabilities were compared using a projected slide series consisting of 40 slides of clinical examples of gingival inflammation and plaque accumulation. Time between assessments was three weeks. Using the slide technique, intraexaminer reliability was established for: (1) Gingival Indices and (2) Plaque Indices. Interexaminer reliability was also established for Gingival Indices. Interexaminer reliability could not be established for Plaque Indices on the first assessment but was established on the post-assessment. Intraexaminer reliability was also determined through clinical examinations of patients. A third clinician was used to manipulate the tissue while investigators evaluated bleeding on provocation and plaque accumulation. Significant results were established for the Gingival Indices and Plaque Indices. Results of this investigation suggest that significant inter-/intraexaminer reliabilities may be obtained for gingival indices using the slide technique. In addition, the clinic technique appeared useful for assessing interexaminer reliability for Gingival Indices. Plaque Indices using the slide technique required more practice than those using the clinic technique.

Dental Health Surveys↗

Measures of reliability in sports medicine and science.

Reliability refers to the reproducibility of values of a test, assay or other measurement in repeated trials on the same individuals. Better reliability implies better precision of single measurements and better tracking of changes in measurements in research or practical settings. The main measures of reliability are within-subject random variation, systematic change in the mean, and retest correlation. A simple, adaptable form of within-subject variation is the typical (standard) error of measurement: the standard deviation of an individual's repeated measurements. For many measurements in sports medicine and science, the typical error is best expressed as a coefficient of variation (percentage of the mean). A biased, more limited form of within-subject variation is the limits of agreement: the 95% likely range of change of an individual's measurements between 2 trials. Systematic changes in the mean of a measure between consecutive trials represent such effects as learning, motivation or fatigue; these changes need to be eliminated from estimates of within-subject variation. Retest correlation is difficult to interpret, mainly because its value is sensitive to the heterogeneity of the sample of participants. Uses of reliability include decision-making when monitoring individuals, comparison of tests or equipment, estimation of sample size in experiments and estimation of the magnitude of individual differences in the response to a treatment. Reasonable precision for estimates of reliability requires approximately 50 study participants and at least 3 trials. Studies aimed at assessing variation in reliability between tests or equipment require complex designs and analyses that researchers seldom perform correctly. A wider understanding of reliability and adoption of the typical error as the standard measure of reliability would improve the assessment of tests and equipment in our disciplines.

Algorithms↗

Short- and long-term reliability of information on previous illness and family history as compared with that on smoking and drinking habits in questionnaire surveys.

To assess the reliability of responses to questionnaires regarding previous illness, family history of cancer, and smoking and drinking habits, we repeated questionnaire surveys four times at intervals of 2 weeks and 1 year (short-term), and 4.5 years (long-term) among 440 subjects aged 40-69. The reliability was assessed using kappa statistic. Kappa was calculated both for complete data and data including missing values. The changes of mode of pre-after paired responses were also investigated. Our results from complete data showed both short- and long-term reliabilities of replies regarding smoking or drinking were excellent (mean kappa 0.85-0.99). The reliability of previous illness was excellent except for stroke for short-term intervals (mean kappa 0.85-1.00), but varied depending on the kinds of illness with long-term intervals (mean kappa -0.01-0.75). Responses to family history had fair to excellent short-term reliability (mean kappa 0.54-0.85). Inclusion of missing value as an independent category reduced reliability remarkably. Subjects stating absence of medical history were more likely to have missing values for this item than subjects with some history. In conclusion, the reliability for information given on previous illness was as good as that on smoking and drinking for a short interval, but was lower for a long-interval probably due to the development of new cases. The reliability of a family history on cancer was slightly poorer than that of individual's previous cancer or other illnesses and that of smoking and drinking even for a short interval.

Adult↗

Reliability of health information on the Internet: an examination of experts' ratings.

BACKGROUND: The use of medical experts in rating the content of health-related sites on the Internet has flourished in recent years. In this research, it has been common practice to use a single medical expert to rate the content of the Web sites. In many cases, the expert has rated the Internet health information as poor, and even potentially dangerous. However, one problem with this approach is that there is no guarantee that other medical experts will rate the sites in a similar manner. OBJECTIVES: The aim was to assess the reliability of medical experts' judgments of threads in an Internet newsgroup related to a common disease. A secondary aim was to show the limitations of commonly-used statistics for measuring reliability (eg, kappa). METHODS: The participants in this study were 5 medical doctors, who worked in a specialist unit dedicated to the treatment of the disease. They each rated the information contained in newsgroup threads using a 6-point scale designed by the experts themselves. Their ratings were analyzed for reliability using a number of statistics: Cohen's kappa, gamma, Kendall's W, and Cronbach's alpha. RESULTS: Reliability was absent for ratings of questions, and low for ratings of responses. The various measures of reliability used gave conflicting results. No measure produced high reliability. CONCLUSIONS: The medical experts showed a low agreement when rating the postings from the newsgroup. Hence, it is important to test inter-rater reliability in research assessing the accuracy and quality of health-related information on the Internet. A discussion of the different measures of agreement that could be used reveals that the choice of statistic can be problematic. It is therefore important to consider the assumptions underlying a measure of reliability before using it. Often, more than one measure will be needed for "triangulation" purposes.

Delivery of Health Care↗

Intramachine and intermachine reliability for selected dynamic muscle performance tests.

The Cybex 6000 isokinetic dynamometer is a new isokinetic device for which no published reports of reliability have been presented in the literature. In addition, the manufacturer not only claims that the new Cybex 6000 is reliable but that torque data obtained from the Cybex 6000 are consistent with data obtained from past Cybex systems, such as the Cybex II. The purpose of this study was to investigate the intramachine reliability of the Cybex 6000 to itself and the intermachine reliability of the Cybex 6000 and the Cybex II. Data on peak torque, work, and power were collected using the Cybex 6000, and data on peak torque were obtained using the Cybex II for knee flexion and extension in 20 volunteers (10 males, 10 females). Subjects were tested three times, twice on the Cybex 6000 and once on the Cybex II, approximately 1 week apart across a 3-week period of time at angular velocities of 60, 180, and 300 degrees/sec. Data were analyzed using intraclass correlations. Results indicated that the majority of test-retest correlation coefficients for all parameters for intramachine reliability of the Cybex 6000 were above .90. Comparing peak torque obtained with the Cybex 6000 to that obtained with the Cybex II (intermachine reliability), correlation coefficients ranged from .72 to .89. In conclusion, information obtained on the Cybex 6000 appears to be quite reliable in a test-retest situation using the same equipment and moderately reliable when compared to the Cybex II. Clinical implications for these results are discussed.

Adult↗