PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Reliability of binocular vision measurements used in the classification of convergence insufficiency.

PURPOSE: To evaluate the reliability of binocular vision measurements used in the classification of convergence insufficiency. METHODS: Two examiners tested 20 fifth and sixth graders in a school setting who passed a screening of visual acuity, refraction, and binocularity. The tests, conducted using a standard protocol, consisted of von Graefe near heterophoria (NH), phorometric positive fusional vergence (PFV), nearpoint of convergence (NPC), and monocular pushup accommodative amplitude (AA). Each examiner measured each child three consecutive times for each test, on two separate occasions, spaced approximately 1 week apart. Intraexaminer and interexaminer agreement was assessed using intraclass correlation coefficients (ICC), the median absolute difference (MAD), and the coefficient of repeatability (COR). RESULTS: The within-session reliability of the NH (ICC: 0.95 to 0.99), NPC (ICC: 0.94 to 0.98), and AA (ICC: 0.88 to 0.95) were good, whereas the PFV was less reliable (ICC: 0.71 to 0.94). The intraexaminer reliability between sessions was good for the NPC (ICC: 0.92 and 0.89), less reliable for NH (ICC: 0.81 and 0.81) and AA (ICC: 0.89 and 0.69), and much less reliable for PFV break (ICC: 0.59 and 0.53). Typical between-session PFV differences (MAD) were between 3 and 4 delta, whereas the COR differences were as large as 12 delta. CONCLUSIONS: Three of the four measures (NH, NPC, and AA) often used in the classification of convergence insufficiency generally have good within-session and between-session reliability. The PFV break was found to have only fair reliability with clinically significant differences between sessions. The large potential test-retest differences found could complicate clinical decision-making in regards to diagnosis and treatment.

Accommodation, Ocular↗

Interrater and intrarater reliability in the measurement of kyphosis in postmenopausal women with osteoporosis.

STUDY DESIGN: A reliability study was performed using repeated random measurements involving three observers, 26 subjects and three instruments. OBJECTIVES: To determine the most reliable, cost-effective, noninvasive, and clinically feasible method of measuring spinal kyphosis. SUMMARY OF BACKGROUND DATA: The most clinically useful, noninvasive and reliable method of measuring postural deformity in spinal osteoporosis (kyphosis) remains unqualified. Despite traditional use of costly, invasive roentgenographs for the evaluation of spinal kyphosis, the reliability of this method remains questionable. METHODS: Twenty-six postmenopausal women with known bone mineral density and a diagnosis of osteoporosis were recruited from the Osteoporosis Program at Women's College Hospital, Toronto, Canada. Non-invasive measurements of thoracic kyphosis were obtained by three trained examiners using the DeBrunner's kyphometer and the flexicurve ruler. The intrarater and interrater reliability of and between each method was compared, using roentgenographic films obtained in the sagittal plane. Spinal posture was classified according to the method of Itoi (1990). Statistical computations were performed using SAS statistical software. RESULTS: Consistent measurements were obtained with the DeBrunner's kyphometer and the flexicurve ruler by each observer, according to the results of critical two-way analysis of variance (Intraclass Correlation Coefficient 2, 1). Measurements in two subgroups, healthy backs (n = 11) and rounded backs (n = 13), showed consistent use of each noninvasive instrument with some examiner preference for specific tools. There was marginally better intrarater and interrater reliability using the DeBrunner's kyphometer compared with that obtained with the flexicurve ruler. Two-way analysis of variance (Intraclass Correlation Coefficient 2, 1) of collapsed data showed no significant difference in the reliability of the kyphometer, flexicurve ruler, or roentgenographs in the measurement of thoracic kyphosis. CONCLUSIONS: The flexicurve ruler and DeBrunner's kyphometer had the closest agreement in the measurement of spinal kyphosis. The kyphometer demonstrated the least variation in intrarater and interrater reliability when compared with the flexicurve ruler and roentgenographs. The flexicurve ruler permits qualitative assessment of posture, however, and is the most cost-effective instrument. The results of this study challenge the traditional belief that roentgenographic analysis is the best method for evaluating spinal kyphosis. The DeBrunner's kyphometer and flexible ruler may represent viable, cost-effective and noninvasive alternatives to roentgenographic evaluation of spinal kyphosis.

Aged↗

Range of motion and lordosis of the lumbar spine: reliability of measurement and normative values.

STUDY DESIGN: Repeated measures for intratester reliability were performed. OBJECTIVES: To investigate the intratester reliability of a new measurement technique that evaluates lumbar range of motion in three planes using a pelvic restraint device, and to examine the reliability of lumbar lordosis measurement by inclinometer technique. Preliminary normative data on lumbar range of motion and lumbar lordosis were collected for comparison with the findings of previous studies. SUMMARY OF BACKGROUND DATA: Various noninvasive measurement methods have been developed for recording lumbar range of motion. However, pelvic movement was not effectively restricted during the use of these measurement techniques. The use of the pelvic restraint device to measure lumbar range of motion has not been investigated previously. Very few studies have investigated the reliability of quantifying lumbar lordosis by the inclinometer technique. METHODS: Normative values were measured in 35 healthy men, and 12 of these subjects were included for the reliability study. Pelvic motion was limited by the pelvic restraint device during lumbar range of motion measurement in standing. An inclinometer was used for evaluation of lumbar flexion, extension, lateral flexion, and lumbar lordosis, whereas a lumbar rotameter was used to measure axial rotation. RESULTS: Good intratester reliability was shown in the lumbar range of motion and lordosis measurement. Most of the intraclass correlation coefficient and Pearson's r values (accompanied with nonsignificant paired t tests) were greater than 0.9, and most of the intrasubject coefficients of variation were less than 10%. The values of lumbar range of motion in three planes and lumbar lordosis found in the current study were comparable with those from most of the previous studies on these measurements in the normal population. CONCLUSIONS: Inclinometer and lumbar rotameter measurements with the use of a pelvic restraint device are reliable for measuring lumbar spine range of motion. Use of the inclinometer technique to record lumbar lordosis also is a reliable measure.

Adult↗

Shoulder abduction strength measurement in football players: reliability and validity of two field tests.

Musculoskeletal and neurologic injuries affecting shoulder strength are common in contact sports. Full-strength recovery is desired before resumption of competition. On-field assessment of shoulder strength is usually done by manual muscle testing, which lacks sensitivity and reliability. Our objective was to determine the reliability and validity of two field instruments capable of quantifying shoulder abduction strength. Twenty junior football players underwent bilateral isokinetic (60 degrees/s) and isometric shoulder abduction strength measurements using a Cybex 340 isokinetic dynamometer. Test-retest measurements of both shoulders of each player were made using strain gauge (SG) and handheld dynamometer (HHD) instruments. Players were tested during rested and competition conditions. Within and between session reliabilities were calculated using the intraclass coefficient, and validity was assessed using Pearson's correlation coefficient. Overall reliability for each device was calculated using Lisrel analysis. SG was found to be superior to HHD in overall reliability and validity. Within-session reliability in the rested and competition states was 0.75 and 0.78, respectively, for SG and 0.60 and 0.81, respectively, for HHD. Between-session reliability in the rested and competition states dropped to 0.51 and 0.63, respectively, for SG and 0.55 and 0.70, respectively, for HHD. Validity was 0.41 and 0.70 for SG when correlated with Cybex at 0 degree and 60 degrees/s respectively. Validity for HHD was 0.28 and 0.42 for Cybex speeds of 0 degree and 60 degrees/s, respectively. SG reliability and validity were similar when testing was done one shoulder at a time or both shoulders concurrently.(ABSTRACT TRUNCATED AT 250 WORDS)

Adolescent↗

Reliability of computed tomography measurements of paraspinal muscle cross-sectional area and density in patients with chronic low back pain.

STUDY DESIGN: A reliability study was conducted. OBJECTIVE: To estimate measurement errors related to equipment and the observer in computed tomography measurements of cross-sectional area and density of paraspinal muscles. Interobserver reliability was not investigated in the current study. SUMMARY OF BACKGROUND DATA: Computer tomography (CT) had been used to measure the cross-sectional area and degeneration of the back muscles in patients with low back pain. METHODS: This study included 31 patients, mean age 47 years, with chronic low back pain. The measurements comprised cross-sectional area (cm2) and density (Hounsfield units [HU]) of the paraspinal muscles at Th12-L1, L3-L4, and L4-L5. To measure the reliability of the equipment and the observer (total reliability), two independent CT scans were performed for each patient. The radiologist traced the cross-sectional area twice within 2 weeks for measurement of the intraobserver reliability. RESULTS: There were no significant differences in the assessments between the first and second CT scans, or between the radiologist's two measurements of the identical slices. The critical difference for the total reliability ranged from 11.3 to 22.8 for the density and from 10.0 to 16.0 for the cross-sectional area. For the cross-sectional area, the measurement error associated with the observer was higher than for the equipment. For the density, the measurement error related to the equipment was higher. The main measurement error was associated with the radiologist for the cross-sectional area and with the CT scanner for the density. CONCLUSIONS: The reliability of the CT scan for measuring the cross-sectional area and density of the back muscles is acceptable. The authors do not know definitely whether their results can be generalized because the interobserver and intermachine reliabilities were not investigated.

Adult↗

An examination of the reliability of a classification algorithm for subgrouping patients with low back pain.

STUDY DESIGN: Test-retest design to examine interrater reliability. OBJECTIVE: Examine the interrater reliability of individual examination items and a classification decision-making algorithm using physical therapists with varying levels of experience. SUMMARY OF BACKGROUND DATA: Classifying patients based on clusters of examination findings has shown promise for improving outcomes. Examining the reliability of examination items and the classification decision-making algorithm may improve the reproducibility of classification methods. METHODS: Patients with low back pain less than 90 days in duration participating in a randomized trial were examined on separate days by different examiners. Interrater reliability of individual examination items important for classification was examined in clinically stable patients using kappa coefficients and intraclass correlation coefficients. The findings from the first examination were used to classify each patient using the decision-making algorithm by clinicians with varying amounts of experience. The reliability of the classification algorithm was examined with kappa coefficients. RESULTS: A total of 123 patients participated (mean age 37.7 [+/-10.7] years, 44% female), 60 (49%) remained stable between examinations. Reliability of range of motion, centralization/peripheralization judgments with flexion and extension, and the instability test were moderate to excellent. Reliability of centralization/peripheralization judgments with repeated or sustained extension or aberrant movement judgments were fair to poor. Overall agreement on classification decisions was 76% (kappa = 0.60, 95% confidence interval 0.56, 0.64), with no significant differences based on level of experience. CONCLUSION: Reliability of the classification algorithm was good. Further research is needed to identify sources of disagreements and improve reproducibility.

Adolescent↗

Measurement of lumbar lordosis: inter-rater reliability, minimum detectable change and longitudinal variation.

STUDY DESIGN: Repeated measures design to examine reliability and longitudinal variation of lumbar lordosis measurement. OBJECTIVES: To determine the interrater reliability, minimum detectable change (MDC) and longitudinal variation of the Cobb method for measuring lumbar lordosis using standardized rules. SUMMARY OF BACKGROUND DATA: The reliability of the 4-line Cobb method for measuring lumbar lordosis was not examined when standardized rules were instituted for drawing the lines. METHODS: A random sample of participants was selected from the Pittsburgh clinic of the multicenter Study of Osteoporotic Fractures for radiographic measurement of lumbar lordosis reliability (n=48) and stability (n=109). A standardized version of the 4-line Cobb method was used for all measurements of lordosis. The Intraclass Correlation Coefficient (ICC) was used to calculate interrater reliability for lordosis and to measure the stability of this measure over an approximate 2-year-time period. The standard error of measurement and MDC were calculated for lordosis measurement based on the ICC value. RESULTS: The interrater reliability coefficient for lumbar lordosis was in the excellent range (ICC=0.98; 95% CI: 0.95, 0.99). The MDC based on measurements between raters was 3.90 degrees. The ICC value for the stability, or reliability from time 1 to time 2, of lordosis measurement over time was 0.81 (95% CI: 0.74, 0.87). CONCLUSION: This study demonstrates that the 4-line Cobb method can be a highly reliable and precise method for measuring lumbar lordosis if standardized procedures are used. The Cobb method has an MDC that is appropriate for clinical use. Also, there is minimal longitudinal variation in lordosis measurements over a 2-year period.

Aged↗

Lifetime DSM-IV diagnosis of alcohol, cannabis, cocaine and opiate dependence: six-month reliability in a multi-site clinical sample.

Psychiatric research increasingly emphasizes the diagnosis of symptoms and syndromes on a longitudinal basis. This study tests the reliability of lifetime DSM-IV diagnoses of alcohol, cannabis, cocaine and opiate dependence. The CIDI-SAM was administered at intervals not less than six months apart to a multi-site sample of 201 clinical respondents. The reliability of lifetime diagnosis of the syndromes, of the criteria which constitute the syndromes, and of the ages of onset reported for the criteria and for the dependence syndromes as a whole, were studied and the effects of patient characteristics suspected to degrade reliability were examined. There was generally good agreement, statistically, at both the syndrome and criterion level between the two interviews. Lifetime diagnoses for three of the drugs--alcohol, cannabis and opiates--were made at or near levels of agreement generally considered excellent under less strict testing conditions, and cocaine dependence was only marginally below this level. Most criteria showed good reliability and all delivered about equal results when averaged across the four substances, although a relationship between reliability and centrality of the symptom to the individual drug abuse pattern was found. Age of onset was almost uniformly highly reliable. Most patient characteristics bore no detectable relationship to reliability, although patients with multiple drug use patterns may warrant more careful probing by interviewers. Overall, these data indicate that lifetime symptoms and diagnoses can be queried reliably, although they must be reported with less confidence than current state diagnoses.

Age of Onset↗

Determinants of reliability in psychiatric surveys of children aged 6-12.

The reliability of young children's self reports of psychiatric information is a concern of epidemiologists and clinicians alike. This paper explores the determinants of test-retest reliability in a sample of children from the general population using reliability coefficients constructed from a kappa statistic. Age, cognitive ability, and gender are related to consistency of reports in a test-retest paradigm. Controlling for age, cognitive ability and gender, children report more reliably on observable behaviors, and less reliably on questions involving unspecified time, reflections of one's own thoughts, and comparison of themselves with others. The reliability of reports of emotions lies between these two extremes. Surprisingly, sentence length of up to 40 words and psychiatric impairment of the child as measured by the Child Global Assessment Scale did not influence reliability. As might be expected, parents' reports of their children are more reliable than their children's reports.

Affective Symptoms↗

Reliability, validity, and clinical utility of the Migraine-ACT questionnaire.

BACKGROUND: The 4-item Migraine-ACT questionnaire is an assessment tool for use by primary care physicians to identify patients who require a change in their current acute migraine treatment. It has been shown to be easy to use, and to be reliable and accurate in its assessments. OBJECTIVES: To further analyze the Migraine-ACT study database, providing additional information on the reliability, validity, and potential clinical utility of the questionnaire. METHODS: Reliability was assessed by recording the distribution of Migraine-ACT scores recorded at baseline and 1 week later (test-retest reliability). Analyses of consistency of Migraine-ACT scores were conducted on the total sample of patients and for the separate centers, using Pearson and Spearman correlations. Validity was assessed by comparing the t-discrimination values for clinically relevant questions within domains of the original 27-item questionnaire. Reliability and validity were also assessed by constructing an "alternative" (Form B) Migraine-ACT questionnaire, derived from an analysis of the second-best items in each domain in the original study data. Clinical utility was assessed using Pearson pairwise correlations to compare Migraine-ACT scores with clinically defined criteria as analyzed by the SF-36 Quality of Life questionnaire, the Migraine Disability Assessment (MIDAS) questionnaire, and the Migraine Therapy Assessment (MTAQ) questionnaire. RESULTS: The distribution of Migraine-ACT scores between the 2 completions of the questionnaire was consistent for the total sample (test-retest reliability, r= .81) and between the individual countries (r= .61 to .92). In this study, the validity (assessed as t-discrimination) of the Migraine-ACT "impact" and "global assessment of relief" questions were markedly higher than those of other endpoints used in migraine clinical studies. The Form B Migraine-ACT questionnaire was almost as reliable and accurate as the original Form A questionnaire. The distribution of Migraine-ACT scores was: 0 = 12.6%, 1 = 13.7%, 2 = 14.7%, 3 = 20.5%, and 4 = 38.4%. The change in Migraine-ACT score correlated with, and had a linear relationship with changes in SF-36, MIDAS, and MTAQ scores, and indicated that a Migraine-ACT score of <or=2 corresponded with a need to consider changing the patient's acute medication. About 40% of the migraine patients in the study scored <or=2 and may have had significant unmet treatment needs. CONCLUSIONS: These data confirm the excellent reliability and validity of the Migraine-ACT questionnaire and provide further evidence for its utility in clinical practice.

Acute Disease↗

Effects of cognitive impairment on the reliability of geriatric assessments in nursing homes.

OBJECTIVE: To explore the relationship between an elderly subject's cognitive status and the reliability of multidimensional assessment data. DESIGN: Survey, with cognitive status as the independent variable and interrater reliability as dependent variable. SETTING: Medicare/Medicaid-certified nursing homes. PARTICIPANTS: 147 residents age 65 or older. MEASUREMENTS: Dual assessments of elderly nursing home residents were performed by nurse assessors using the Health Care Financing Administration's new Minimum Data Set for Nursing Home Resident Assessment and Care Screening (MDS). Assessments were classified on the basis of residents' cognitive status, and levels of disagreement between assessors were analyzed. MAIN RESULTS: Overall assessment reliability, agreement concerning a resident's activities of daily living status, and the reliability of estimates of his or her communication skills and sensory abilities were significantly affected by a resident's cognitive status. The presence of cognitive impairment made these measurements less reliable--especially those related to communication skills, vision, and hearing. CONCLUSIONS: Assessments of residents suffering from cognitive impairment were significantly less reliable than assessments of cognitively intact residents. However, these differences in reliability were not uniform across all assessment domains. When treating the cognitively impaired elderly, clinicians must exercise caution in their reliance on standardized measurements that may be less reliable for this population.

Activities of Daily Living↗

The reliability and validity of birth certificates.

OBJECTIVES: To summarize the reliability and validity of birth certificate variables and encourage nurses to spearhead data improvement. DATA SOURCES: A Medline key word search of reliability and validity of birth certificate, and a reference review of more than 60 articles were done. STUDY SELECTION: Twenty-four primary research studies of U.S. birth certificates that involved validity or reliability assessment. DATA EXTRACTION: Studies were reviewed, critiqued, and organized as either a reliability or a validity study and then grouped by birth certificate variable. DATA SYNTHESIS: The reliability and validity of birth certificate data vary considerably by item. Insurance, birthweight, Apgar score, and delivery method are more reliable than prenatal visits, care, and maternal complications. Tobacco and alcohol use, obstetric procedures, and delivery events are unreliable. Birth certificates are not valid sources of information on tobacco and alcohol use, prenatal care, maternal risk, pregnancy complications, labor, and delivery. CONCLUSIONS: Birth certificates are a key data source for identifying causes of increasing U.S. infant mortality but have serious reliability and validity problems. Nurses are with mothers and infants at birth, so they are in a unique position to improve data quality and spread the word about the importance of reliable and valid data. Recommendations to improve data are presented.

Alcohol Drinking↗

Reliability of a standardized and expanded Brief Psychiatric Rating Scale: a replication study.

This study aimed to determine the replicability of the interrater reliability coefficients obtained with a standardized and expanded Brief Psychiatric Rating Scale (BPRS-E) in a 1991 psychometric evaluation. Furthermore, intrarater reliability was assessed. At item level, interrater concordance turned out to be satisfactory for most of the BPRS-E items. However, only a few of the items reached acceptable chance-corrected coefficients. In contrast to the previous study, the anxiety-depression subscale met the standard of acceptable interrater reliability in the present study. As in the 1991 study, the 10-item psychotic disintegration scale as well as BPRS-18 global scores met (or closely approximated) this standard. The 6 additional items of BPRS-E did not contribute to the scale's reliability. Joining the samples of the 1991 and replication studies (to cover the range of symptoms' severity and heterogeneity more fully) did not improve interrater reliability. Intrarater reliability coefficients were globally comparable to interrater reliability coefficients. In all, the results of this replication study suggest that only the anxiety-depression subscale, the 10-item psychotic disintegration scale and the BPRS-18 global scale can be used reliably in unselected groups of psychiatric inpatients in acute distress.

Adolescent↗

Factors affecting reliability coefficients of health attitude scales.

This study determined the minimum number of health attitude items and minimum sample size required to achieve maximum scale reliability coefficients, using different methods of estimating reliability. A 54-item alcohol attitude scale was administered to 700 participants. The scale produced .96 and .91 reliability coefficients, using the Cronbach Alpha (CA) and the Split-half (S-B) methods, respectively. A computer program randomly selected groups of participants and items from the pool of participants and items using different increments. A matrix of coefficients of reliability for both methods was calculated for different groups of items and sample size. To replicate the study, a 30-item cancer attitude scale was administered to more than 1,000 representative participants and produced reliability coefficients of .94 (using CA) and .82 (using S-B). The same computer and statistical procedures were repeated for the second data set. Results from both analyses consistently demonstrated that sample size has an insignificant effect on the coefficient values of reliability. Reliability increased as the number of items reached 18. Adding more items only negligibly increased the coefficients. Overall, the CA method consistently produced higher coefficient values of reliability compared to the S-B method.

Attitude to Health↗

Effect of a patient training video on visual field test reliability.

AIMS: To evaluate the effect of a visual field test educational video on the reliability of the first automated visual field test of new patients. METHODS: A prospective, randomised, controlled trial of an educational video on visual field test reliability of patients referred to the hospital eye service for suspected glaucoma was undertaken. Patients were randomised to either watch an educational video or a control group with no video. The video group was shown a 4.5 minute audiovisual presentation to familiarize them with the various aspects of visual field examination with particular emphasis on sources of unreliability. Reliability was determined using standard criteria of fixation loss rate less than 20%, false positive responses less than 33%, and false negative responses less than 33%. RESULTS: 244 patients were recruited; 112 in the video group and 132 in the control group with no significant between group difference in age, sex, and density of field defects. A significant improvement in reliability (p=0.015) was observed in the group exposed to the video with 85 (75.9%) patients having reliable results compared to 81 (61.4%) in the control group. The difference was not significant for the right (first tested) eye with 93 (83.0%) of the visual fields reliable in the video group compared to 106 (80.0%) in the control group (p = 0.583), but was significant for the left (second tested) eye with 97 (86.6 %) of the video group reliable versus 97 (73.5%) of the control group (p = 0.011). CONCLUSIONS: The use of a brief, audiovisual patient information guide on taking the visual field test produced an improvement in patient reliability for individuals tested for the first time. In this trial the use of the video had most of its impact by reducing the number of unreliable fields from the second tested eye.

Aged↗

Reliability of the National Institutes of Health Stroke Scale. Extension to non-neurologists in the context of a clinical trial.

BACKGROUND AND PURPOSE: The reliability of the National Institutes of Health Stroke Scale (NIHSS) has been established through testing its use in live and videotaped patients. This reliability testing has primarily focused on the use of the scale by neurologists. We sought to determine the reliability of the NIHSS as used by non-neurologists in the context of a clinical trial. METHODS: In anticipation of the initiation of a randomized trial of a new therapy for patients with acute ischemic stroke, 30 physician investigators (30% of whom were not neurologists) and 29 non-physician study coordinators were trained in the use of the NIHSS at an informational and training conference using standardized videotaped patient examinations. A series of 4 patients were rated initially. After 3 months, the same 4 patients were rerated, providing a measure of intraobserver reliability. An additional series of 4 new patients were also rated after 3 months and, with the initial 4 ratings, provided data for assessment of interobserver reliability. RESULTS: Overall, 28% of the raters had previous experience with the NIHSS, and 22% had previously used the videotapes as used in the present trial. The coefficients of determination (r2) were each greater than .95 when the means of the two ratings of the same 4 cases were compared between (1) neurologists and other types of physicians, (2) physicians and study coordinators, (3) raters who had prior experience with the NIHSS and those without prior experience, and (4) raters who had used the videotapes in the past and those who had never viewed the tapes. The calculated r2s were greater than .98 for the initial rating of the first 4 cases and for the later rating of the 4 new cases. The slopes of the regression lines were all near 1, indicating that the raters were similarly calibrated. The intraclass correlation coefficients were .93 and .95, reflecting high levels of intraobserver and interobserver reliability. CONCLUSIONS: These data extend the previously demonstrated reliability of the NIHSS to non-neurologists and show that both a variety of physician investigators and nurse study coordinators can be rapidly trained to reliably apply the scale in the context of an actual clinical trial.

Cerebrovascular Disorders↗

Reliability of visual analog and verbal descriptor scales for "objective" measurement of temporomandibular disorder pain.

Eight dentists viewed standardized videotapes showing palpations of the temporomandibular joint and muscles of mastication and recorded their judgments concerning the amount of pain the patient was experiencing. Judgments were recorded using a four-point verbal descriptor scale (VDS) ("none", "mild", "moderate", "severe" pain) or a 100-mm visual analog scale (VAS) anchored with the terms "no pain" and "worst pain possible". Test/re-test reliability over a one-week period and interjudge reliabilities were calculated for each scale; reliabilities of the two scales were directly compared based on the statistical equivalence of weighted kappa and the Intraclass Correlation Coefficient. Neither scale showed satisfactory reliability. Median test/re-test reliabilities were k = 0.590 for the VDS and r = 0.822 for the VAS. Interjudge reliabilities averaged k = 0.394 for the VDS and r = 0.735 for the VAS. Direct comparison of reliabilities for the two scales showed no clear advantage for either scale. The marginal reliabilities of these scales, when used by dentists to quantify the patient's pain, suggest that neither scale should be regarded as an "objective" pain measure.

Decision Making↗

The reliability of the Diabetes Care Profile for African Americans.

The Diabetes Care Profile (DCP) is an instrument used to assess social and psychological factors related to diabetes and its treatment. The reliability of the DCP was established in populations consisting primarily of Caucasians with type 2 diabetes. This study tests whether the DCP is a reliable instrument for African Americans with type 2 diabetes. Both African American (n = 511) and Caucasian (n = 235) patients with type 2 diabetes were recruited at six sites located in the metropolitan Detroit area. Scale reliability was calculated by Cronbach's coefficient alpha. The scale reliabilities ranged from .70 to .97 for African Americans. These reliabilities were similar to those of Caucasians, whose scale reliabilities ranged from .68 to .96. The Feldt test was used to determine differences between the reliabilities of the two patient populations. No significant differences were found. The DCP is a reliable survey instrument for African American and Caucasian patients with type 2 diabetes.

Black or African American↗