PubMed HealthSearch

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Reliability of seven measures of social intelligence in a sample of adolescents with mental retardation.

This study evaluates the reliability of seven measures, selected to assess the social-cognitive variables hypothesized by Greenspan to define social intelligence. Responses from 75, 30 and 20 adolescents with mental retardation were used to assess each test's internal, interrater, and test-retest reliabilities, respectively. Interrater reliability coefficients were high to very high (.76 to .98), internal reliabilities were moderate to very high (.66 to .90), and test-retest reliabilities were moderate to high (.54 to .74). Internal and test-retest reliability coefficients compared favourably with those reported for the subtests of the Revised Wechsler Intelligence Scale for Children.

Activities of Daily Living

Assessing the utility of reliability indices for automated visual fields. Testing ocular hypertensives.

Monocular (right eye) visual fields were recorded with the Humphrey Visual Field Analyzer (30-2 Program) at baseline as well as 6 and 12 months later in 120 patients with established ocular hypertension. Indices of field reliability (fixation loss, less than 20%; false-positives and false-negatives, less than 33%) and field sensitivity (mean deviation [MD] and pattern standard deviation [PSD]) were examined. At baseline, 35% of patients exhibited low reliability (LR) fields, a figure which decreased to approximately 25% at 6 and 12 months, respectively. During this period, over 50% of patients produced at least one LR field, whereas 8.3% were unable to produce even one reliable field. Exhibition of a LR field appeared to be independent of patient age. Fixation errors, the major cause of LR fields, decreased by approximately 10% over the 12-month period; most patients had between 20 and 32% fixation errors. The incidence of significant defects identified by PSD was greater than that for MD; this was true for both reliable and LR fields. It is suggested that increasing the fixation loss criteria for assessing patient reliability to a 33% cutoff might substantially increase the percentage of fields graded reliable with minimal effect on the sensitivity or specificity of the test.

Adult

Interobserver reliability and perceptual ratings: more than meets the ear.

The purpose of this study was to examine the reliability of ratings of perceptual characteristics for 10 ataxic dysarthric subjects. The influence of the occurrence of "deviant" speech parameters on the calculation of reliability coefficients was also explored. Results indicated that overall interobserver agreement levels for minimally trained judges compared favorably to reliability coefficients reported in previous studies. Furthermore, levels of overall agreement were above levels of agreement expected on the basis of chance alone. In contrast to overall interobserver agreement, much lower levels of interobserver agreement were obtained when "occurrence reliability" coefficients were calculated for deviant dimensions alone. However, occurrence reliability coefficients surpassed the level of agreement expected on the basis of chance alone for all subjects. Based on the results of this investigation, recommendations are made for modifying standard practices for obtaining interobserver reliability for perceptual ratings of speech characteristics.

Adolescent

The reliability of passive smoking histories reported in a case-control study of lung cancer.

A test-retest design has been used to examine the reliability of passive smoking histories reported in personal interviews. A total of 117 control subjects initially interviewed in a lung cancer case-control study conducted in metropolitan Toronto, Canada, between 1983 and 1984 were reinterviewed on average six months later. Responses to initial screening questions used to detect a person's exposure to passive smoke were more reliable for residential than for occupational exposure. Respondents also more reliably reported residential exposure to spouse's passive smoke than to the passive smoke of others at home. Quantitative measures of exposure to passive smoke, i.e., number and duration of exposure, were even less reliably reported. Nonsmoking respondents gave the most reliable information. The low reliability of self-reported duration of exposure to passive smoke is consistent with the inability of several studies to detect a significant dose-response relation with lung cancer risk when measures of dose that depend solely on duration are used.

Data Collection

Reliability of the attraction method for measuring lumbar spine backward bending.

The distraction method is one method used to measure forward bending of the spine. Although this technique, which requires the use of a tape measure held over the spine and the location of anatomical landmarks, appears to be highly practical, previous studies have not examined its use for measuring backward bending. The purpose of our study was to determine the reliability of a similar technique, the attraction method, for measuring backward bending of the lumbar spine and to examine whether subjects with low back pain (LBP) could perform similar motion as subjects without LBP. Two groups composed of 100 subjects each, one with "significant" limiting low back pain (SLBP) and the other without "significant" limiting low back pain (NSLBP), were evaluated twice by a physical therapist to assess intrarater reliability. To assess interrater reliability, 11 subjects from the NSLBP Group were evaluated by a second therapist. For the total sample of 200 subjects, the intraclass correlation coefficient (ICC) for intrarater reliability was .95; for the SLBP Group, the ICC was .93; and for the NSLBP Group, the ICC was .90. For the sample of 11 NSLBP Group subjects examined for interrater reliability, the ICC was .94. Using a Kolmogorov-Smirnov test, we found the distribution for backward bending of the two groups to be significantly different. The attraction method, thus, appears to be a reliable method for measuring backward bending of the lumbar spine.

Adolescent

Intrarater reliability of manual muscle testing and hand-held dynametric muscle testing.

Physical therapists require an accurate, reliable method for measuring muscle strength. They often use manual muscle testing or hand-held dynametric muscle testing (DMT), but few studies document the reliability of MMT or compare the reliability of the two types of testing. We designed this study to determine the intrarater reliability of MMT and DMT. A physical therapist performed manual and dynametric strength tests of the same five muscle groups on 11 patients and then repeated the tests two days later. The correlation coefficients were high and significantly different from zero for four muscle groups tested dynametrically and for two muscle groups tested manually. The test-retest reliability coefficients for two muscle groups tested manually could not be calculated because the values between subjects were identical. We concluded that both MMT and DMT are reliable testing methods, given the conditions described in this study. Both testing methods have specific applications and limitations, which we discuss.

Adult

Reliability of the Modified Motor Assessment Scale and the Barthel Index.

Many physical therapists use descriptive and functional assessments of motor recovery for patients with stroke. The purpose of this study was to establish the reliability of two such assessments. The Modified Motor Assessment Scale (MMAS) assesses motor recovery; the Barthel Index assesses functional independence. Interrater and intrarater reliability were determined for the total scores and individual item ratings using videotaped MMAS and Barthel Index assessments of seven patients with stroke. Therapists viewed and rated the videotaped assessments on two occasions separated by one month. The intrarater reliability results were higher than the interrater reliability results for total scores, and both results were acceptable statistically. Interrater and intrarater reliability of the individual item ratings were also determined. The MMAS and Barthel Index are reliable assessments of motor recovery and function for patients with stroke. Physical therapists are encouraged to use the two scales to document changes in the motor recovery and functional independence of patients with stroke.

Activities of Daily Living

Reliability of clinical measurements of lumbar lordosis taken with a flexible rule.

The purpose of this study was to examine the intratester and intertester reliability of lumbar lordosis measurements taken with a flexible rule. Two physical therapists (Tester 1 and Tester 2) took measurements on 40 subjects without low back pain (LBP) and on 40 subjects with LBP. Intraclass correlation coefficients (ICCs) were used to determine the degree of agreement between repeated measurements taken by the same therapist and between measurements taken by the two therapists. The ICC values for intratester reliability of Tester 1 were .84 for subjects without LBP and .94 for subjects with LBP. The ICC values of Tester 2 were .73 for subjects without LBP and .83 for subjects with LBP. Intertester reliability generally was poor, with ICC values of .41 for subjects without LBP and .50 for subjects with LBP. The results suggest that measurements of lumbar lordosis with a flexible rule may be reliable if taken by the same physical therapist. The degree of reliability, however, may vary from therapist to therapist. The intertester reliability of these measurements appears to be poor, but these conclusions must be interpreted carefully because of the limited number of therapists participating in this study.

Adult

Intrarater reliability of manual muscle test (Medical Research Council scale) grades in Duchenne's muscular dystrophy.

The purpose of this study was to document the intrarater reliability of manual muscle test (MMT) grades in assessing muscle strength in patients with Duchenne's muscular dystrophy (DMD). Subjects were 102 boys, aged 5 to 15 years, who were participating in a double-blind, multicenter trial to document the effects of prednisone on muscle strength in patients with DMD. Four physical therapists participated in the study. Two identical (duplicate) evaluations were performed within 5 days of each other by the same examiner initially and after 6 and 12 months of treatment. A total of 18 muscle groups were tested on each patient, 16 of them bilaterally, using a modification of the Medical Research Council scale. Reliability of muscle strength grades obtained for individual muscle groups and of individual muscle strength grades was analyzed using Cohen's weighted Kappa. The reliability of grades for individual muscle groups ranged from .65 to .93, with the proximal muscles having the higher reliability values. The reliability of individual muscle strength grades ranged from .80 to .99, with those in the gravity-eliminated range scoring the highest. We conclude the MMT grades are reliable for assessing muscle strength in boys with DMD when consecutive evaluations are performed by the same physical therapist.

Adolescent

Reliability of lumbar isometric torque in patients with chronic low back pain.

In this study, the test-retest reliability of lumbar isometric strength testing in patients with chronic low back pain (CLBP) was assessed. Isometric torque measurements were obtained from 89 patients with CLBP at seven different angles of lumbar flexion. Because previous studies have demonstrated significant strength differences between male and female subjects, separate data analyses were performed for each gender. Results indicated moderate to high reliability for patients with CLBP when tested at individually determined angles of flexion within their idiosyncratic range of motion (ROM) (female subjects: r = .59-.96, P less than .05, SEE = 12.0-24.2 N.m; male subjects: r = .71-.93, P less than .05, SEE = 25.1-62.1 N.m). For comparison with previously published data on asymptomatic controls, an additional set of analyses was conducted for subjects with full lumbar ROM. Similar reliability was demonstrated for this subsample (female subjects: r = .57-.93, P less than .05, SEE = 12.4-27.9 N.m; male subjects: r = .63-.93, P less than .05, SEE = 34.2-44.2 N.m). The authors concluded that isometric lumbar extension torque could be reliably measured in patients with CLBP at multiple positions within the full ROM, although reliability decreased at the most extended positions. The demonstrated reliability will allow researchers to assess treatment effects and group differences without undue concern for artifact attributable to measurement error.

Back Pain

On the methods and theory of reliability.

This paper reviews the most frequently used and misused reliability measures appearing in the mental health literature. We illustrate the various types of data sets on which reliability is assessed (i.e., two raters, more than two raters, and varying numbers of raters with dichotomous, polychotomous, and quantitative data). Reliability statistics appropriate for each data format are presented, and their pros and cons illustrated. Inadequancies of some methods are highlighted. The meaning of different levels of reliability obtained with various statistics is discussed. This critique is intended for the reading professional and the investigator who has an occasional need for reliability assessment. Statistical expertise is not required and theoretical material is referenced for the interested reader. Necessary formulas for computations are presented in the appendices. A summary table of some suitable reliability measures is presented.

Humans

The diagnosis of hypersensitivity to ingested foods. Reliability of skin prick testing and the radioallergosorbent test with different materials.

The diagnostic reliability in food allergy of skin prick tests (SPT) and the radio-allergosorbent test (RAST) was investigated in paediatric patients with respiratory and skin allergies. SPT and RAST were found to be reliable for the diagnosis of allergy to codfish, peas, nuts, peanuts and egg white. Positive SPT and RAST to cereals were common, but were most often without clinical significance or were correlated with respiratory allergy to the inhalation of flour dust. SPT and RAST were only partly reliable with regard to allergy to cow's milk, and were mostly reliable when used together and showing corresponding results. Experimental allergosorbents for RAST with soy beans and white beans were not reliable. The study shows the need to improve the diagnostic materials and to establish the diagnostic reliability of the material and tests used for each food item in question.

Adolescent

The validity of reliability assessments.

This paper focuses on reliability and evaluation of health education programs in school settings. Reliability is a concept that guides researchers in selecting or developing instruments, and is used as a standard, with validity and acceptability, for judging the credibility of research findings and inferences. Reliability is defined within the context of research design, and methods for estimating the reliability of cognitive measures are reviewed. Using data gathered in a school health education curriculum evaluation as an example, possible errors in hypotheses testing that may occur when estimating internal consistency of cognitive test scores obtained in quasi-experimental designs are examined. The appropriateness of internal consistency as a measure of reliability of cognitive measures is discussed and suggestions for reliability assessment and related issues such as power analysis are presented.

Attitude to Health

Instrument validity and reliability in three health education journals, 1980-1987.

Investigators examined how often validity and reliability measures were reported for research articles in three health education journals: Health Education, Health Education Quarterly, and the Journal of School Health. Articles published from 1980 to 1987 were considered in the analysis. Of the 611 articles published by Health Education during the period used for analysis, 128 (21%) met the criteria of a research article. Reliability was reported for 22 (17%) articles, and validity was reported for 78 (61%) articles. Health Education Quarterly published 212 articles; 74 (35%) were research articles. Reliability was reported for 16 (21%) articles and validity was reported for 40 (54%) articles. The Journal of School Health published 778 articles, of which 243 (31%) were research articles. Reliability was reported for 62 (25%), and validity was reported for 164 (67%) of the research articles. A chi-square test found a significant difference among the number of research articles published by the journals. Chi-square tests also found significant differences among the journals in the proportion of research articles that reported reliability information and the proportion that reported validity. A significant trend was noted for Health Education Quarterly and the Journal of School Health; the proportion of research articles that reported validity and reliability increased over time for both publications.

Health Education

Reliability of seizure diaries in adult epileptic patients.

Daily diaries are used widely in neurologic research and clinical practice to assess alterations in seizure frequency among patients with epilepsy. However, no formal tests of the reliability of this data collection method have been performed. We investigated the reliability of seizure recall in adult patients participating in a longitudinal study of stress, mood and seizure frequency. Patients maintained daily diaries for 10-36 weeks. The reliability study entailed completion of a single additional diary on the evening of a randomly selected day with reference to the preceding day. This design produced two diaries, completed 1 day apart, for the same 24-hour period. Measuring reliability with the Pearson correlation coefficient, overall reliability of seizure recall was 0.95 and was not markedly influenced by the subjects' sociodemographic characteristics, neurological or psychological status. In sum, the assumption in the literature that the daily diary is a reliable method for securing data on seizure counts appears warranted.

Adolescent

Variability and reliability of joint measurements.

The purpose of this study was to determine the variability and reliability of joint measurements as carried out by three physician observers. The intratester variation and reliability of nine different joint measurements was determined in eight healthy subjects. The measurements were taken in eight sessions by each tester. In this population also the intertester variation and reliability was determined by the three observers. This was also done in a population of middle-aged athletes over a period of 2.5 years. The results indicate that it is difficult to show either an improvement or worsening of a joint motion of less than 5 degrees to 10 degrees for most joints measured by the same tester. The intertester variation is not consistent over a longer period of time, so differences between observers during long-term studies cannot be corrected on the basis of a single study at a single point in time. The reliability of all nine joint measurements is not very high, but is probably sufficient if the results are used to compare groups within a single population and for large studies with experienced observers. Because the reliability strongly depends on the interindividual variation, it is preferable to determine the reliability for each study population.

Adult

Diagnostic interviewing with children: the use and reliability of the diagnostic coding form.

There have been few attempts to standardize assessment methods in Child Psychiatry. This paper describes a semi-structured approach to diagnostic interviewing of the child. Thirty-four children six to 13 years of age, and their parents, were interviewed two weeks apart by two different psychiatrists. A diagnostic coding form consisting of 29 clinical symptom items, eight summary items, and nine positive health ratings was used. Three diagnostic items were also included: "severity of clinical condition," "probability of disorder," and "adjustment status." Twelve of the Time 2 interviews with the child and parent were videotaped and rated by three different psychiatrists. Results indicated that summary items had higher reliability than individual symptom items and the three diagnostic items had the highest reliability, suggesting reliability is better for broad classes of behaviour. Interrater reliability was higher for the face-to-face rating than videotaped ratings. This suggests first that face-to-face interviews are reasonably stable over a two week period and second, since videotaped ratings had lowest reliability on items that depended on inferences about the child's feedlings, an important source of variance in assessment may be the clinician's ability to empathize with the child and draw inferences about internal feeling-states. It was concluded that this interview schedule can be a part of routine clinical practice. It ensures a reasonably standard, yet flexible and reliable approach to diagnostic interviewing.

Adaptation, Psychological

The reliability of alcohol abusers' self-reports of drinking and life events that occurred in the distant past.

This study investigated the test-retest reliability of 69 alcohol abusers' current reports about their past (approximately 8 years prior to interview) drinking behavior and life events. Drinking behavior was assessed by the Lifetime Drinking History (LDH) questionnaire and life events were assessed using the Recent Life Changes Questionnaire (RLCQ). Reliability coefficients for LDH variables were generally moderate to high (r = .52 to .81). Using empirical criteria, the diagnostic power of the two LDH interviews to classify correctly subjects as either having had or not having had a drinking problem was quite high. The reliability coefficient for the RLCQ was r = .85 and 91.7% of the identified events were reported in both interviews. Similarly high test-retest reliabilities and individual event agreement rates were obtained for the six homogeneous subscales of the RLCQ. Subjects were also asked why they had given inconsistent answers to life events questions in the two interviews. Inconsistencies often resulted from errors in the temporal placement of events or from misunderstanding items, rather than from failure to recall an event; this suggests that some sources of error in recalling life events can be reduced. It is concluded that alcohol abusers' reports of drinking and life events occurring many years prior to the date of interview are generally reliable. This finding is consistent with previous studies showing high test-retest reliabilities for reports of recent drinking and related events.

Adult