PubMed HealthSearch

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

On the methods and theory of reliability.

This paper reviews the most frequently used and misused reliability measures appearing in the mental health literature. We illustrate the various types of data sets on which reliability is assessed (i.e., two raters, more than two raters, and varying numbers of raters with dichotomous, polychotomous, and quantitative data). Reliability statistics appropriate for each data format are presented, and their pros and cons illustrated. Inadequancies of some methods are highlighted. The meaning of different levels of reliability obtained with various statistics is discussed. This critique is intended for the reading professional and the investigator who has an occasional need for reliability assessment. Statistical expertise is not required and theoretical material is referenced for the interested reader. Necessary formulas for computations are presented in the appendices. A summary table of some suitable reliability measures is presented.

Humans

The diagnosis of hypersensitivity to ingested foods. Reliability of skin prick testing and the radioallergosorbent test with different materials.

The diagnostic reliability in food allergy of skin prick tests (SPT) and the radio-allergosorbent test (RAST) was investigated in paediatric patients with respiratory and skin allergies. SPT and RAST were found to be reliable for the diagnosis of allergy to codfish, peas, nuts, peanuts and egg white. Positive SPT and RAST to cereals were common, but were most often without clinical significance or were correlated with respiratory allergy to the inhalation of flour dust. SPT and RAST were only partly reliable with regard to allergy to cow's milk, and were mostly reliable when used together and showing corresponding results. Experimental allergosorbents for RAST with soy beans and white beans were not reliable. The study shows the need to improve the diagnostic materials and to establish the diagnostic reliability of the material and tests used for each food item in question.

Adolescent

The validity of reliability assessments.

This paper focuses on reliability and evaluation of health education programs in school settings. Reliability is a concept that guides researchers in selecting or developing instruments, and is used as a standard, with validity and acceptability, for judging the credibility of research findings and inferences. Reliability is defined within the context of research design, and methods for estimating the reliability of cognitive measures are reviewed. Using data gathered in a school health education curriculum evaluation as an example, possible errors in hypotheses testing that may occur when estimating internal consistency of cognitive test scores obtained in quasi-experimental designs are examined. The appropriateness of internal consistency as a measure of reliability of cognitive measures is discussed and suggestions for reliability assessment and related issues such as power analysis are presented.

Attitude to Health

Instrument validity and reliability in three health education journals, 1980-1987.

Investigators examined how often validity and reliability measures were reported for research articles in three health education journals: Health Education, Health Education Quarterly, and the Journal of School Health. Articles published from 1980 to 1987 were considered in the analysis. Of the 611 articles published by Health Education during the period used for analysis, 128 (21%) met the criteria of a research article. Reliability was reported for 22 (17%) articles, and validity was reported for 78 (61%) articles. Health Education Quarterly published 212 articles; 74 (35%) were research articles. Reliability was reported for 16 (21%) articles and validity was reported for 40 (54%) articles. The Journal of School Health published 778 articles, of which 243 (31%) were research articles. Reliability was reported for 62 (25%), and validity was reported for 164 (67%) of the research articles. A chi-square test found a significant difference among the number of research articles published by the journals. Chi-square tests also found significant differences among the journals in the proportion of research articles that reported reliability information and the proportion that reported validity. A significant trend was noted for Health Education Quarterly and the Journal of School Health; the proportion of research articles that reported validity and reliability increased over time for both publications.

Health Education

Reliability of seizure diaries in adult epileptic patients.

Daily diaries are used widely in neurologic research and clinical practice to assess alterations in seizure frequency among patients with epilepsy. However, no formal tests of the reliability of this data collection method have been performed. We investigated the reliability of seizure recall in adult patients participating in a longitudinal study of stress, mood and seizure frequency. Patients maintained daily diaries for 10-36 weeks. The reliability study entailed completion of a single additional diary on the evening of a randomly selected day with reference to the preceding day. This design produced two diaries, completed 1 day apart, for the same 24-hour period. Measuring reliability with the Pearson correlation coefficient, overall reliability of seizure recall was 0.95 and was not markedly influenced by the subjects' sociodemographic characteristics, neurological or psychological status. In sum, the assumption in the literature that the daily diary is a reliable method for securing data on seizure counts appears warranted.

Adolescent

Variability and reliability of joint measurements.

The purpose of this study was to determine the variability and reliability of joint measurements as carried out by three physician observers. The intratester variation and reliability of nine different joint measurements was determined in eight healthy subjects. The measurements were taken in eight sessions by each tester. In this population also the intertester variation and reliability was determined by the three observers. This was also done in a population of middle-aged athletes over a period of 2.5 years. The results indicate that it is difficult to show either an improvement or worsening of a joint motion of less than 5 degrees to 10 degrees for most joints measured by the same tester. The intertester variation is not consistent over a longer period of time, so differences between observers during long-term studies cannot be corrected on the basis of a single study at a single point in time. The reliability of all nine joint measurements is not very high, but is probably sufficient if the results are used to compare groups within a single population and for large studies with experienced observers. Because the reliability strongly depends on the interindividual variation, it is preferable to determine the reliability for each study population.

Adult

Diagnostic interviewing with children: the use and reliability of the diagnostic coding form.

There have been few attempts to standardize assessment methods in Child Psychiatry. This paper describes a semi-structured approach to diagnostic interviewing of the child. Thirty-four children six to 13 years of age, and their parents, were interviewed two weeks apart by two different psychiatrists. A diagnostic coding form consisting of 29 clinical symptom items, eight summary items, and nine positive health ratings was used. Three diagnostic items were also included: "severity of clinical condition," "probability of disorder," and "adjustment status." Twelve of the Time 2 interviews with the child and parent were videotaped and rated by three different psychiatrists. Results indicated that summary items had higher reliability than individual symptom items and the three diagnostic items had the highest reliability, suggesting reliability is better for broad classes of behaviour. Interrater reliability was higher for the face-to-face rating than videotaped ratings. This suggests first that face-to-face interviews are reasonably stable over a two week period and second, since videotaped ratings had lowest reliability on items that depended on inferences about the child's feedlings, an important source of variance in assessment may be the clinician's ability to empathize with the child and draw inferences about internal feeling-states. It was concluded that this interview schedule can be a part of routine clinical practice. It ensures a reasonably standard, yet flexible and reliable approach to diagnostic interviewing.

Adaptation, Psychological

The reliability of alcohol abusers' self-reports of drinking and life events that occurred in the distant past.

This study investigated the test-retest reliability of 69 alcohol abusers' current reports about their past (approximately 8 years prior to interview) drinking behavior and life events. Drinking behavior was assessed by the Lifetime Drinking History (LDH) questionnaire and life events were assessed using the Recent Life Changes Questionnaire (RLCQ). Reliability coefficients for LDH variables were generally moderate to high (r = .52 to .81). Using empirical criteria, the diagnostic power of the two LDH interviews to classify correctly subjects as either having had or not having had a drinking problem was quite high. The reliability coefficient for the RLCQ was r = .85 and 91.7% of the identified events were reported in both interviews. Similarly high test-retest reliabilities and individual event agreement rates were obtained for the six homogeneous subscales of the RLCQ. Subjects were also asked why they had given inconsistent answers to life events questions in the two interviews. Inconsistencies often resulted from errors in the temporal placement of events or from misunderstanding items, rather than from failure to recall an event; this suggests that some sources of error in recalling life events can be reduced. It is concluded that alcohol abusers' reports of drinking and life events occurring many years prior to the date of interview are generally reliable. This finding is consistent with previous studies showing high test-retest reliabilities for reports of recent drinking and related events.

Adult

Diabetologists' judgments of diabetic control: reliability and mathematical simulation.

In study 1, laboratory and supervised blood or urine test data from actual cases were used to develop patient profiles. Seven diabetologists from the same institution rated the diabetic control of 125 profiles on a four-point scale (1 = poor, 2 = fair, 3 = good, 4 = excellent). Six of the 7 diabetologists demonstrated adequate intra- and interrater reliability. Study 2 assessed the reliability of judgments of diabetic control made by diabetologists working in two different settings. There were 9 raters from institution 1 and 8 from institution 2. The impact of the amount and type of information on judgment reliability was evaluated by developing two types of profiles. The test form contained only laboratory and supervised blood or urine test data similar to that utilized in study 1. The history form contained this information as well as other descriptive data typically available to diabetologists. The 17 diabetologists rated 125 anonymous profiles on each of two separate occasions approximately 1 wk apart. On one occasion they rated profiles presented on the test form. On the other occasion they rated profiles presented on the history form. As in study 1, the diabetologist raters demonstrated adequate intra- and interrater reliability. Intrarater reliability was somewhat better when rating test form profiles compared with history form profiles. Reliability was not higher within than between institutions. An analysis of the relative contribution of different diabetes control indices to the diabetologists' judgments indicated that HbA1 influenced raters' judgments at both institutions more than any other single variable.(ABSTRACT TRUNCATED AT 250 WORDS)

Adolescent

Comparative clinical reliability of fasting plasma glucose and glycosylated hemoglobin in non-insulin-dependent diabetes mellitus.

Because accurate determination of glycosylated hemoglobin (GHb) is difficult and relatively expensive in comparison with the modest cost and ready availability for tests of fasting plasma glucose (FPG), we examined the reliability of repeated measurements of FPG and GHb in typical diabetic outpatients taken in the usual clinical setting. We determined FPG and GHb concurrently on three separate occasions spanning 4 wk in 41 patients with non-insulin-dependent diabetes mellitus (NIDDM) and, for contrast, 5 with insulin-dependent diabetes mellitus (IDDM). Most of the NIDDM subjects were obese, with initial FPG levels ranging from 93 to 355 mg/dl. The reliability of each test was estimated by calculating two measures: the intraclass correlation coefficient (rho I) and the coefficient of variation (CV) for the repeated test values. For NIDDM patients treated with diet or oral hypoglycemic agents (OHA), rho I for FPG, log(FPG), and GHb were very similar. For insulin-treated NIDDM patients, rho I for FPG was somewhat lower than the coefficient in other treatment groups, and the reliability of FPG by this measure did not match the reliability of GHb within the limits of statistical significance. By analyzing the CV of test values repeated within subject, the reliability of FPG did not differ from GHb in any of the NIDDM treatment groups. Although patients were recruited sequentially to minimize sample selection bias, caution must be exercised in the interpretation of the statistical analyses of reliability with either rho I or CV due to limitations imposed by small sample size.(ABSTRACT TRUNCATED AT 250 WORDS)

Analysis of Variance

Percent of agreement among raters and rater reliability of the copying subtest of the Stanford-Binet Intelligence Scale: Fourth Edition.

The purpose of this study was to investigate the interrater reliability of the visual-motor portion of the Copying subtest of the Stanford-Binet Intelligence Scale: Fourth Edition. Eight raters independently scored 11 protocols completed by children aged 5 through 10 years, using the scoring criteria and guidelines in the manual. The raters marked each of 10 items pass or fail and computed a total raw score for each protocol. Interrater reliability coefficients were obtained for each child's protocol, and the Kappa coefficient was computed for each item. Significant raters' reliability coefficients ranged from .82 to .91, which were low in comparison to test-retest reliability and Kuder-Richardson-20 coefficients for this and other subtests of the Stanford-Binet in the technical manual. Percent agreement among 8 raters also indicated weak reliability. Although the obtained results suggested some interrater reliability coefficients within acceptable levels, questions were raised about the scoring criteria for individual items. Caution is warranted in the use of cognitive measures which include subjective judgement of the examiner in applying scoring criteria.

Child

Number of stimuli as a reliability parameter in perimetry.

Catch trials test patient performance during automated, static perimetry, but their adequacy to estimate reliability is uncertain even though up to 10% of the test time is reserved for catch trials. The 308 visual fields (program G1, all 3 phases, Octopus 201) of 308 eyes of 308 glaucoma, suspected glaucoma, and normal subjects were studied. The 108 visual fields (mean sensitivity > 10 dB; corrected loss variance < 50 dB2) without false responses to catch trials were considered reliable. A multiple linear regression analysis of these 108 fields was performed and revealed the following result (r2 = 0.751): Number of stimuli = 480 + (40.short-term fluctuation) + (8.8.the square root of the index corrected loss variance) - (2.2.mean sensitivity). This equation was used to estimate the number of stimuli required of a reliable subject to complete an examination. Excess stimuli would thus be a sign of reduced reliability. The difference between the estimated and the actual number of stimuli was called the 'stimulus discrepancy'. In 169 fields with false-positive and 58 fields with false-negative responses, the false-positive and false-negative responses correlated with the 'stimulus discrepancy' (r = 0.19, P = 0.014; r = 0.29, P < 0.026, respectively). The number of stimuli depends not only on reliability but also on the software and hardware of the perimeter. 'Stimulus discrepancy' may be an additional useful perimetric reliability parameter which does not require extra testing time.

Adult

Reliability of individual differences for H-reflex recordings.

Complementary data from two closely-related experiments were analyzed to investigate the reliability of individual differences of H-reflex amplitude. Some concomitant analysis of M-reflex was also included. The purpose was to do this in a manner that was more direct and more detailed than found hitherto in the literature. Fundamental techniques espoused particularly by F.M. Henry (1969) formed the basis of this approach. The aims included a clearer quantification of reliability with a distinction between inter- and intra-individual sources of variation, an analysis of the effects different numbers of trials had on reliability, and an investigation of the progression of reliability over a 20-trial sequence of observations. The control conditions of both experiments were designed to be essentially identical with respect to subject position, stimulus and recording configurations. Experiment 1 had 20 subjects, four control conditions of 10-trials in each while experiment 2 had 18 subjects, five control conditions and 20 trials. The posterior tibial nerve was percutaneously stimulated every 10 sec. and surface electromyographic recordings were made from the soleus muscle. H-reflex and M-reflex amplitudes were measured on every trial. It was found that the reliability of individual differences of both H-reflex and M-reflex was extremely robust with the majority of coefficients being above .950. In addition, the reliabilities remained high when as few as four trials were examined. The pattern of individual differences over the longer series of trials confirmed the stability of interindividual differences and showed that subjects became more consistent in their responses as trials progressed.(ABSTRACT TRUNCATED AT 250 WORDS)

H-Reflex

Reliability and variability of heart rate monitoring in 3-, 4-, or 5-yr-old children.

We describe the daily heart rate patterns and the between day and within day reliabilities of several heart rate variables measured in 159 Anglo-, African-, and Mexican-American children aged 3-5 yr. Heart rates were measured over 12 waking hours with a Quantum XL Telemetry heart rate monitor. There were no significant ethnic, gender, day of week, or season of the year differences in either mean resting heart rate, mean daily heart rate, mean longest duration of the heart rate sustained above 120 bpm for the day, nor percent of minutes of daily heart rate above 120 bpm. The reliabilities for these variables for 2 d of observation separated by 3-6 months ranged from 0.65 to 0.66. At this level of reliability, just over 4 d of recording are necessary to achieve a reliability of 0.80. All within-day across-hour reliabilities were greater than 0.80. However, for mean hourly heart rate and the longest duration of heart rate sustained above 120 bpm each hour, a principal components analysis revealed three distinct time components during the day. This suggests that monitoring heart rate during limited portions of the day will provide a biased estimate of overall heart rate. For the morning component, there were significant ethnic and gender differences in the children's heart rates and younger children had longer durations of heart rate sustained above 120 bpm than older children. Although daily heart rate monitoring is not a perfect indicator of children's physical activity, these data suggest that it may be a reliable measure among younger children from different ethnic and gender groups.

Black People

Measuring severity of illness: a comparison of interrater reliability among severity methodologies.

Methods for measuring illness severity are receiving increasing attention from payers, purchasers, and others interested in the equity and financial incentives of prospective payment systems, as well as from those concerned with the use of mortality rates and other outcomes to measure quality of care. Several methodologies have been proposed for measuring the severity of illness of patients admitted to hospitals. When choosing among the available measures, one characteristic of interest is reliability. In this paper, we present a comparative evaluation of interrater reliability among four severity measures--APACHE II, MedisGroups, Patient Management Categories (PMCs), and Disease Staging Q-Scale--as well as for the Diagnosis Related Groups (DRG) classification system. The results show APACHE II, MedisGroups, and DRGs to be highly reliable, with Inter-Rater Reliability Coefficient (RI) values greater than .8. PMCs and Disease Staging Q-Scale were able to achieve fair-to-good levels of reliability. Results are consistent regardless of which reliability statistics are used.

Evaluation Studies as Topic

Reliability of self-reported sexual behavior risk factors for HIV infection in homosexual men.

This study was undertaken to determine the reliability of self-reported sexual behavior using the test and retest technique when used with self-reported sexual behavior. The subjects were 116 asymptomatic homosexual men who participated in another study (an examination of behavioral and demographic determinants of HIV antibody status). The subjects were asked to complete two questionnaires. The first contained demographic and sexual behavior questions. The second, administered an average of 6 weeks later, used a subset of the questions in the first questionnaire. The reliability of the test-retest procedure was measured by the Kappa statistic, which assesses the proportion of agreement between two data items, accounting for the amount of agreement expected by chance. The highest degree of reliability as measured by Kappa was found with demographic information, smoking history, and sexual orientation. Self-reported sexual behaviors for the previous 6 months generally had the next highest degree of reliability as measured by Kappa. Questions examining change over the previous 5 years had the lowest reliability. Behavior changes during the time between questionnaires, subjectivity of the answer categories, and social desirability of the answers are three factors that may result in a lack of reliability in this self-reported sexual behavior questionnaire. This raises methodological concerns about the measurement of behavioral risk factors for AIDS and the ability to assess meaningfully subjective reports of behavioral change.

Acquired Immunodeficiency Syndrome

Inter- and intra-examiner reliability of the upper cervical X-ray marking system: a second look.

To determine the degree of reliability (stability over time) for six Pettibon practitioners, the scores resulting from the reading and re-reading of 30 X rays were analyzed using bivariate scattergrams, Pearson Product-moment correlation coefficient estimates and correlated samples t tests. To examine reliability (equivalence over experts) across the practitioners, a repeated measures analysis of variance approach was used. Liberal and conservative reliability coefficients for the upper angle and lower angle were computed. Examination of the data suggest that the reliability (stability over time) for the practitioners is very good. The data on reliability (equivalence over experts) across the practitioners also suggests reliability is very good.

Cervical Vertebrae

Reliability of provocative tests of motion sickness susceptibility.

Accurate prediction of space motion sickness is dependent, in part, upon the reliability of terrestrial-based motion sickness susceptibility tests. In the present study, test-retest reliability values were derived from motion sickness susceptibility scores obtained from two successive exposures to each of three tests: 1) Coriolis Sickness Sensitivity Index (CSSI); 2) Staircase Velocity Movement Test (SVMT); and 3) Parabolic Flight Static Chair Test (PSCT). The reliability of the three tests ranged from 0.70 to 0.88. Normalizing values from predictors with skewed distributions improved the reliability. The apparent inconsistency between our finding of high reliability of predictive tests, and previous reports of low correlations between ground-based predictors and space motion sickness may be due to unreliability in assessment of the sickness criterion. Issues of reliability and validity of predictor and criterion measures, and their implications for future development of ground-based predictive tests are discussed in some detail.

Acceleration