PubMed HealthSearch

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

An Assessment of Reliability Estimation Methods for Binomial Health Care Quality Measures.

We evaluated the performance of commonly used methods for estimating the reliability of binomial health care quality measures using simulated datasets spanning a range of performance score means and variances, numbers of entities, and patient sample sizes. For each simulation, reliability was estimated for all selected methods and compared with the known true reliability derived from the simulation parameters, with methods assessed on their accuracy and precision. Logistic regression with reliability estimated on the outcome scale demonstrated the highest accuracy and precision among all methods evaluated. The widely used Adams beta-binomial method performed poorly, although a modification recommended by Nieser and Harris substantially improved its performance. These approaches are applicable only to binomial measures. Among methods that can be applied to both binomial and continuous measures, permutation resampling of the Spearman rank correlation coefficient was the most accurate and precise, outperforming other commonly used approaches. Overall, for binomial quality measures, logistic regression on the outcome scale is the preferred method for reliability estimation, followed closely by the modified beta-binomial approach, while for non-binomial measures, permutation-based Spearman rank correlation appears to be the most suitable method.

Reproducibility of Results

The interrater reliability of the Psychotic Inpatient Profile.

The interrater reliabilities of the 12 scales of the Psychotic Inpatient Profile (PIP) were assessed in three independent samples. These reliabilities were found to be consistently low for the three samples. Several possible sources of low reliability were investigated, and evidence to support the hypothesis that consistent rater biases might have contributed to this low reliability was found. The authors make several suggestions of ways to improve PIP's reliability.

Adjustment Disorders

Reliability of the WAIS-R with psychiatric inpatients.

WAIS-R subtest and composite scale reliabilities, standard errors of measurement, and standard errors of estimate were determined for a sample of psychiatric inpatients (N = 100). For Digit Span and Digit Symbol, test-retest stability coefficients were obtained; split-half reliability coefficients were calculated for all other subtests. With the exception of Object Assembly (rxx = .38), all subtest and composite scale reliability coefficients were large and acceptable. Based on the standard error of measure, the most reliable WAIS-R subtests were Digit Symbol (.77), Information (1.04), and Picture Completion (1.07). Reliability coefficients for the psychiatric inpatient sample were, in general, comparable to those values reported for the standardization group (Wechsler, 1981). Significant differences were obtained only on the Object Assembly and Vocabulary subtests.

Adjustment Disorders

The inter-rater reliability and internal consistency of a clinical evaluation exercise.

OBJECTIVE: To assess the internal consistency and inter-rater reliability of a clinical evaluation exercise (CEX) format that was designed to be easily utilized, but sufficiently detailed, to achieve uniform recording of the observed examination. DESIGN: A comparison of 128 CEXs conducted for 32 internal medicine interns by full-time faculty. This paper reports alpha coefficients as measures of internal consistency and several measures of inter-rater reliability. SETTING: A university internal medicine program. Observations were conducted at the end of the internship year. PARTICIPANTS: Participants were 32 interns and observers were 12 full-time faculty in the department of medicine. The entire intern group was chosen in order to optimize the spectrum of abilities represented. Patients used for the study were recruited by the chief resident from the inpatient medical service based on their ability and willingness to participate. INTERVENTION: Each intern was observed twice and there were two examiners during each CEX. The examiners were given a standardized preparation and used a format developed over five years of previous pilot studies. MEASUREMENTS AND MAIN RESULTS: The format appeared to have excellent internal consistency; alpha coefficients ranged from 0.79 to 0.99. However, multiple methods of determining inter-rater reliability yielded similar results; intraclass correlations ranged from 0.23 to 0.50 and generalizability coefficients from a low of 0.00 for the overall rating of the CEX to a high of 0.61 for the physical examination section. Transforming scores to eliminate rater effects and dichotomizing results into pass-fail did not appear to enhance the reliability results. CONCLUSIONS: Although the CEX is a valuable didactic tool, its psychometric properties preclude reliable assessment of clinical skills as a one-time observation.

Clinical Competence

Comparative reliability of categorical and analogue rating scales in the assessment of psychiatric symptomatology.

The reliability of 26 items from the ninth edition of the Present State Examination (PSE) was assessed using both the conventional categorical scales and separately constructed analogue scales. Reliability was also calculated when the analogue responses were rescaled down to 2, 3 and 4 categories. The levels of inter-rater agreement obtained were comparable to those achieved in previous studies of PSE reliability, although as expected the levels of agreement on audiotapes were greater than those for independent interviews performed on the same day. These levels were not significantly affected by any of the changes in scale format, but there were apparent differences in reliability depending on the statistics used. In selecting or constructing a psychiatric rating scale, the question of reliability should not influence the choice of a categorical or continuous scale, or the number of scored points in the scale.

Psychiatric Status Rating Scales

Assessing the reliability of clinical scales when the data have both nominal and ordinal features: proposed guidelines for neuropsychological assessments.

The purpose of this article is to present, for the first time, a comprehensive methodology for assessing the reliability of a clinical scale that is frequently utilized in neuropsychological research and in biomedical studies, more generally. The dichotomous-ordinal scale is characterized by a single category of "absence" and two or more ordinalized categories of "presence" of a symptom trait, state, or behavior, and it also has special properties that need to be understood in order for its reliability to be appropriately assessed. Using the Brief Psychiatric Rating Scale (BPRS) as a clinical example, we cover the principles of expressing scale reliability in terms of a dichotomy ("absence" - "presence" of a given BPRS symptom); as a trichotomy ("none"; "mild to moderate" symptomatology; and "severe" symptomatology); and as the full 7-category dichotomous-ordinal scale: "none," "very mild," "mild," "moderate," "moderately severe," "severe," and "extremely severe." Criteria are presented that can be used to evaluate which of these three formats produces the most reliable results. Finally, we address, with a second sample, the important issue of replication, or whether the original reliability findings generalize to other independent populations.

Adult

Reliability of dual diagnosis. Substance dependence and psychiatric disorders.

The Structured Clinical Interview for DSM-III-R was used to examine the effects of the co-occurrence of psychiatric and substance dependence disorders on diagnostic reliability. The test-retest reliability over a 1-week period was studied in groups of: a) individuals with current substance abuse diagnoses (N = 97), b) individuals with past, but not current, drug histories (N = 146), and c) individuals without substance abuse diagnoses (N = 356; primarily psychiatric patients). A measurement of reliability (Kappa coefficients) was estimated for four general psychiatric categories (psychotic, mood, anxiety, and eating disorders), along with specific most-frequent diagnoses in each category (schizophrenia, major depression, panic disorders, and bulimia nervosa, respectively). Past use and non-drug-use groups were similar in their generally reliable reporting of current and past psychiatric disorders. However, current mood and psychotic disorders were less reliably diagnosed in the group with current substance use disorders.

Adult

The Sickness Impact Profile: reliability of a health status measure.

This report describes the results of research conducted on the reliability of the Sickness Impact Profile (SIP). The SIP is a questionnaire instrument designed to measure sickness-related behavioral dysfunction and is being developed for use as an outcome measure in the evaluation of health care. The test-retest reliability of the SIP in terms of several reliability measures was investigated using different interviewers, forms, administration procedures, and a variety of subjects who differed in terms of type and severity of dysfunction. The results provided evidence for the feasibility of collecting reliable data using the SIP under these various conditions. In addition, subject variability in relation to reliability is discussed.

Disability Evaluation

Reliability of goniometric measurements of hip motion in spastic cerebral palsy.

Reliability of goniometric measurements of ranges of motion in the right hip of four children with mild or moderate spastic diplegia was studied by comparing results when physical therapists used specific and non-specific measurement instructions. Measurements of hip extension, abduction and external rotation were made. The use of specific measurement instructions improved inter-rater reliability only in the case of external rotation. Inter-session reliability was not improved by the use of specific measurement instructions. It is concluded that goniometric measurements have such a low level of reliability that they can only be used to assist clinical judgment, and that they are not sufficiently reliable to be used in studied of cerebral palsy.

Adolescent

Reliability of caries data in three clinical trials.

The reliability of caries data obtained during three clinical trials is presented. Each of the three trials, conducted by different examiners, lasted 3 years and involved children aged 12 to 15 years. Reliability was quantified in terms of reliability coefficient and error variance, which allowed the effect of error upon the efficiency of each study to be measured. Reliability tended to be reduced when precaviation (initial) lesions were included in the DMFS count, but there was little difference between the reliability coefficients for each of four different surface-types. Error was more important in 3-year incremental rather than prevalence data, but nevertheless had a rather small influence upon sample size estimation or confidence limits of percent caries reductions.

Adolescent

Diagnostic reliability of the percutaneous ultrasonic Doppler technique for vertebral arterial occlusive diseases.

There is little data on the diagnostic reliability of the ultrasonic Doppler technique for vertebral arterial occlusive lesions. Percutaneous vertebral Doppler examination and the vertebral angiograms were compared to determine the diagnostic reliability of this technique in 64 vertebral arteries of 53 patients with cerebrovascular disease. The percutaneous vertebral Doppler findings were quantitatively analyzed using a sound spectrograph and were classified into three types: no flow signal type, poor flow type and normal flow type. In nine patients with the no flow type the angiograms revealed vertebral occlusion or a missing vertebral artery in six, giving a diagnostic reliability of 67%. In 17 patients with poor flow type the angiograms revealed vertebral occlusion or a missing vertebral artery in five, terminal narrowing of the artery in nine, and hypoplasia in two giving a diagnostic reliability of 94%. For all vertebral arteries examined with this technique, including normal ones, the diagnostic reliability was 92% (59/64). Percutaneous vertebral Doppler examination has clinical usefulness as a screening test for occlusive vertebral arterial diseases.

Adult

Further studies of a measure of adience-abience: reliability.

Previous studies have suggested adequate reliability for a measure of perceptual adience-abience based on the HABGT in view of significant evidence concerning its construct and predictive validity. The present study explored the test-retest reliability of this scale with 40 process schizophrenics over a two-week interval. Reliability was found to be adequate for both males and females (rho equals .84) for the total scale. The four components of the scale were similarly found to be reliable. No subject changed on retest with respect to adient or abient orientation. Interjudge reliability was very high (rho = .912).

Adult

Psychopathology scale of the Hutt adaptation of the Bender-Gestalt Test: reliability.

The test-retest reliability of the Hutt Adaptation of the Bender-Gestalt test was explored with a population of 40 process schizophrenics over a two-week interval. The total Psychopathology Scale Score was found to have high retest reliability for both male and female patients (rho = .87 for males and .83 for females). Moreover the three major components for the Scale were found to have high reliability, and fairly high reliabilities were obtained for patients scoring high as well as low on the Scale. Interjudge reliability was also found to be very high (rho = .895), confirming previous studies in this respect. On these grounds, the Scale offers promise both for clinical and research purposes.

Adult

Establishing reliability of the Community Oriented Program Environment Scale on chemically-dependent black females.

Research instruments are often used on samples without determining their reliability for that group of subjects. As a result, information is disseminated that really has no scientifically sound base. Thirty chemically-dependent Black females were chosen as subjects to determine reliability of the Community Oriented Program Environment Scale on this population. They were administered the instrument, and reliability was obtained using the Kuder-Richardson 20 internal consistency test. Findings showed low overall reliability and extremely low subscale reliability on this sample. It was concluded that this instrument would yield uninterpretable data in looking at attrition and treatment environment with this population.

Adult

Clinical measures of shoulder subluxation: their reliability.

The purpose of this report is to describe the reliability of three clinical measures used to evaluate changes in shoulder subluxation. The three methods include measuring the subacromial space in fingers breadth, using calipers, or a plexiglass jig. Thirty-six patients with shoulder subluxation who had experienced a cerebrovascular accident were the subjects. Four occupational therapists with experience in stroke rehabilitation were divided into two teams of two therapists and rated the subjects independently. Each rater repeated her assessments on nine subjects to test intrarater agreement. Inter-rater agreement was assessed both between members of the same team (27 subjects per rater pair) and members of different teams (18 subjects per rater pair). The measure of reliability was the intraclass correlation coefficient (ICC2, 1) as derived from two-way analysis of variance. The highest intra-rater reliability was displayed by the finger breadth method, the second highest by the caliper method and the lowest by the plexiglass jig. The coefficients in the former two cases were always above .8. Using the jig only one rater achieved this level. Agreement between the two members of the same team were above .75 for the fingers and caliper methods, but less than this for the jig. Between members of different teams however, only the finger breadth method attained reliabilities above .7, and the plexiglass jig, in particular, showed very low reliability. These results demonstrate the difficulty of achieving consistent clinical measurement for a condition like shoulder subluxation.(ABSTRACT TRUNCATED AT 250 WORDS)

Adult

Absolute and relative reliability of several response parameters used in vestibular assessment.

This experiment investigated the reliability of five indices used to express the magnitude of induced vestibular nystagmus. The investigation was undertaken because information concerning the reliability of the various response parameters when caloric testing is repeated on the same subject is limited. Electronystagmographic assessments of the nystagmus response to the Fitzgerald-Hallpike test were obtained from 16 subjects on three occasions with equal intervals of time between occasions. The data were statistically treated with an analysis of variance technique that allows the generation of correlation coefficients. The findings showed that speed of the slow phase, total beats, culmination frequency, and total amplitude reliably express the magnitude of induced nystagmus on repeated tests. Duration of the induced reaction was not as reliable as the other indices on repeated testing. In addition, difference scores were more repeatable than absolute scores and the caution that must be maintained when using correlation coefficient analyses to estimate reliability is highlighted.

Adolescent

Response reliability in a longitudinal survey in Thailand.

The two rounds of the National Longitudinal Study in Thailand provide a useful opportunity to explore response reliability in a large-scale social and demographic survey in a developing country. The results indicate that nonrandom reliability at the individual level ranged from quite high (for several straightforward, facual questions) to quite low (for most attitudinal questions). There was considerable distributional stability, however, even for many of the variables with low individual-level reliability. In terms of its response reliability, the Thai study compares reasonably well with several leading US fertility surveys. However, in both countries response reliability at the individual level for attitudinal questions is distressingly low. This clearly should be a matter of major concern for social scientists using survey results.

Demography

Levels of reliability in fertility survey data.

A number of factual and attitudinal questions asked in the 1973 Taiwan KAP-4 survey were repeated in a postenumeration survey one month later in order to assess the reliability of responses of the 286 women reinterviewed. The level of reliability is found to vary depending on the measures used and on whether the focus is aggregate data or individual responses. Analysis of consistency of responses shows that while overall reliability at both aggregate and individual levels is reasonably good, there is greater reliability for factual than for attitudinal data. Nevertheless, consistency of responses on factual questions varies considerably depending on the salience of the topic to the respondent. Estimates of reliability are shown to depend on the measure used and on the skewness of the distributions of the responses.

Contraception