PubMed HealthSearch

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Assessing the likelihood of reliable workplace behavior: further contributions to the validation of the Employee Reliability Inventory.

This paper summarizes a number of studies in which the validity of the Employee Reliability Inventory, a preemployment screening instrument designed to assess the likelihood of reliable and productive workplace behavior, was examined. Criterion-related studies compared the scores of a broadly diverse group of job applicants with those obtained from an array of criterion and comparison groups, for whom there was documented evidence of reliable or unreliable behavior. Criterion-related evidence indicates that the six scales are effective in differentiating a variety of criterion groups with unreliable behavior from a number of different job applicant comparison groups. Construct-related evidence for the validity of an emotional adjustment scale is reported as well. The issue of response distortion in preemployment inventories is discussed, and data are reported which indicate that scores on all six scales appear to be functionally free from the potentially confounding effects of response distortion. These results are consistent with the original validation and cross-validation findings, which supported the validity of the six scales when assessing the likelihood of reliable behavior in a population of job applicants.

Adult

Reliability of psychiatric diagnosis. II. The test/retest reliability of diagnostic classification.

In a study of interrater diagnostic reliability, 101 psychiatric inpatients were independently interviewed by physicians using a structured interview. Newly admitted patients were randomly selected and examined by one of three psychiatrists. A second psychiatrist reexamined the same patient about 24 hours later. Interviews from the two examinations were evaluated independently and diagnoses were made on the basis of objective criteria. The degree of diagnostic agreement for the two examinations were calculated using the kappa statistic. Agreement was found to be high as compared to other studies in the psychiatric literature, despite the fact that in most previous investigations diagnoses were not made independently. The results were also compared to studies of reliability of medical judgments. Possible reasons for the high interrater reliability are discussed and include the use of a structured interview and objective diagnostic criteria.

Attitude of Health Personnel

Reliability, Device Agreement and Validity of Load-Velocity Profiles: A Systematic Review with Meta-analysis.

BACKGROUND: For a valid one-repetition maximum (1RM) prediction via load-velocity (LV) relationships, high reliability and accuracy must be assumed. OBJECTIVE: Since individual study results indicate ambivalent prediction, this systematic review and meta-analysis was designed to provide a updated and comprehensive overview, extending knowledge about the validity and reliability of commercially available velocity sensors in Part I and the validity and reliability of velocity-based 1RM prediction models in Part II. METHODS: A systematic literature search was conducted in PubMed/MEDLINE, Web of Science, and Scopus. Validity and/or reliability studies or velocity-based 1RM prediction evaluations were included. Methodological quality was assessed using adapted COSMIN. The analysis was performed for intraclass correlation coefficient (ICC), Lin's concordance correlation coefficient (CCC), and Pearson's correlation coefficient (r). The review was preregistered in PROSPERO (CRD42025634595). RESULTS: Sixty-three studies were included for sensor validity and reliability and 38 for 1RM prediction models. Part I: Velocity sensors demonstrated good-to-excellent pooled validity and device agreement (ICC = 0.91-0.92 [0.83-0.97]; k = 55 and 439, respectively); intra- and inter-day reliability were classified as good to excellent with ICC = 0.90-0.91 [0.85-0.95] (k = 228 and 608, respectively), with sensor technology moderating the results. However, substantial heterogeneity and wide ranges of study-level estimates indicated considerable variability across moderators, linear position transducer (LPT) generally showing more consistent performance than inertial measurement units (IMU). Part II: Velocity-based 1RM prediction showed ICCs = 0.90 [0.83-0.94] (k = 124) and ICC = 0.91 [0.72-0.98] (k = 9); for reliability and validity, respectively. DISCUSSION: Commercial velocity sensors generally provide high relative validity and reliability. Results varied depending on exercise complexity, intensity, sensor technology, and modeling approach. While velocity-based 1RM prediction demonstrated high average validity, large heterogeneity in lower body exercises significantly biased the results. Furthermore, the dearth of measurement error and agreement analyses prohibits final conclusions. CONCLUSION: Therefore, velocity-based monitoring and 1RM prediction require cautious interpretation, as sensor- and exercise-specific evidence remains limited.

Load–velocity relationship

The reliability of distance run tests for children in grades K-4.

The purpose of this study was to determine test-retest reliability for the 1-mile, 3/4-mile, and 1/2-mile distance run/alk tests for children in Grades K-4. Fifty-one intact physical education classes were randomly assigned to one of the three distance run conditions. A total of 1,229 (621 boys, 608 girls) completed the test-retests in the fall (October), with 1,050 of these students (543 boys, 507 girls) repeating the tests in the spring (May). Results indicated that the 1-mile run/walk distance, as recommended for young children in most national test batteries, has acceptable intraclass reliability (.83 less than R less than .90) for both boys and girls in Grades 3 and 4, has minimal (fall) to acceptable (spring) reliability for Grade 2 students (.70 less than R less than .83), but is not reliable for children in Grades K and 1 (.34 less than R less than .56). The 1/2 mile was the only distance meeting minimal reliability standards for boys and girls in Grades K and 1 (.73 less than R less than .82). Results also indicated that reliability estimates remained fairly stable across gender and age groups from the fall to spring testing periods, with the exception of the noticeably improved values for Grade 2 students on the 1-mile run/walk test. Criterion-referenced reliability (P, percent agreement) was also estimated relative to Physical Best and Fitnessgram run/walk standards. Reliability coefficients for all age group standards were acceptable to high (.70 less than P less than .95), except for Fitnessgram standards for 5-year-old girls on the 1-mile test for both fall and spring and for 6-year-old boys and girls on the 1-mile test administered in the spring.

Child

Getting the story straight: evaluating the test-retest reliability of a university health history questionnaire.

This study was designed to establish the reliability of a health history questionnaire used as a screening tool for incoming university students. The authors used a test-retest design, with a test interval of 6 months, on a sample of medical and nursing students. The analysis focused on overall reliability of the questionnaire and reproducibility of specific items, based on question format. Questionnaire items of specific interest were those with dichotomous yes/no response options versus open-ended format questions, those using the words frequently or recently, or those that asked multiple questions. Demographic characteristics of the subjects were considered in the evaluation of reliability. Overall reliability of the questionnaire (93.6%) was above the anticipated level of 90%, and subject sex or program of study did not show any significant differences in reproducibility of responses. Although wording of questions did not affect item reliability, dichotomous format questions demonstrated a higher degree of reliability (96.4%) than the overall reliability of the questionnaire. Recommendations for enhancing the reliability of the questionnaire are based on item analysis and information gathered from interviews with subjects.

Adult

Factors affecting the reliability of ratings of students' clinical skills in a medicine clerkship.

OBJECTIVE: To determine the overall reliability and factors that might affect the reliability of ratings of students' clinical skills in a medicine clerkship. DESIGN: A nine-item instrument was used to evaluate students' clinical skills. Raters were also asked to provide a grade of each student's overall clinical performance. Generalizability studies were performed to estimate the reliability of the ratings. The effects of rater experience and clerkship setting were investigated by regression analysis. SETTING: Teaching hospitals and community-based sites in three Northwestern states. PARTICIPANTS: All students (328) who had completed the 12-week clerkship in internal medicine at one medical school during the academic years 1987-1989. Raters included attending physicians, chief residents, and other residents. RESULTS: Seven observations were needed to provide a reliable rating of the overall clinical grade. More observations were needed to obtain reliable ratings for individual items, ranging from seven observations needed for the rating of data gathering skills to 27 observations needed for the rating of interpersonal relationships with patients. Rater experience and clerkship setting (i.e., teaching hospitals vs. community-based clinics) were found, in general, not to affect significantly the ratings received by students. CONCLUSIONS: Reliable ratings of students' overall clinical skills, including overall clinical grades, can be achieved by collecting a minimum of seven observations. More observations are needed to measure reliably the interpersonal aspects of clinical performance. These findings support the use of performance ratings to evaluate clinical skills and knowledge of students in clerkship settings.

Clinical Clerkship

Reliability of the TIP and DIP speech-hearing tests for children.

The reliability of SRT and speech intelligibility tests has been studied on adults. Reliability estimates for SRT are between .60 and .90, with standard error estimates from 1.5 to over five dB. For speech intelligibility tests the reliability estimates range from .50 to 90, with standard error of estimates from 2.5% to over 10%. Little has been reported on test reliability with children. For this study the TIP and DIP tests, for threshold and discrimination, respectively, were given to 295 normal and 138 hypacusic children three through twelve years of age. Subjects were retested within one week. TIP test-retest reliability was .72 for normals, and .89 to .99 for hypacusics. DIP test-retest reliability was .46 to .51 for normals and .60 to .93 for hypacusics. Standard error of estimate was about 3 dB for TIP, and 10% for DIP. These values are about the same as the reliability values for adults.

Child

Validity and reliability of lupus activity measures in the routine clinic setting.

As part of a cohort study of 150 patients with systemic lupus erythematosus (SLE), we investigated the validity and reliability of several indices of lupus activity, including the UCSF/JHU Lupus Activity Index (LAI), the SLE Disease Activity Index (SLEDAI), and a simple Core Index combining common elements. Validity was assessed by measuring correlations of these indices at the first cohort visit with the physician's global assessment (PGA) of SLE activity. The correlation of M-LAI (LAI modified so as not to contain PGA) and SLEDAI with PGA was 0.64 (95% CI 0.50, 0.70) and 0.55 (95% CI 0.42, 0.64), respectively. Reliability was assessed in a study of 6 patients seen twice, one week apart, by 9 physicians. The interrater reliability and test-retest reliability was greater for LAI (or M-LAI) than for SLEDAI. The Core Index performed better in its correlation with PGA (R = 0.78), although it contained no treatment data or serologic tests. Its interrater reliability and test-retest reliability were comparable with LAI. We conclude that (1) all indices have high validity; (2) LAI and the Core Index have higher reliability; and (3) these indices can be readily assimilated into routine clinic practice.

Adult

AIDS knowledge and attitudes among injection drug users: the issue of reliability.

Among injection drug users (IDUs), AIDS-related knowledge and attitudes have not consistently predicted AIDS risk behavior. This may be due in part to the limited reliability of indexes used to measure drug users' AIDS knowledge and attitudes. In addition, the substantive interpretation of findings is confounded if index reliability is lower for particular demographic groups (e.g., ethnic populations and women). This report is based on 8 measures of AIDS-related knowledge and attitudes in a sample of 332 injection drug users in Los Angeles. The reliability of knowledge and attitude indexes for the overall sample is generally acceptable for the purpose of group comparison (average alpha = .60). But reliability is consistently lower for respondents who are Hispanic (average alpha = .49) and respondents with less formal education (alpha = .56). The reliability of 2 measures of sex-related attitudes is lower for female respondents. It is therefore important that the reliability of knowledge and attitude indexes be assessed not just for drug-user samples as a whole, but also within demographic groups of substantive interest.

Acquired Immunodeficiency Syndrome

Reliability and validity in binary ratings: areas of common misunderstanding in diagnosis and symptom ratings.

Confusion may exist between the reliability of a binary rating (for example, schizophrenia versus not-schizophrenia) and its implications for validity. High reliability does not guarantee validity, but paradoxically, low reliability does not imply poor validity in all contexts. Changes in the base rate or in experimental design may indicate high validity even when the reliability was thought to be low. Attempts to improve the psychiatric nomenclature by increasing only reliability run the risk of the "attenuation paradox" where further increases in reliability will make the ratings less valid. Finally, the assumption of random error in making diagnoses does not always hold, so that statistical analyses must be adjusted accordingly. New statistical methods are needed to index only false-positive or false-negative rates in order to quantify the error that will reduce some validity coefficients.

Bipolar Disorder

Reliability of seven measures of social intelligence in a sample of adolescents with mental retardation.

This study evaluates the reliability of seven measures, selected to assess the social-cognitive variables hypothesized by Greenspan to define social intelligence. Responses from 75, 30 and 20 adolescents with mental retardation were used to assess each test's internal, interrater, and test-retest reliabilities, respectively. Interrater reliability coefficients were high to very high (.76 to .98), internal reliabilities were moderate to very high (.66 to .90), and test-retest reliabilities were moderate to high (.54 to .74). Internal and test-retest reliability coefficients compared favourably with those reported for the subtests of the Revised Wechsler Intelligence Scale for Children.

Activities of Daily Living

Intrarater reliability of manual muscle test (Medical Research Council scale) grades in Duchenne's muscular dystrophy.

The purpose of this study was to document the intrarater reliability of manual muscle test (MMT) grades in assessing muscle strength in patients with Duchenne's muscular dystrophy (DMD). Subjects were 102 boys, aged 5 to 15 years, who were participating in a double-blind, multicenter trial to document the effects of prednisone on muscle strength in patients with DMD. Four physical therapists participated in the study. Two identical (duplicate) evaluations were performed within 5 days of each other by the same examiner initially and after 6 and 12 months of treatment. A total of 18 muscle groups were tested on each patient, 16 of them bilaterally, using a modification of the Medical Research Council scale. Reliability of muscle strength grades obtained for individual muscle groups and of individual muscle strength grades was analyzed using Cohen's weighted Kappa. The reliability of grades for individual muscle groups ranged from .65 to .93, with the proximal muscles having the higher reliability values. The reliability of individual muscle strength grades ranged from .80 to .99, with those in the gravity-eliminated range scoring the highest. We conclude the MMT grades are reliable for assessing muscle strength in boys with DMD when consecutive evaluations are performed by the same physical therapist.

Adolescent

Reliability of lumbar isometric torque in patients with chronic low back pain.

In this study, the test-retest reliability of lumbar isometric strength testing in patients with chronic low back pain (CLBP) was assessed. Isometric torque measurements were obtained from 89 patients with CLBP at seven different angles of lumbar flexion. Because previous studies have demonstrated significant strength differences between male and female subjects, separate data analyses were performed for each gender. Results indicated moderate to high reliability for patients with CLBP when tested at individually determined angles of flexion within their idiosyncratic range of motion (ROM) (female subjects: r = .59-.96, P less than .05, SEE = 12.0-24.2 N.m; male subjects: r = .71-.93, P less than .05, SEE = 25.1-62.1 N.m). For comparison with previously published data on asymptomatic controls, an additional set of analyses was conducted for subjects with full lumbar ROM. Similar reliability was demonstrated for this subsample (female subjects: r = .57-.93, P less than .05, SEE = 12.4-27.9 N.m; male subjects: r = .63-.93, P less than .05, SEE = 34.2-44.2 N.m). The authors concluded that isometric lumbar extension torque could be reliably measured in patients with CLBP at multiple positions within the full ROM, although reliability decreased at the most extended positions. The demonstrated reliability will allow researchers to assess treatment effects and group differences without undue concern for artifact attributable to measurement error.

Back Pain

On the methods and theory of reliability.

This paper reviews the most frequently used and misused reliability measures appearing in the mental health literature. We illustrate the various types of data sets on which reliability is assessed (i.e., two raters, more than two raters, and varying numbers of raters with dichotomous, polychotomous, and quantitative data). Reliability statistics appropriate for each data format are presented, and their pros and cons illustrated. Inadequancies of some methods are highlighted. The meaning of different levels of reliability obtained with various statistics is discussed. This critique is intended for the reading professional and the investigator who has an occasional need for reliability assessment. Statistical expertise is not required and theoretical material is referenced for the interested reader. Necessary formulas for computations are presented in the appendices. A summary table of some suitable reliability measures is presented.

Humans

The diagnosis of hypersensitivity to ingested foods. Reliability of skin prick testing and the radioallergosorbent test with different materials.

The diagnostic reliability in food allergy of skin prick tests (SPT) and the radio-allergosorbent test (RAST) was investigated in paediatric patients with respiratory and skin allergies. SPT and RAST were found to be reliable for the diagnosis of allergy to codfish, peas, nuts, peanuts and egg white. Positive SPT and RAST to cereals were common, but were most often without clinical significance or were correlated with respiratory allergy to the inhalation of flour dust. SPT and RAST were only partly reliable with regard to allergy to cow's milk, and were mostly reliable when used together and showing corresponding results. Experimental allergosorbents for RAST with soy beans and white beans were not reliable. The study shows the need to improve the diagnostic materials and to establish the diagnostic reliability of the material and tests used for each food item in question.

Adolescent

Percent of agreement among raters and rater reliability of the copying subtest of the Stanford-Binet Intelligence Scale: Fourth Edition.

The purpose of this study was to investigate the interrater reliability of the visual-motor portion of the Copying subtest of the Stanford-Binet Intelligence Scale: Fourth Edition. Eight raters independently scored 11 protocols completed by children aged 5 through 10 years, using the scoring criteria and guidelines in the manual. The raters marked each of 10 items pass or fail and computed a total raw score for each protocol. Interrater reliability coefficients were obtained for each child's protocol, and the Kappa coefficient was computed for each item. Significant raters' reliability coefficients ranged from .82 to .91, which were low in comparison to test-retest reliability and Kuder-Richardson-20 coefficients for this and other subtests of the Stanford-Binet in the technical manual. Percent agreement among 8 raters also indicated weak reliability. Although the obtained results suggested some interrater reliability coefficients within acceptable levels, questions were raised about the scoring criteria for individual items. Caution is warranted in the use of cognitive measures which include subjective judgement of the examiner in applying scoring criteria.

Child

Number of stimuli as a reliability parameter in perimetry.

Catch trials test patient performance during automated, static perimetry, but their adequacy to estimate reliability is uncertain even though up to 10% of the test time is reserved for catch trials. The 308 visual fields (program G1, all 3 phases, Octopus 201) of 308 eyes of 308 glaucoma, suspected glaucoma, and normal subjects were studied. The 108 visual fields (mean sensitivity > 10 dB; corrected loss variance < 50 dB2) without false responses to catch trials were considered reliable. A multiple linear regression analysis of these 108 fields was performed and revealed the following result (r2 = 0.751): Number of stimuli = 480 + (40.short-term fluctuation) + (8.8.the square root of the index corrected loss variance) - (2.2.mean sensitivity). This equation was used to estimate the number of stimuli required of a reliable subject to complete an examination. Excess stimuli would thus be a sign of reduced reliability. The difference between the estimated and the actual number of stimuli was called the 'stimulus discrepancy'. In 169 fields with false-positive and 58 fields with false-negative responses, the false-positive and false-negative responses correlated with the 'stimulus discrepancy' (r = 0.19, P = 0.014; r = 0.29, P < 0.026, respectively). The number of stimuli depends not only on reliability but also on the software and hardware of the perimeter. 'Stimulus discrepancy' may be an additional useful perimetric reliability parameter which does not require extra testing time.

Adult