PubMed HealthSearch

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Reliability of psychiatric diagnosis. II. The test/retest reliability of diagnostic classification.

In a study of interrater diagnostic reliability, 101 psychiatric inpatients were independently interviewed by physicians using a structured interview. Newly admitted patients were randomly selected and examined by one of three psychiatrists. A second psychiatrist reexamined the same patient about 24 hours later. Interviews from the two examinations were evaluated independently and diagnoses were made on the basis of objective criteria. The degree of diagnostic agreement for the two examinations were calculated using the kappa statistic. Agreement was found to be high as compared to other studies in the psychiatric literature, despite the fact that in most previous investigations diagnoses were not made independently. The results were also compared to studies of reliability of medical judgments. Possible reasons for the high interrater reliability are discussed and include the use of a structured interview and objective diagnostic criteria.

Attitude of Health Personnel

Reliability, Device Agreement and Validity of Load-Velocity Profiles: A Systematic Review with Meta-analysis.

BACKGROUND: For a valid one-repetition maximum (1RM) prediction via load-velocity (LV) relationships, high reliability and accuracy must be assumed. OBJECTIVE: Since individual study results indicate ambivalent prediction, this systematic review and meta-analysis was designed to provide a updated and comprehensive overview, extending knowledge about the validity and reliability of commercially available velocity sensors in Part I and the validity and reliability of velocity-based 1RM prediction models in Part II. METHODS: A systematic literature search was conducted in PubMed/MEDLINE, Web of Science, and Scopus. Validity and/or reliability studies or velocity-based 1RM prediction evaluations were included. Methodological quality was assessed using adapted COSMIN. The analysis was performed for intraclass correlation coefficient (ICC), Lin's concordance correlation coefficient (CCC), and Pearson's correlation coefficient (r). The review was preregistered in PROSPERO (CRD42025634595). RESULTS: Sixty-three studies were included for sensor validity and reliability and 38 for 1RM prediction models. Part I: Velocity sensors demonstrated good-to-excellent pooled validity and device agreement (ICC = 0.91-0.92 [0.83-0.97]; k = 55 and 439, respectively); intra- and inter-day reliability were classified as good to excellent with ICC = 0.90-0.91 [0.85-0.95] (k = 228 and 608, respectively), with sensor technology moderating the results. However, substantial heterogeneity and wide ranges of study-level estimates indicated considerable variability across moderators, linear position transducer (LPT) generally showing more consistent performance than inertial measurement units (IMU). Part II: Velocity-based 1RM prediction showed ICCs = 0.90 [0.83-0.94] (k = 124) and ICC = 0.91 [0.72-0.98] (k = 9); for reliability and validity, respectively. DISCUSSION: Commercial velocity sensors generally provide high relative validity and reliability. Results varied depending on exercise complexity, intensity, sensor technology, and modeling approach. While velocity-based 1RM prediction demonstrated high average validity, large heterogeneity in lower body exercises significantly biased the results. Furthermore, the dearth of measurement error and agreement analyses prohibits final conclusions. CONCLUSION: Therefore, velocity-based monitoring and 1RM prediction require cautious interpretation, as sensor- and exercise-specific evidence remains limited.

Load–velocity relationship

Reliability of the TIP and DIP speech-hearing tests for children.

The reliability of SRT and speech intelligibility tests has been studied on adults. Reliability estimates for SRT are between .60 and .90, with standard error estimates from 1.5 to over five dB. For speech intelligibility tests the reliability estimates range from .50 to 90, with standard error of estimates from 2.5% to over 10%. Little has been reported on test reliability with children. For this study the TIP and DIP tests, for threshold and discrimination, respectively, were given to 295 normal and 138 hypacusic children three through twelve years of age. Subjects were retested within one week. TIP test-retest reliability was .72 for normals, and .89 to .99 for hypacusics. DIP test-retest reliability was .46 to .51 for normals and .60 to .93 for hypacusics. Standard error of estimate was about 3 dB for TIP, and 10% for DIP. These values are about the same as the reliability values for adults.

Child

Reliability and validity in binary ratings: areas of common misunderstanding in diagnosis and symptom ratings.

Confusion may exist between the reliability of a binary rating (for example, schizophrenia versus not-schizophrenia) and its implications for validity. High reliability does not guarantee validity, but paradoxically, low reliability does not imply poor validity in all contexts. Changes in the base rate or in experimental design may indicate high validity even when the reliability was thought to be low. Attempts to improve the psychiatric nomenclature by increasing only reliability run the risk of the "attenuation paradox" where further increases in reliability will make the ratings less valid. Finally, the assumption of random error in making diagnoses does not always hold, so that statistical analyses must be adjusted accordingly. New statistical methods are needed to index only false-positive or false-negative rates in order to quantify the error that will reduce some validity coefficients.

Bipolar Disorder

On the methods and theory of reliability.

This paper reviews the most frequently used and misused reliability measures appearing in the mental health literature. We illustrate the various types of data sets on which reliability is assessed (i.e., two raters, more than two raters, and varying numbers of raters with dichotomous, polychotomous, and quantitative data). Reliability statistics appropriate for each data format are presented, and their pros and cons illustrated. Inadequancies of some methods are highlighted. The meaning of different levels of reliability obtained with various statistics is discussed. This critique is intended for the reading professional and the investigator who has an occasional need for reliability assessment. Statistical expertise is not required and theoretical material is referenced for the interested reader. Necessary formulas for computations are presented in the appendices. A summary table of some suitable reliability measures is presented.

Humans

The diagnosis of hypersensitivity to ingested foods. Reliability of skin prick testing and the radioallergosorbent test with different materials.

The diagnostic reliability in food allergy of skin prick tests (SPT) and the radio-allergosorbent test (RAST) was investigated in paediatric patients with respiratory and skin allergies. SPT and RAST were found to be reliable for the diagnosis of allergy to codfish, peas, nuts, peanuts and egg white. Positive SPT and RAST to cereals were common, but were most often without clinical significance or were correlated with respiratory allergy to the inhalation of flour dust. SPT and RAST were only partly reliable with regard to allergy to cow's milk, and were mostly reliable when used together and showing corresponding results. Experimental allergosorbents for RAST with soy beans and white beans were not reliable. The study shows the need to improve the diagnostic materials and to establish the diagnostic reliability of the material and tests used for each food item in question.

Adolescent

An Assessment of Reliability Estimation Methods for Binomial Health Care Quality Measures.

We evaluated the performance of commonly used methods for estimating the reliability of binomial health care quality measures using simulated datasets spanning a range of performance score means and variances, numbers of entities, and patient sample sizes. For each simulation, reliability was estimated for all selected methods and compared with the known true reliability derived from the simulation parameters, with methods assessed on their accuracy and precision. Logistic regression with reliability estimated on the outcome scale demonstrated the highest accuracy and precision among all methods evaluated. The widely used Adams beta-binomial method performed poorly, although a modification recommended by Nieser and Harris substantially improved its performance. These approaches are applicable only to binomial measures. Among methods that can be applied to both binomial and continuous measures, permutation resampling of the Spearman rank correlation coefficient was the most accurate and precise, outperforming other commonly used approaches. Overall, for binomial quality measures, logistic regression on the outcome scale is the preferred method for reliability estimation, followed closely by the modified beta-binomial approach, while for non-binomial measures, permutation-based Spearman rank correlation appears to be the most suitable method.

Reproducibility of Results

The interrater reliability of the Psychotic Inpatient Profile.

The interrater reliabilities of the 12 scales of the Psychotic Inpatient Profile (PIP) were assessed in three independent samples. These reliabilities were found to be consistently low for the three samples. Several possible sources of low reliability were investigated, and evidence to support the hypothesis that consistent rater biases might have contributed to this low reliability was found. The authors make several suggestions of ways to improve PIP's reliability.

Adjustment Disorders

Comparative reliability of categorical and analogue rating scales in the assessment of psychiatric symptomatology.

The reliability of 26 items from the ninth edition of the Present State Examination (PSE) was assessed using both the conventional categorical scales and separately constructed analogue scales. Reliability was also calculated when the analogue responses were rescaled down to 2, 3 and 4 categories. The levels of inter-rater agreement obtained were comparable to those achieved in previous studies of PSE reliability, although as expected the levels of agreement on audiotapes were greater than those for independent interviews performed on the same day. These levels were not significantly affected by any of the changes in scale format, but there were apparent differences in reliability depending on the statistics used. In selecting or constructing a psychiatric rating scale, the question of reliability should not influence the choice of a categorical or continuous scale, or the number of scored points in the scale.

Psychiatric Status Rating Scales

Reliability of goniometric measurements of hip motion in spastic cerebral palsy.

Reliability of goniometric measurements of ranges of motion in the right hip of four children with mild or moderate spastic diplegia was studied by comparing results when physical therapists used specific and non-specific measurement instructions. Measurements of hip extension, abduction and external rotation were made. The use of specific measurement instructions improved inter-rater reliability only in the case of external rotation. Inter-session reliability was not improved by the use of specific measurement instructions. It is concluded that goniometric measurements have such a low level of reliability that they can only be used to assist clinical judgment, and that they are not sufficiently reliable to be used in studied of cerebral palsy.

Adolescent

Diagnostic reliability of the percutaneous ultrasonic Doppler technique for vertebral arterial occlusive diseases.

There is little data on the diagnostic reliability of the ultrasonic Doppler technique for vertebral arterial occlusive lesions. Percutaneous vertebral Doppler examination and the vertebral angiograms were compared to determine the diagnostic reliability of this technique in 64 vertebral arteries of 53 patients with cerebrovascular disease. The percutaneous vertebral Doppler findings were quantitatively analyzed using a sound spectrograph and were classified into three types: no flow signal type, poor flow type and normal flow type. In nine patients with the no flow type the angiograms revealed vertebral occlusion or a missing vertebral artery in six, giving a diagnostic reliability of 67%. In 17 patients with poor flow type the angiograms revealed vertebral occlusion or a missing vertebral artery in five, terminal narrowing of the artery in nine, and hypoplasia in two giving a diagnostic reliability of 94%. For all vertebral arteries examined with this technique, including normal ones, the diagnostic reliability was 92% (59/64). Percutaneous vertebral Doppler examination has clinical usefulness as a screening test for occlusive vertebral arterial diseases.

Adult

Absolute and relative reliability of several response parameters used in vestibular assessment.

This experiment investigated the reliability of five indices used to express the magnitude of induced vestibular nystagmus. The investigation was undertaken because information concerning the reliability of the various response parameters when caloric testing is repeated on the same subject is limited. Electronystagmographic assessments of the nystagmus response to the Fitzgerald-Hallpike test were obtained from 16 subjects on three occasions with equal intervals of time between occasions. The data were statistically treated with an analysis of variance technique that allows the generation of correlation coefficients. The findings showed that speed of the slow phase, total beats, culmination frequency, and total amplitude reliably express the magnitude of induced nystagmus on repeated tests. Duration of the induced reaction was not as reliable as the other indices on repeated testing. In addition, difference scores were more repeatable than absolute scores and the caution that must be maintained when using correlation coefficient analyses to estimate reliability is highlighted.

Adolescent

Response reliability in a longitudinal survey in Thailand.

The two rounds of the National Longitudinal Study in Thailand provide a useful opportunity to explore response reliability in a large-scale social and demographic survey in a developing country. The results indicate that nonrandom reliability at the individual level ranged from quite high (for several straightforward, facual questions) to quite low (for most attitudinal questions). There was considerable distributional stability, however, even for many of the variables with low individual-level reliability. In terms of its response reliability, the Thai study compares reasonably well with several leading US fertility surveys. However, in both countries response reliability at the individual level for attitudinal questions is distressingly low. This clearly should be a matter of major concern for social scientists using survey results.

Demography

Levels of reliability in fertility survey data.

A number of factual and attitudinal questions asked in the 1973 Taiwan KAP-4 survey were repeated in a postenumeration survey one month later in order to assess the reliability of responses of the 286 women reinterviewed. The level of reliability is found to vary depending on the measures used and on whether the focus is aggregate data or individual responses. Analysis of consistency of responses shows that while overall reliability at both aggregate and individual levels is reasonably good, there is greater reliability for factual than for attitudinal data. Nevertheless, consistency of responses on factual questions varies considerably depending on the salience of the topic to the respondent. Estimates of reliability are shown to depend on the measure used and on the skewness of the distributions of the responses.

Contraception

Reliability of psychiatric diagnosis. I. A methodological review.

This article reviews some methodological aspects of studies of diagnostic reliability in psychiatry. We define and discuss the concept of interrater reliability and review some of the ways in which the design of the reliability study can influence the results. Three basic methodological issues are raised, including: importance of structured interviews and objective diagnostic criteria, the importance of a test/retest vs an interviewer/observer design, and the calculation of reliability in a way that takes chance agreement into account.

Attitude of Health Personnel

Reliability of the Group for Advancement of Psychiatry Diagnostic Categories in Child Psychiatry.

A total of 403 multiple diagnoses were independently assigned to 41 patient protocols by 73 psychiatrists, psychologists, and social workers to determine the levels of interrater reliability of the Group for the Advancement of Psychiatry (GAP) diagnostic categories. With the exception of the psychotic disorders category, these diagnostic categories were found to have low levels of interdiagnostician reliability. Differences in the reliabilities across disciplines and levels of training were found. It is noted, however, that neither years of experience, kind of training, nor direct contact with the patient can be regarded as a substitute for improvements in the classification system itself. The importance of a reliable classification system for child psychiatry is emphasized and suggestions for improvements in the present GAP system are made.

Child

Reliability and stability of the Comrey Personality Scales in a clinical setting.

Examined split-half reliabilities, retest reliabilities, and stability of the Comrey Personality Scales (CPS) (Comrey, 1970) for a Navy sample of 200 young drug abusers tested at the beginning and end of rehabilitation in a no-feedback, compulsory participation setting. The pre-rehabilitation split-half reliabilities for the eight personality scales ranged from .73 to .94, with an average of .85, and ranged from .74 to .91, with an average of .83, on the post-rehabilitation administration. Retest reliabilities for the eight scales were between .39 and .64, with an average of .52. Post-rehabilitation means were significantly higher on four of the eight scales, and one scale mean was significantly lower. It is concluded that the CPS is appropriate in a clinical setting in which test participation is compulsory and test feedback is not made available to the respondents.

Adolescent

The reliability of reaction time determinations.

The reliability of simple and two-choice reaction time (RT) as a function of number of trials and measure of central tendency was examined. It was found that, within the particular conditions of the study, 18 trials were sufficient to obtain a highl reliable determination of simple RT but 30 trials were required for a highly reliable determination of two-choice RT. The selection of measure of central tendency was not related to degree of reliability.

Adult