PubMed HealthSearch

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Item reliability of the Milani-Comparetti Motor Development Screening Test.

The purpose of this study was to determine the level of interobserver and test-retest reliability of the Milani-Comparetti Motor Development Screening Test. Sixty healthy children, aged 1 through 16 months, were videotaped during administration of the Milani-Comparetti test. Four pediatric physical therapists independently viewed each videotape and scored the responses. Interobserver reliability was determined by calculation of percentage of agreement and the G statistic between a primary observer and each therapist. Forty-three children were retested within one week by the initial tester to examine test-retest reliability. Test-retest reliability was determined by percentage of agreement of items between the two test sessions and using the Kappa statistic. Interobserver percentage of agreement for the individual items on the Milani-Comparetti test ranged from 79% to 98%. The G statistic was significant for all items indicating the high percentage-of-agreement values were not due merely to chance agreement. Test-retest agreement ranged from 80% to 100%. Using Kappa statistic guidelines, excellent test-retest reliability (K greater than .75) was found for 82% of the test items, with good reliability of the remaining items. Acceptable interobserver and test-retest reliability was found for all items on the Milani-Comparetti test. Use of the Milani-Comparetti test as a clinical screening tool for prediction or follow-up of motor development in children at risk for developmental delays requires further evaluation.

Child Development

Reliability of goniometric measurements and visual estimates of knee range of motion obtained in a clinical setting.

The purpose of this study was to examine the intratester and intertester reliability for goniometric measurements of knee flexion and extension passive range of motion (PROM). In addition, parallel-forms reliability for PROM measurements of the knee obtained by use of a goniometer and by visual estimation was examined. The intertester reliability for visual estimates of the PROM of the knee was also examined. Repeated measurements were obtained on 43 patients in a clinical setting. The intraclass correlation coefficients (ICCs) for intratester reliability of measurements obtained with a goniometer were .99 for flexion and .98 for extension. Intertester reliability for measurements obtained with a goniometer was .90 for flexion and .86 for extension. The ICCs for parallel-forms reliability for measurements obtained with a goniometer and by visual estimation ranged from .82 to .94. The intertester reliability for measurements obtained by visual estimation was .83 for flexion and .82 for extension. Results suggest clinicians should use a goniometer to take repeated PROM measurements of a patient's knee to minimize the error associated with these measurements.

Adult

Facioscapulohumeral dystrophy natural history study: standardization of testing procedures and reliability of measurements. The FSH DY Group.

BACKGROUND AND PURPOSE: The natural history of facioscapulohumeral muscular dystrophy (FSHD) has not been studied prospectively. Knowledge of the natural progression of any disease provides essential information for the design of clinical trials. We present a protocol for the study of the natural history of FSHD using quantitative muscle testing (QMT), manual muscle testing (MMT), and functional testing. SUBJECTS: Thirty-two persons with FSHD (mean age = 36.1 years, SD = 9.6, range = 17-49) and 32 age- and gender-matched volunteer controls (mean age = 35.8 years, SD = 8.0, range = 23-50) served as subjects. METHODS: Using standardized testing procedures, we examined intrarater reliability of the MMT, QMT, and functional testing measurements in both groups. We also examined interrater reliability in 7 subjects with FSHD. Eighteen muscle groups were tested for each subject using QMT and MMT. RESULTS: Intraclass correlation coefficient (ICC) values ranged from .86 to .99 for intrarater reliability and from .86 to .99 for interrater reliability of QMT measurements. Weighted kappa values of .81 to .98 for intrarater reliability and .50 to 1.00 for interrater reliability were obtained for MMT measurements. Intrarater ICCs for various functional testing measures ranged from .60 to .97. In addition, the comparability of the two QMT machines used in the study was demonstrated by testing the same set of volunteer controls on each machine's linear force transducer (ICC = .89-.98). CONCLUSION AND DISCUSSION: We conclude that this standardized testing protocol produces reliable measurements of muscle strength and functional ability in subjects with FSHD.

Adolescent

Description and interobserver reliability of the Tufts Assessment of Motor Performance.

This paper describes the conceptual basis for the development of a new clinical evaluation instrument, the Tufts Assessment of Motor Performance (TAMP). The TAMP is a 32-item, diagnosis-independent, criterion-referenced test that samples physical performance items in the areas of mobility, activities of daily living and physical aspects of communication. The administrative and scoring criteria of the TAMP are presented, and the multiple measurement dimensions are described. The documentation of patient status and progress, as described in the functional and performance profiles, is outlined. The paper also reports initial interobserver reliability on the intraitem tasks and the summary indexes of the two profiles. Forty individuals (20 adults and 20 children) with neurologic and musculoskeletal disorders comprised the reliability sample. Kappa and intraclass correlations were used to estimate the reliability of three independent raters on individual tasks and aggregate scores, respectively. Task reliability for the assistance and approach measurement dimensions were generally higher than for the more qualitative pattern and proficiency dimensions. Yet over 90% of all the tasks had acceptable reliability, while all the summary indexes had high interobserver reliability. Determination of interobserver reliability data is the initial phase of defining the most appropriate and technically valuable items, and will serve as a basis for item revision and reduction to enhance the clinical utility of the test.

Activities of Daily Living

Reliability of transient-evoked otoacoustic emissions.

OBJECTIVE: This investigation addressed four factors affecting transient-evoked otoacoustic emission (TEOAE) reliability: 1) The effect of evoking-stimulus level, 2) the effect of analyzing bandwidth, 3) the effect of slight-mild hearing loss, and 4) the effect of variability in the stimulus spectrum. DESIGN: TEOAEs at 80, 74, 68, and 62 dB pSPL evoking-stimulus levels were measured in 25 ears spanning a range of hearing levels from normal to mild hearing loss for a minimum of 10 test sessions. Reliability was assessed for 1/6-, 1/3-, 1/2-, and 1-octave analyzing bandwidths. RESULTS: Evoking-stimulus level, hearing loss, and center frequency did not significantly affect reliability. With decreasing analyzing bandwidth, reliability decreased. Intrasubject test-retest standard deviations were 1.2 dB for a broadband analyzing bandwidth and 1.4, 1.5, 1.6, and 1.8 dB for 1-, 1/2-, 1/3-, and 1/6-octave analyzing bandwidths, respectively. Stimulus variability within narrower bandwidths was of sufficient magnitude to influence test-retest reliability, and attempts to correct for the variations in stimulus spectrum were unsuccessful. Slopes of the input-output functions differed across frequencies, with shallower slopes at higher frequencies. CONCLUSIONS: In general, TEOAE amplitude is highly reliable. For those individuals in this study who were more variable, the variability was at low frequencies or across the entire frequency spectrum. For clinical applications, the choice of analyzing bandwidth should be based on consideration of both frequency specificity (where narrow analyzing bandwidths are optimal) and reliability (where wide analyzing bandwidths are optimal).

Acoustic Impedance Tests

Reliability of variables in the kinematic analysis of spring hurdles.

The purpose of this study was to investigate the reliability of kinematic variables in spring hurdles and to find out how many trials are needed to achieve reliable data. Seven British National level athletes in sprint hurdles were videotaped and all eight trials of each athlete were digitized from two camera views to produce three dimensional coordinates. The reliability of 28 kinematic variables across eight trials ranged from 0.54 to 1.00 for females and from 0.00 to 0.99 for males. The number of trials needed to reach a certain reliability level was evaluated using Spearman-Brown prophecy formula, and in the worst case (horizontal velocity lost for males) 78 trials would be needed to reach 0.90 reliability. The results showed reasonably high reliability, and the values for the female trials were generally higher than the male trials. The relative height of the hurdles enforces a more demanding clearance for males that can lead to increased variation within the subjects and thus lowered reliability. Subsequently, the results indicate that often more than one trial is needed to provide accurate quantitative results of the technique.

Adult

Interexaminer reliability in physical examination of patients with low back pain.

STUDY DESIGN: Seventy-one patients with low back pain were examined by two physiotherapists (50 patients) and two physicians (21 patients). The two physiotherapists had worked together for many years, but the two physicians had not. The interexaminer reliability of the clinical tests included in the physical examination was evaluated. OBJECTIVES: To evaluate the interexaminer reliability of clinical tests used in the physical examination of patients with low back pain under ideal circumstances, which was the case for the physiotherapists. SUMMARY OF BACKGROUND DATA: Numerous clinical tests are used in the evaluation of patients with low back pain. To reach the correct diagnosis, only tests with an acceptable validity and reliability should be used. Previous studies have mainly shown low reliability. It is important that clinical tests not be rejected because of low reliability caused by differences between examiners in performance of the examination and in their definition of normal results. METHODS: Two examiners, either two physiotherapists or two physicians, independently examined patients with low back pain. RESULTS: In approximately half of the clinical tests studied, an acceptable reliability was demonstrated. CONCLUSION: On the basis of the physiotherapists series, the reliability was acceptable for a number of clinical tests that are used in the evaluation of patients with low back pain. The results suggest that clinical tests should be standardized to a much higher degree than they are today.

Adolescent

Apophysial joint degeneration, disc degeneration, and sagittal curve of the cervical spine. Can they be measured reliably on radiographs?

STUDY DESIGN: Interexaminer reliability study. OBJECTIVES: To determine the reliability of grading apophysial joint and disc degenerative changes and the reliability of measuring sagittal curves on lateral cervical spine radiographs. SUMMARY OF BACKGROUND DATA: Several authors have proposed that the presented of degenerative changes and the absence of lordosis in the cervical spine are indicators of poor recovery from neck injuries caused by motor vehicle collisions. The validity of those conclusions is questionable because the reliability of the methods used in their studies to measure the presence of degenerative changes and the absence of lordosis has not been determined. METHODS: Kellgren's classification system for apophysial joint and disc degeneration, as well as the pattern and magnitude of the sagittal curve on 30 lateral cervical spine radiographs were assessed independently by three examiners. RESULTS: Moderate reliability was demonstrated for classifying apophysial joint degeneration with an intraclass correlation coefficient of 0.45 (95% confidence interval, 0.09-0.71). Classifying degenerative disc disease had substantial reliability, with an intraclass correlation coefficient of 0.71 (95% confidence interval, 0.23-0.88). Measuring the magnitude of the sagittal curve from C2 to C7 had excellent interexaminer agreement, with an intraclass correlation coefficient of 0.96 (95% confidence interval, 0.88-0.98) and an interexaminer error of 8.3 degrees. CONCLUSIONS: The classification system for degenerative disc disease proposed by Kellgren et al and the method of measurement of sagittal curves from C2 to C7 demonstrated an acceptable level of reliability and can be used in outcomes research.

Analysis of Variance

Teaching DSM-III to clinicians. Some problems of the DSM-III system reducing reliability, using the diagnosis and classification of depressive disorders as an example.

Experiences from teaching DSM-III to more than three hundred Norwegian psychiatrists and clinical psychologists suggest that reliable DSM-III diagnoses can be achieved within a few hours training with reference to the decision trees and the diagnostic criteria only. The diagnoses provided are more reliable than the corresponding ICD diagnoses which the participants were more familiar with. The three main sources of reduced reliability of the DSM-III diagnoses are related to: poor knowledge of the criteria which often is connected with failure of obtaining diagnostic key information during the clinical interview; unfamiliar concepts and vague or ambiguous criteria. The two first issues are related to the quality of the teaching of DSM-III. The third source of reduced reliability reflects unsolved validity issues. By using the classification of five affective case stories as examples, these sources of diagnostic pitfalls, reducing reliability and ways to overcome these problems when teaching the DSM-III system, are discussed. It is concluded that the DSM-III system of classification is easy to teach and that the system is superior to other classification systems available from a reliability point of view. The current version of the DSM-III system, however, partly owes a high degree of reliability to broad and heterogeneous diagnostic categories like the concept major depression, which may have questionable validity. Thus, the future revisions of the DSM-III system should, above all, address the issue of validity.

Depressive Disorder

Action potential propagation through embryonic dorsal root ganglion cells in culture. II. Decrease of conduction reliability during repetitive stimulation.

1. The reliability of the propagation of action potentials (AP) through dorsal root ganglion (DRG) cells in embryonic slice cultures was investigated during repetitive stimulation at 1-20 Hz. Membrane potentials of DRG cells were recorded intracellularly while the axons were stimulated by an extracellular electrode. 2. In analogy to the double-pulse experiments reported previously, either one or two types of propagation failures were recorded during repetitive stimulation, depending on the cell morphology. In contrast to the double-pulse experiments, the failures appeared at longer interpulse intervals and usually only after several tens of stimuli with reliable propagation. 3. In the period with reliable propagation before the failures, a decrease in the conduction velocity and in the amplitude of the afterhyperpolarization (AHP), an increase in the total membrane conductance, and the disappearance of the action potential "shoulder" were observed. 4. The reliability of conduction during repetitive stimulation was improved by lowering the extracellular calcium concentration or by replacing the extracellular calcium by strontium. The reliability of conduction decreased by the application of cadmium, a calcium channel blocker, 4-amino pyridine, a fast potassium channel blocker, or apamin or muscarine, the blockers of calcium-dependent potassium channels. The reliability of conduction was not effected by blocking the sodium potassium pump with ouabain or by replacing extracellular sodium with lithium. 5. In the period with reliable propagation cadmium, apamin, and muscarine reduced the amplitude of the AHP. The shoulder of the action potential was more pronounced and not sensitive to repetitive stimulation when extracellular calcium was replaced by strontium. It disappeared when cadmium was applied. 6. In DRG somata changes of the intracellular Ca2+ concentration were monitored by measuring the fluorescence of the Ca2+ indicator Fluo-3 with a laser-scanning confocal microscope. During repetitive stimulation, an accumulation of intracellular calcium occurred that recovered very slowly (tens of seconds) after the AP trains. 7. Computer model simulations performed in analogy to the experimental protocols produced conduction failures during repetitive stimulation only when the calcium currents during the APs were reduced. 8. From these findings it is concluded that conduction failures during repetitive stimulation are dependent on an accumulation of intracellular calcium leading to an inactivation of calcium currents, combined with small contributions of an accumulation of extracellular potassium and a summation of slow potassium conductances.

Acetylcholine

An intervention to improve the reliability of manuscript reviews for the Journal of the American Academy of Child and Adolescent Psychiatry.

OBJECTIVE: The effects of methods used to improve the interrater reliability of reviewers' ratings of manuscripts submitted to the Journal of the American Academy of Child and Adolescent Psychiatry were studied. METHOD: Reviewers' ratings of consecutive manuscripts submitted over approximately 1 year were first analyzed; 296 pairs of ratings were studied. Intraclass correlations and confidence intervals for the correlations were computed for the two main ratings by which reviewers quantified the quality of the article: a 1-10 overall quality rating and a recommendation for acceptance or rejection with four possibilities along that continuum. Modifications were then introduced, including a multi-item rating scale and two training manuals to accompany it. Over the next year, 272 more articles were rated, and reliabilities were computed for the new scale and for the scales previously used. RESULTS: The intraclass correlation of the most reliable rating before the intervention was 0.27; the reliability of the new rating procedure was 0.43. The difference between these two was significant. The reliability for the new rating scale was in the fair to good range, and it became even better when the ratings of the two reviewers were averaged and the reliability stepped up by the Spearman-Brown formula. The new rating scale had excellent internal consistency and correlated highly with other quality ratings. CONCLUSIONS: The data confirm that the reliability of ratings of scientific articles may be improved by increasing the number of rating scale points, eliciting ratings of separate, concrete items rather than a global judgment, using training manuals, and averaging the scores of multiple reviewers.

Algorithms

A reliability study of the universal goniometer, fluid goniometer, and electrogoniometer for the measurement of ankle dorsiflexion.

This study investigated the reliability of three goniometers, the universal, fluid, and electro-goniometers, in the measurement of ankle dorsiflexion. Intra- and interobserver reliability were assessed using 10 healthy volunteers and five observers. A standardized ankle position was used to measure full range of active dorsiflexion. Intraobserver reliability was assessed using one observer over two successive occasions. Interobserver reliability was assessed among five observers over five separate occasions. A one-factor analysis of variance to examine intraobserver reliability demonstrated no significant difference between each of the devices on the two occasions. A multifactorial analysis of variance demonstrated significant differences among observers and again among devices (P < 0.001). Secondary analysis for interdevice reliability demonstrated significant differences among the three devices (p < 0.1). The study suggests that each device cannot be used reliably among observers or be used interchangeably, and clinical judgment based on angular changes of less than 10 degrees are invalid if rigid protocols are not followed.

Adult

Assessing reliability of a measure of self-rated health.

The test-retest reliability of self-rated health is analysed and compared with the reliability of health questions phrased more as well as less precisely. Differences in reliability between men and women and between age groups are also assessed. The study is based on 204 and 409 re-interviews from the 1991 Swedish Level of Living Survey and the 1989 Survey of Living Conditions respectively. The results show that the reliability of self-rated health is as good as or even better than that of most of the more specific questions. Only an indicator of high blood pressure showed significantly higher reliability. The reliability of self-rated health is good in all subgroups studied, and is even excellent among older men. It is concluded that the good overall reliability of self-rated health found in this study is in line with previous results concerning the validity of people's assessments of their general health as well as results concerning the basis upon which they make these judgements.

Adult

Interrater reliability of Alzheimer's disease diagnosis.

To determine interrater reliability of dementia diagnosis, 4 physicians experienced in the evaluation of dementia patients applied 3 sets of diagnostic criteria to each of 62 patients, based on a standardized set of medical record information. All patients had undergone similar examinations and follow-up to establish the initial clinical diagnosis (76% had autopsy). Raters were blind to the diagnosis and to follow-up information after the initial evaluation period. This paper presents interrater agreement (kappa values) for a diagnosis of Alzheimer's disease using the American Psychiatric Association diagnostic criteria from the Diagnostic and Statistical Manual (DSM-III), the National Institute of Neurological and Communicative Disorders and Stroke (NINCDS) criteria for the clinical diagnosis of Alzheimer's disease, and the Eisdorfer and Cohen Research Diagnostic Criteria (ECRDC) for primary neuronal degeneration. The NINCDS showed somewhat higher average interrater reliability (kappa = 0.64) than the DSM-III (kappa = 0.55) and considerably higher interrater reliability than the ECRDC (kappa = 0.37). One rater displayed conspicuously lower levels of interrater reliability than the other 3, especially in DSM-III and ECRDC. This study indicates that interrater reliability of DSM-III and NINCDS criteria are comparable. Documentation of interrater reliability and, if necessary, training to improve reliability is an important consideration in research where different observers are diagnosing dementing illnesses.

Aged

Reliability of the NINDS Myotatic Reflex Scale.

The assessment of deep tendon reflexes is useful for localization and diagnosis of neurologic disorders, but only a few studies have evaluated their reliability. We assessed the reliability of four neurologists, instructed in two different countries, in using the National Institute of Neurological Disorders and Stroke (NINDS) Myotatic Reflex Scale. To evaluate the role of training in using the scale, the neurologists randomly and blindly evaluated a total of 80 patients, 40 before and 40 after a training session. Inter- and intraobserver reliability were measured with kappa statistics. Our results showed substantial to near-perfect intraobserver reliability, and moderate-to-substantial interobserver reliability of the NINDS Myotatic Reflex Scale. The reproducibility was better for reflexes in the lower than in the upper extremities. Neither educational background nor the training session influenced the reliability of our results. The NINDS Myotatic Reflex Scale has sufficient reliability to be adopted as a universal scale.

Adult

Validity and reliability of self-reported drinking behavior: dealing with the problem of response bias.

This work assesses the validity and reliability of self-reported survey data on drinking behavior. There is evidence to suggest that data are adversely affected by bias from underreporting. This bias affects the validity of measures of consumption of alcohol and can have deleterious effects on the results of some forms of statistical estimation. Data for this study were collected at an isolated military base. The remoteness of this site and the fact that it is a military station made it possible to estimate the actual level of consumption of alcohol for the population by assessing apparent consumption through officially recorded sales of alcohol. The results of eight measures of consumption of alcohol were compared with apparent consumption, as established by documented sales, and the validity and reliability of the various measures were determined using the classical correlational approach. The validity and reliability of the data generated by the self-report survey were also analyzed using LISREL, the measurement model in particular. The results indicate that various instruments used to assess the consumption of alcohol produce very different outcomes in terms of their validity and reliability, some questions being considerably more valid and reliable than others. Two of the more salient characteristics of questions that affect validity and reliability were isolated, namely a question's ability to aid recall and its ability to mitigate the effects of persons providing socially desirable responses. The LISREL results show that these are two underlying factors for the measurement of the consumption of alcohol. It is concluded that questions that produce valid and reliable responses do so for identifiable reasons, and measurement instruments can be improved by incorporating particular features.

Alcohol Drinking

Two measurement techniques for assessing subtalar joint position: a reliability study.

Proper assessment of the subtalar joint is critical for foot and ankle evaluation. Yet, reliability of open kinetic chain goniometric measurements of the subtalar joint has been poor. Two alternative techniques, navicular height and calcaneal position with an inclinometer, have been reported in the literature but lack reliability assessment. The purpose of this study was to determine the intertester and intratester reliability of navicular height and calcaneal position using an inclinometer. Thirty healthy, volunteer subjects (22 females, age 24 +/- 3.6 years; eight males, age 25 +/- 5.1 years) participated in this study. Two testers performed repeated measures on both feet of each subject (N = 60) during two testing sessions. Testers determined the 1) subtalar neutral position, 2) resting position, and 3) difference between these two measurements using an inclinometer for calcaneal position and navicular height. Intratester and intertester reliabilities (ICC 2, 1), standard errors of measurement, and 95% confidence intervals were determined. Intertester and intratester reliability for calcaneal position ranged from .68 to .91 for all measurements. Intertester and intratester reliability for navicular height ranged from .73 to .96 for all measurements. We conclude that these weight-bearing measurement techniques are reliable and acceptable for clinical and research purposes as measured. In addition, we hypothesize that these measurement techniques are simpler than previously described open kinetic chain methods.

Adult

The influence of experience on the reliability of goniometric and visual measurement of forefoot position.

Goniometric measurement of forefoot position relative to the rearfoot is a routine procedure used by rehabilitation specialists. This measurement is also frequently made by visual estimation. The influence of tester experience on the reliability of these two techniques at the forefoot is unknown. The purpose of this investigation was to directly examine the reliability of goniometric and visual estimation of forefoot position measurements when experienced and inexperienced testers perform the evaluation. Two clinicians (> or = 10 years experience) and two physical therapy students were recruited as testers. Ten subjects (20-31 years old), free from pathology, were measured. Each foot was evaluated twice with the goniometer and twice with visual estimation by each tester. Intraclass correlation coefficient (ICC) and coefficients of variation method error were used as estimates of reliability. There was no dramatic difference in the intratester or intertester reliability between experienced and inexperienced testers, regardless of the evaluation used. Estimates of intratester reliability (ICC 2,1), when using the goniometer, ranged from 0.08 to 0.78 for the experienced examiners and from 0.16 to 0.65 for the inexperienced examiners. When using visual estimation, ICC (2,1) values ranged from 0.51 to 0.76 for the experienced examiners and 0.53 to 0.57 for the inexperienced examiners. The estimate of intertester reliability [ICC (2,2)] for the goniometer was 0.38 for the experienced examiners and 0.42 for the inexperienced examiners. When using visual estimation, ICC (2,2) values were 0.81 for the experienced examiners and 0.72 for the inexperienced examiners. Although experience does not appear to influence forefoot position measurements, of the two evaluation techniques, visual estimation may be the more reliable.

Adult