PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Reliability of the TIP and DIP speech-hearing tests for children.

The reliability of SRT and speech intelligibility tests has been studied on adults. Reliability estimates for SRT are between .60 and .90, with standard error estimates from 1.5 to over five dB. For speech intelligibility tests the reliability estimates range from .50 to 90, with standard error of estimates from 2.5% to over 10%. Little has been reported on test reliability with children. For this study the TIP and DIP tests, for threshold and discrimination, respectively, were given to 295 normal and 138 hypacusic children three through twelve years of age. Subjects were retested within one week. TIP test-retest reliability was .72 for normals, and .89 to .99 for hypacusics. DIP test-retest reliability was .46 to .51 for normals and .60 to .93 for hypacusics. Standard error of estimate was about 3 dB for TIP, and 10% for DIP. These values are about the same as the reliability values for adults.

Child↗

Assessing clinical signs of temporomandibular disorders: reliability of clinical examiners.

Data on interrater reliability in assessing a number of clinical signs commonly evaluated in the diagnosis and treatment of temporomandibular disorders (TMD) is presented in this article. Four experienced dental hygienists who were field examiners for a large epidemiologic study of TMD and three experienced clinical TMD specialists (dentists) who are coinvestigators in the same study followed carefully detailed specifications and criteria for examination of TMD patients and pain-free controls. Excellent reliability was found for vertical range of motion measures and for summary indices measuring the overall presence of a clinical sign that could arise from several sources (for example, summary indices of muscle palpation pain). However, many clinical signs important in the differential diagnosis of subtypes of TMD were not measured with high reliability. In particular, assessment of pain in response to muscle palpation and identification of specific temporomandibular joint sounds seemed to be possible only with modest, sometimes marginal, reliability. These modest reliabilities could arise from examiner error because the clinical signs are themselves unreliable, changing spontaneously over time and making it difficult to find the same sign on successive examinations. The finding that, without calibration, experienced clinicians showed low reliability with other clinicians suggests the importance of establishing reliable clinical standards for the examination and diagnostic classification of TMD.

Adult↗

PRES- and orthostatic-induced heart-rate changes as markers of labile hypertension: magnitude and reliability measures.

Split-half and test-retest reliabilities of heart-rate responses to a baroreceptor manipulation and an orthostatic maneuver were compared between subjects with either normal or elevated blood-pressure. Ten subjects showing elevated resting blood-pressure and II normotensive subjects participated in two experimental sessions, each including heart-rate recordings during baroreceptor manipulation and orthostatic challenge. Carotid baroreceptors were manipulated by applying the baroreceptor-specific phase-related external suction (PRES) technique. The orthostatic stimulation procedure (OSP) was a change of body position from lying to standing. Heart rate responses evoked by OSP failed to discriminate significantly between the groups either in the magnitude or the (test, retest) reliability measure. The PRES procedure also failed to discriminate with the conventional magnitude measure, but the reliability measures showed significant differences. Paradoxically, the high-blood-pressure group manifested the higher baroreceptor reliability. The present findings are consistent with the view that operant conditioning produces phasic blood-pressure increases. In this view, blood-pressure increases activate the arterial baroreceptors which, in turn, dampen pain and/or stress sensitivity. Individuals showing high consistency (reliability) in their cardiovascular responses are more likely to learn this form of conditioning, and hence to eventually increase their tonic blood-pressure. High reliability of cardiovascular responses may therefore constitute a risk for hypertension. Aside from such theoretical considerations, the findings indicate that less conventional dependent variables like reliability may be worth exploring in the search for the etiology of essential hypertension, and that, in this search, specificity (relative to baroreceptor function) is more important than the magnitude of the heart-rate changes that are produced.

Adult↗

Reliability of in vivo volume measures of hippocampus and other brain structures using MRI.

Volume reductions of the hippocampus are associated with Alzheimer's disease, schizophrenia, and epilepsy. We used clinically available MRI methods (2D acquisition; inversion recovery and calculated T2 images; 3 mm contiguous slices) that optimize image contrast, quality, and resolution and standardized positioning protocols to maximize the in vivo accuracy (test-retest reliability) of brain volume measurements in volunteers who were scanned two or three times. Volunteers were scanned in the same MRI instrument (intrascanner reliability) as well as in two different instruments (interscanner reliability). A single rater obtained brain volume measures of seven contiguous slices centered on the anterior commissure. The in vivo intrascanner reliability for measures of anterior hippocampus and ventricular volumes was very good, with reliability coefficients [intraclass r (rxx)] ranging between .855 and .997, and a median coefficient of variation (CV) of 6.4%. Reliability was good for amygdala (rxx of .740 and .764) and for total frontal and temporal lobe volumes and white matter volume measures (rxx ranging between .640 and .823, median coefficient of variation was 3.2%). Overall, interscanner reliability was also good. We discuss the implications of our results relative to the possible clinical utility of hippocampal quantification and the feasibility of prospective studies aimed at quantifying progressive neurodegeneration.

Adult↗

The Irvine-Minnesota inventory to measure built environments: reliability tests.

BACKGROUND: Inter-rater reliability is an important element of environmental audit tools. This paper presents results of reliability tests of the Irvine-Minnesota Inventory, an extensive audit tool aimed at measuring a broad range of built environment features that may be linked to active living. METHODS: Inter-rater reliability was measured by percentage agreement between observers. Reliability was tested on a broad range of sites in both California and Minnesota. RESULTS: For the variables that remained in the inventory, in tests conducted at the University of California-Irvine, 76.8% of the variables had >80% agreement among the three raters. In tests conducted at the University of Minnesota, 99.2% of the variables had >80% agreement among the two raters. CONCLUSIONS: Reliability was high for most items. The inventory was modified to eliminate items with low reliability. Differences in the use of the inventory and the goals of the research led to generally higher reliability in Minnesota. Those differences, limitations, and directions for future research are discussed.

California↗

Alarm mistrust in automobiles: how collision alarm reliability affects driving.

As roadways become more congested, there is greater potential for automobile accidents and incidents. To improve roadway safety, automobile manufacturers are now designing and incorporating collision avoidance warning systems; yet, there has been little investigation of how the reliability of alarm signals might impact driver performance. We measured driving and alarm reaction performances following alarms of various reliability levels. In Experiment One, 70 participants operated a driving simulator while being presented console emitted collision alarms that were 50%, 75%, or 100% reliable. In Experiment Two, the same participants were presented spatially generated collision alarms of the same reliability levels. The results were similar in both experiments: alarm and automobile swerving reactions were significantly better when alarms were more reliable; however, drivers still failed to avoid collisions following reliable alarms. These results emphasize that alarm designers should maximize alarm reliability while minimizing alarm invasiveness.

Accidents, Traffic↗

The reliability of a simplified water displacement instrument: a method for measuring arm volume.

OBJECTIVES: To present a new water displacement measurement, the Simplified Water Displacement Instrument (SWDI), and to evaluate its intra- and intertester reliability. DESIGN: Reliability design. SETTING: Hospital setting. PARTICIPANTS: Fifty-six healthy people were studied. Intratester reliability was evaluated once a week for 4 weeks in 20 women and 10 men. Intertester reliability was assessed by 2 physical therapists in 26 people. INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Coefficients of variation (CVs) and intraclass correlation coefficients (ICCs). RESULTS: The intratester reliability showed a CV range of 2.2% to 2.6% and an ICC range of .98 to .99. The intertester reliability showed a CV of 1.3% and an ICC of .99. There was a significant increase in arm volume in men compared with women. There were no significant differences in changes in volume over the 4 weeks. There was a significant greater right arm volume (3.3%) among the right-handed subjects (P<.001). CONCLUSIONS: Both intra- and intertester reliability were satisfactory for the SWDI.

Aged↗

Measuring standing hindfoot alignment: reliability of goniometric and visual measurements.

OBJECTIVE: To examine the reliability of a functional weight-bearing measure of hindfoot alignment, the standing tibiocalcaneal angle (STCA), and to compare the relative reliabilities of goniometrically and visually estimated STCAs. DESIGN: Prospective blinded comparison. SETTING: Sports medicine center. PARTICIPANTS: Eighteen asymptomatic volunteer subjects (10 men, 8 women; age range, 22-41y). INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: Two experienced examiners completed 2 blinded goniometric STCA and 2 blinded visual STCA measurements on each subject's right and left ankles in random order. RESULTS: Quantitative visual and goniometric STCAs were similar (visual mean range, 5.61 degrees -6.50 degrees valgus vs goniometric mean range, 5.50 degrees -6.94 degrees valgus), and both measurements exhibited good to excellent intrarater reliabilities (intraclass correlation coefficient [ICC] range, .80-.94; 95% prediction limits, 1.51 degrees -2.06 degrees ). Interrater ICCs were only fair for both measurement methods (.50-.75; 95% prediction limits, 2.2 degrees -4.1 degrees ). In terms of relative reliability, the visual STCA and goniometric STCA exhibited good to excellent agreement (ICC range, .64-.95). CONCLUSIONS: The STCA as described herein exhibited acceptable intrarater reliability for clinical use but may not be acceptably reliable between experienced examiners. The visual and goniometric STCA measurements were quantitatively similar and exhibited similar reliability. Using either method, changes of up to 2 degrees over time may be attributable to measurement error. Clinicians may consider using either STCA measurement in evaluating patients with lower-limb injuries or during screening of high-risk populations.

Adult↗

Test-retest reliability of isokinetic dynamometry for the assessment of spasticity of the knee flexors and knee extensors in children with cerebral palsy.

OBJECTIVE: To assess test-retest reliability of the peak resistance torque and slope of work methods of spasticity measurement of the knee flexors and extensors in children with cerebral palsy (CP). DESIGN: Test-retest reliability study. SETTING: Pediatric orthopedic hospital. PARTICIPANTS: Fifteen children with CP. INTERVENTION: Knee extensor and flexor spasticity was assessed with an isokinetic dynamometer using passive movements at 15 degrees, 90 degrees, and 180 degrees/s taken 1 hour apart. MAIN OUTCOME MEASURES: Peak resistive torque and work were calculated. The relative and absolute test-retest reliability was calculated by using intraclass correlation coefficients (ICCs) and Bland-Altman plots, respectively. RESULTS: Relative reliability was good (ICC>.75) for slope-of-work and peak resistance torque measurements at a velocity of 180 degrees/s, whereas reliability of peak torque measurements was decreased (ICC<.51) at slower velocities for both muscle groups. The 95% limits of agreement of Bland-Altman plots contained most data points for both methods, but the width of the limits of agreement were wide. CONCLUSIONS: The measurement of spasticity of the knee extensors and flexors in children with CP using peak-resistance torque at 180 degrees/s and the slope of work method has acceptable relative test-retest reliability. However, the absolute reliability of spasticity data should be considered cautiously.

Adolescent↗

Increased reliability for single-case research results: is the bootstrap the answer?

There is need for objective and reliable single-case research (SCR) results in the movement toward evidence-based interventions (EBI), for inclusion in meta-analyses, and for funding accountability in clinical contexts. Yet SCR deals with data that often do not conform to parametric data assumptions and that yield results of low reliability. A resampling technique, the bootstrap, largely bypasses statistical assumptions and usually yields more reliable results. This study answers questions about the extent of need for the bootstrap in SCR and its impact on effect size reliability. The bootstrap was applied in Allison et al. mean shift analyses (Faith, Allison, & Gorman, 1997) to data from 166 published AB graphs. Results showed the bootstrap improved reliability of 88% of the analyses and reduced reliability of only 3%. The reliability improvement was large enough to be practically useful. The bootstrap was paired with a method for cleansing data of autocorrelation, which also proved effective. Pending replication, the findings encourage broad application within SCR of both the bootstrap and autocorrelation cleansing.

Case-Control Studies↗

Korean version of the diagnostic interview for genetic studies: Validity and reliability.

The Diagnostic Interview for Genetic Studies (DIGS), developed in 1994 by the National Institute of Mental Health (NIMH), was translated into Korean and tested for reliability and diagnostic validity. Concurrent validity was tested using the Structured Clinical Interview for DSM-IV (SCID) and clinical diagnoses in 53 patients, most of whom had either schizophrenia or bipolar disorder. Inter-rater reliability was tested in 24 patients. Test-retest reliability was also tested in 17 patients. Overall and specific diagnostic validity for the Korean version of DIGS (DIGS-K) was excellent for most diagnoses. Inter-rater and test-retest reliability for overall and specific diagnoses also ranged from fair to excellent. For schizoaffective disorder, the test-retest reliability of DIGS-K was in a fair range, although the level was lower than that of other diagnoses. However, its diagnostic validity and inter-rater reliability was below fair range. In conclusion, DIGS-K appears to be a reliable interview for major psychiatric disorders.

Genetic Testing↗

Diagnostic reliability of the Semi-structured Assessment for Drug Dependence and Alcoholism (SSADDA).

UNLABELLED: The Semi-structured Assessment for Drug Dependence and Alcoholism (SSADDA) is a diagnostic instrument developed for studies of the genetics of substance use and associated disorders. The SSADDA provides more detailed coverage of specific drug use disorders, particularly cocaine and opioid dependence, than existing psychiatric diagnostic instruments. A computerized version of the SSADDA was developed to permit direct entry of subject responses by the interviewer. This study examines the diagnostic reliability of the SSADDA for substance use disorders and for other DSM-IV disorders that are commonly associated with substance use disorders. METHODS: Two hundred and ninety-three subjects (mean age = 39 yr, 52.2% women) were interviewed twice over a 2-week period in two sub-studies examining the inter-rater (n = 173) or test-retest reliability (n = 120) of the SSADDA. The kappa statistic and Yule's Y were used to measure reliability. RESULTS: The reliability of most substance dependence diagnoses was good to excellent, although the reliability of substance abuse diagnoses was substantially lower. The reliability of the associated psychiatric diagnoses varied from fair to excellent. CONCLUSIONS: The SSADDA yields reliable diagnoses for a variety of psychiatric disorders, including alcohol and drug dependence. Although developed for use in genetic studies, its broad and detailed coverage of disorders and computer-assisted format will allow it to be used in a variety of applications requiring careful diagnostic assessment.

Adult↗

The reliability and reproducibility of foot type measurements using a mirrored foot photo box and digital photography compared to caliper measurements.

UNLABELLED: The height of the medial longitudinal arch (MLA) is thought to be a predisposing factor to various lower extremity injuries. Discrepancy exists as to whether MLA height plays a role in injury prevention. The purpose of this study was to determine the intertester and intratester reliability, and the validity of the mirrored foot photo box (MFPB) and caliper measurements to radiographic measurements. METHODS: Thirty subjects with equal numbers of men and women were recruited. Both feet were tested (n=60) in a 90% weight bearing stance. A set of anatomic landmarks were palpated, marked, and measured using a caliper, MFPB, and radiographs. The protocol was completed by two testers on 2 days approximately 1 week apart. Intertester and intratester reliability were determined using the intraclass correlation coefficient (ICC)(2,k) and the ICC(2,1), respectively. Validity of both measurement techniques to radiographic measurements was determined using the ICC(2,k). RESULTS: The intertester reliability ranged from 0.991 to 0.577, while the intratester reliability ranged from 0.994 to 0.527, with first metatarsal angle being the only variable with poor reliability. Most variables demonstrated acceptable validity between the MFPB and the caliper measurements, and acceptable validity between the MFPB and calipers compared to radiographic measurements. The MFPB took 51.3+/-19.6s per foot while the caliper measurements averaged 227.4+/-68.9s to complete the measurements. DISCUSSION: The MFPB is as reliable as the caliper measurements, and offers better intertester reliability. Both the caliper and MFPB measurements demonstrated acceptable validity to radiographic measurements and testing time was reduced when using the MFPB compared to calipers.

Adolescent↗

Interobserver and intraobserver reliability of therapist-assisted videotaped evaluations of upper-limb hemiplegia.

PURPOSE: Therapist-assisted videotaped sessions have been used to augment physical examinations in the evaluation of hand and arm function in patients with spastic hemiplegia. The purpose of this study was to assess the interobserver and intraobserver reliability of standardized videotaped examinations in the evaluation and functional classification of these patients. METHODS: Three examiners reviewed standardized videotaped examinations of 10 adolescents with spastic hemiplegia on 2 separate occasions. All 10 patients were under consideration for surgical intervention for their upper-extremity dysfunction. Videotapes were used to assess upper-extremity range of motion, finger and thumb deformity, and reach, pinch, and grip function. Upper-extremity function was graded according to the House and Mowery classification systems. Interobserver and intraobserver reliabilities were measured with the kappa coefficient. RESULTS: Range of motion, deformity, and upper-extremity functional strategy assessment showed slight to excellent interobserver reliability and good to almost perfect intraobserver reliability. Interobserver and intraobserver reliability of the consolidated House classification system was more reliable than the Mowery or standard House classification systems. CONCLUSIONS: Evaluations of standardized videotaped examinations in patients with hemiplegia were reliable between and among observers. Such therapist-assisted videotaped evaluations may provide useful data for clinical decision-making and multicenter outcomes studies in patients with upper-extremity involvement with spastic hemiplegia.

Adolescent↗

Subjective and Objective Numerical Outcome Measure Assessment (SONOMA). A combined outcome measure tool: findings on a study of reliability.

OBJECTIVE: To determine the reliability of a combined tool, namely that of Subjective and Objective Numerical Outcome Measure Assessment (SONOMA). METHODS: Testing was conducted, limited to patients with neck, midback, or lower back pain, with or without radiculopathy, in an outpatient chiropractic office setting. Test-retest reliability of the objective analysis of SONOMA was carried out on the same day (n = 50) with an interval time period of less than 60 minutes. Between-day reliability of the subjective analysis of SONOMA was carried out with an interval time period of 24 hours (n = 50). Individual and combined parameter reliability was established for the tool. RESULTS: Short-term objective and between-day subjective reliability coefficients were high. The Pearson correlation coefficient for the combined tool was .96, and the coefficient for the individual parameters ranged from .55 through .93. All these correlations were statistically significant, with a P value not more than .0001. CONCLUSIONS: The SONOMA combined numerical outcome measure tool demonstrated a high degree of reliability. This outcome tool measures directly and therefore reflects patient pain perception, functional status, and provider-driven objective assessment. We believe this tool provides the unique combination of both subjective and objective functional capacity assessment. It should be valuable for day-to-day practical application, as well as considered for future clinical trials and quality-of-care studies. This combined tool shows promise as having a high degree of reliability and, hence, may demonstrate a comprehensive representation of the patient-clinical picture, particularly in regard to functional capacity assessment.

Activities of Daily Living↗

Assessment of neonatal resuscitation skills: a reliable and valid scoring system.

OBJECTIVE: To study the reliability and validity of a scoring instrument for the assessment of neonatal resuscitation skills in a training setting. METHODS: Fourteen paediatric residents performed a neonatal resuscitation on a manikin, while being recorded with a video camera. The videotapes were analysed using an existing scoring instrument with an established face and content validity, adjusted for use in a training setting. Intra- and inter-rater reliability were assessed by comparing the ratings of the videotapes of three raters, one of who rated the videotapes twice. Intra-class coefficients (ICC) were calculated for the sum score, percentages of agreement and kappa coefficients for the individual items. To study construct validity, the performance of a second resuscitation of by residents was assessed after they had received feedback on their first performance. RESULTS: The ICC were 0.95 and 0.77 for intra- and inter-rater reliability, respectively. The median percentage of intra-rater agreement was 100%; inter-rater agreement 78.6-84.0%. The median kappa was 0.85 for intra-rater reliability, and 0.42-0.59 for inter-rater reliability. Residents showed a 10% improvement (95% confidence interval -4; 23%) on performance of a second resuscitation, which supports the instrument's construct validity. CONCLUSION: A useful and valid instrument with good intra-rater and reasonable inter-rater reliability is now available for the assessment of neonatal resuscitation skills in a training setting. Its reliability can be improved by using a more advanced manikin and by training of the raters.

Cardiopulmonary Resuscitation↗

Reliability of self-reported smoking history and age at initial tobacco use.

BACKGROUND: Many studies use questionnaires to determine smoking status and age of smoking onset. This study aimed to determine the reliability of self-reported smoking history and age of smoking initiation. METHOD: The proportion of inconsistent answers and correlation coefficients of reported age of initial smoking were measured by an answer-reanswer analysis of questionnaires in an ongoing, two-step, population-based survey of health behavior. Interviews were conducted on the day of recruitment to and the day of discharge from mandatory military service in Israel among a sample of 25,437 young men and women recruited between 1986 and 2000. RESULTS: Of 7276 participants reporting current or past smoking upon recruitment, 559 (7.7%) reported never having smoked upon discharge, thus demonstrating prima facie inconsistency. Variables significantly associated with reliable reporting in a multivariate logistic regression model were female gender (P = 0.04) and more than 4 years of military service (P < 0.01). 6010 subjects who reported a positive smoking history at both recruitment and discharge were available for analysis of reliability of reported age at smoking onset. Intraclass correlation coefficients for recruitment/discharge consistency in reported age at first cigarette were 0.73 (95% CI: 0.71-0.74) and 0.76 (95% CI: 0.74-0.78) for men and women, respectively. Eastern origin, lower subject education level, and lower paternal education level were also associated with lower reliability. CONCLUSIONS: Our results showed a relatively high level of answer-reanswer reliability, with some variance attributable to personal characteristics. These results suggest that self-reported age at onset of tobacco use is practical and reliable in normative, young adult populations. However, time elapsed between questionnaires and demographic and lifestyle characteristics may affect reliability rates, and thus should be carefully regarded in future studies.

Adolescent↗

Reliability of the dynamic gait index in people with vestibular disorders.

OBJECTIVE: To examine the interrater reliability of the Dynamic Gait Index (DGI) when used with patients with vestibular disorders and with previously published instructions. DESIGN: Correlational study. SETTING: Outpatient physical therapy clinic. PARTICIPANTS: Subjects included 30 patients (age range, 27-88y) with vestibular disorders, who were referred for vestibular rehabilitation. INTERVENTIONS: Subjects' performance on the DGI was concurrently rated by 2 physical therapists experienced in vestibular rehabilitation to determine interrater reliability. MAIN OUTCOME MEASURES: Percentage agreement, kappa statistics, and the ratio of subject variability to total variability were calculated for individual DGI items. Kappa statistics for individual items were averaged to yield a composite kappa score of the DGI. Total DGI scores were evaluated for interrater reliability by using the Spearman rank-order correlation coefficient. RESULTS: Interrater reliability of individual DGI items varied from poor to excellent based on kappa values (kappa range,.35-1.00). Composite kappa values showed good overall interrater reliability (kappa=.64) of total DGI scores. The Spearman rho demonstrated excellent correlation (r=.95) between total DGI scores given concurrently by the 2 raters. CONCLUSION: DGI total scores, administered by using the published instructions, showed moderate interrater reliability with subjects with vestibular disorders. The DGI should be used with caution in this population at this time, because of the lack of strong reliability.

Adult↗