PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Intratester and intertester reliability of the Cybex electronic digital inclinometer (EDI-320) for measurement of active neck flexion and extension in healthy subjects.

This study examined the intratester and intertester reliability of the electronic digital goniometer EDI-320 for the measurement of active neck flexion and extension in healthy subjects. In the context of evidence-based practice, the EDI-320 instrument has the potential to improve patient assessment, provide a clearer picture of patient progress, and confirm the effectiveness of physiotherapy interventions. However, the psychometric properties of the EDI-320 have not yet been documented for cervical spine range of motion. Forty-four individuals with no known history of cervical disorder within the three months prior to the testing, voluntarily consented to participate in this study. Repeated measurements with the EDI-320 were taken by two trained testers (TH1 and TH2) and data were recorded by two separate observers. Subjects performed a standardized warm-up. Testers were required to repeat palpation of bony landmarks prior to each trial. Measurements were taken at the end-range of active cervical flexion and extension for each subject. Both testers measured each subject twice. The intraclass correlation coefficients (ICC) were derived from one-way ANOVA for intratester reliability and a two-way ANOVA for intertester reliability. Paired t -tests were then applied to verify for systematic error. Moderate intratester reliability was found for both testers for flexion (TH1: ICC=0.77; 95% CI: 0.62-0.87; TH2: ICC=0.77; 95% CI: 0.58-0.87). As for extension, high intratester reliability was found for TH1 (ICC=0.79; 95% CI: 0.65-0.88) and moderate for TH2: (ICC=0.83; 95% CI: 0.63-0.92). Intertester reliability results showed a moderate reliability for both flexion and extension (ICC=0.66; 95% CI: 0.24-0.84) on the first trial. On the second trial, reliability was moderate for flexion (ICC=0.73; 95% CI: 0.53-0.85) and high for extension (ICC=0.80; 95% CI: 0.64-0.89). The t -test analysis revealed the inclusion of systematic error by Tester 2 for intratester reliability. This error was also found for all but one of the intertester reliability calculations. This study has shown that the EDI-320 is a moderately reliable instrument for quantifying cervical flexion and extension range of motion. The presence of systematic error in the study highlights the importance of following standardized procedures and suggests that the EDI-320 could be more reliable than reported in this study. Further psychometric studies investigating the validity of the EDI and reliability with subjects affected by cervical pathology is warranted.

Adult↗

Reliability of the Modified Ashworth Scale in the assessment of plantarflexor muscle spasticity in patients with traumatic brain injury.

Although the Modified Ashworth Scale (MAS) is commonly used to assess the severity of muscle spasticity for ankle plantarflexors, its reliability has only been established for elbow muscles. Interrater reliability, intrarater reliability and temporal (between-days) reliability were examined in this study. Also, interrater reliability for use of the scale with plantarflexors was compared with reported results from the measurement of elbow flexors. Thirty adult volunteers with traumatic brain injuries participated. There were 20 men and 10 women; the mean age was 28.3 years (SD = 10.8). Two physical therapists used the MAS to score the subjects independently. Measurements were repeated to yield multiple scores for intrarater reliability assessment. Twenty-one of the subjects returned individually on separate days to be measured again, so that temporal reliability could be assessed. Spearman's correlation coefficients were 0.73 for interrater reliability 0.74 and 0.55 for intrarater reliability, and 0.82 for temporal reliability. Overall, reliability of the MAS for assessing plantarflexor spasticity in patients with traumatic brain injury was found to be minimally adequate to support its continued use. However, interrater reliability was less than that which has been reported for elbow flexors, and intrarater reliability findings were mixed.

Adolescent↗

Inter-rater and test-retest reliability of three contingent valuation question formats in south-east Nigeria.

This paper examines the inter-rater and test-retest reliability of willingness to pay (WTP) for insecticide-treated mosquito nets and net re-treatment using the bidding game (BG), binary with follow-up (BWFU) and a novel structured haggling technique (SH). Inter-rater reliability was evaluated by having two sets of interviewers administer questionnaires to 109 (BG), 110 (BWFU) and 103 (SH) randomly selected household heads. Test-retest reliability was investigated by repeating interviews on 146 (BG), 161 (BWFU) and 139 (SH) household heads one month after an initial survey. Data analysis used testing of means, Spearman's correlation and Pearson's correlation coefficient for test of reliability, while non-parametric analysis was used to determine factors causing a variation in WTP. The study was conducted in Southeast Nigeria. Inter-rater reliability coefficients were estimated for the individual's WTP for own nets, WTP for others and WTP for re-treatment. Using WTP for own nets as the best reliability estimate, the coefficients were high at values of 0.77 (C.I. 0.72-0.86), 0.75 (C.I. 0.64-0.81) and 0.74 (C.I. 0.63-0.82) in the BG, BWFU and SH, respectively. In test-retest reliability coefficients, the coefficients for WTP for own nets were low-to-moderate at values of 0.51 (C.I. 0.40-0.62), 0.41 (C.I. 0.28-0.53) and 0.56 (C.I. 0.41-0.65) for the BG, BWFU and SH groups, respectively. Factors such as gender, change in income, unplanned expenditures, stated WTP in first survey, time-to-think, external information, and subjecting respondents to more than one interview explained the lower test-retest reliability coefficients. We conclude that the CVM was reliable in the study area and the question formats had similar levels of reliability. The lower coefficients in the test-retest reliability were due to the influence of factors affecting demand that had changed in the intervening period. Standard formats for determining reliability within CVM should be developed for easy comparison of results from different studies.

Animals↗

AO or Schatzker? How reliable is classification of tibial plateau fractures?

INTRODUCTION: We compare the intra- and interobserver reproducibility of classifications of tibial plateau fractures most commonly used in our clinical practice. These were the AO and Schatzker classifications. PATIENTS AND METHODS: Agreement was measured using kappa coefficients on the data obtained from three observers reviewing 30 fractures and these values were interpreted according to Landis and Koch. RESULTS: It was found that both classifications were substantially reliable with regards to intraobserver reliability but that the Schatzker system was only fairly reliable and the AO classification moderately reliable with regards to interobserver reliability. Breaking down the AO classification, with regards to intraobserver reliability, the AO group was substantially reliable and the type excellently reliable. For interobserver reliability, the AO group was moderately reliable while the AO type was substantially reliable. CONCLUSION: For tibial plateau fractures seen on plain x-ray, the AO classification is more reliable between observers than the Schatzker classification.

Humans↗

Discussion between reviewers does not improve reliability of peer review of hospital quality.

OBJECTIVES: Peer review is used to make final judgments about quality of care in many quality assurance activities. To overcome the low reliability of peer review, discussion between several reviewers is often recommended to point out overlooked information or allow for reconsideration of opinions and thus improve reliability. The authors assessed the impact of discussion between 2 reviewers on the reliability of peer review. METHODS: A group of 13 board-certified physicians completed a total of 741 structured implicit record reviews of 95 records for patients who experienced severe adverse events related to laboratory abnormalities while in the hospital (hypokalemia, hyperkalemia, renal failure, hyponatremia, and digoxin toxicity). They independently assessed the degree to which each adverse event was caused by medical care and the quality of the care leading up to the adverse event. Working in pairs, they then discussed differences of opinion, clarified factual discrepancies, and rerated the record. The authors compared the reliability of each measure before and after discussion, and between and within pairs of reviewers, using the intraclass correlation coefficient for continuous ratings and the kappa statistic for a dichotomized rating. RESULTS: The assessment of whether the laboratory abnormality was iatrogenic had a reliability of 0.46 before discussion and 0.71 after discussion between paired reviewers, indicating considerably improved agreement between the members of a pair. However, across reviewer pairs, the reviewer reliability was 0.36 before discussion and 0.40 after discussion. Similarly, for the rating of overall quality of care, reliability of physician review went from 0.35 before discussion to 0.58 after discussion as assessed by pair. However, across pairs the reliability increased only from 0.14 to 0.17. Even for prediscussion ratings, reliability was substantially higher between 2 members of a pair than across pairs, suggesting that reviewers who work in pairs learn to be more consistent with each other even before discussion, but this consistency also did not improve overall reliability across pairs. CONCLUSIONS: When 2 physicians discuss a record that they are reviewing, it substantially improves the agreement between those 2 physicians. However, this improvement is illusory, as discussion does not improve the overall reliability as assessed by examining the reliability between physicians who were part of different discussions. This finding may also have implications with regard to how disagreements are resolved on consensus panels, guideline committees, and reviews of literature quality for meta-analyses.

Causality↗

Reliability and validity of the chronic respiratory questionnaire (CRQ).

BACKGROUND: The Chronic Respiratory Questionnaire (CRQ) is frequently applied to assess quality of life in patients with chronic obstructive pulmonary disease (COPD). However, the reliability and validity of this questionnaire have not yet been determined. This study investigates the reliability and validity of the four separate dimensions of the CRQ. METHODS: The CRQ was administered on two consecutive days to 40 patients with COPD (mean FEV1 44% predicted, FEV1/IVC 37% predicted). Internal consistency reliability of each dimension was investigated by Cronbach's alpha reliability coefficient, test retest reliability by the Spearman-Brown reliability coefficient (p), and content validity by Pearson's correlation coefficient between the CRQ and the symptom checklist (SCL-90). RESULTS: Items of the fatigue, emotion, and mastery dimensions showed a high internal consistency reliability (alpha = 0.71-0.88) as well as a high test retest reliability (p above 0.90). These three dimensions correlated with comparable dimensions of the SCL-90. Items of the dyspnoea dimension showed a low internal consistency reliability (alpha = 0.53) and a test retest reliability of p = 0.73. CONCLUSIONS: Items of the dimensions fatigue, emotion, and mastery of the CRQ are reliable and valid and can be used to assess quality of life in patients with severe airways obstruction. Items of the dyspnoea dimension are less reliable and should not be included in the overall score of the CRQ in comparative research. However, by scoring the items of dyspnoea separately they may be useful for the evaluation of the effects of intervention in a specific patient.

Aged↗

Reliability of biomechanical variables of sprint running.

PURPOSE: The purpose of this paper was to report the reliability of variables used in the biomechanical assessment of sprint running and to document how these reliability measures are likely to improve when using the average score of multiple trials. METHODS: Twenty-eight male athletes performed maximal-effort sprints. Video and ground reaction force data were collected at the 16-m mark. The reliability (systematic bias, random error, and retest correlation) for a single score was calculated for 26 kinematic and 7 kinetic variables. In addition, the reliability (random error and retest correlation) for the average score of 2, 3, 4, and 5 trials was predicted from the reliability of a single score. RESULTS: For all variables, there was no evidence of systematic bias. The measures of random error and retest correlation differed widely among the variables. Variables describing horizontal velocity of the body's center of mass were the most reliable, whereas variables based on vertical displacement of the body's center of mass or braking ground reaction force were the least reliable. For all variables, reliability improved notably when the average score of multiple trials was the measurement of interest. CONCLUSION: Although it is up to the researcher to judge whether a measurement is reliable enough for its intended use, some of the lower-reliability variables were possibly too unreliable to monitor small changes in an athlete's performance. Nonetheless, there was a consistent trend for reliability to improve notably when the average score of multiple trials was the measurement of interest. Subsequently, if resources permit, researchers and applied sports-scientists may like to consider using the average score of multiple trials to gain the advantages that improved reliability offers.

Adult↗

Inter-rater and test-retest reliability of the Revised Diagnostic Interview for Borderlines.

The baseline inter-rater reliability, test-retest reliability, follow-up inter-rater reliability, and follow-up longitudinal reliability (interrater reliability between generations of raters) of borderline symptoms and the diagnosis of borderline personality disorder (BPD) were assessed using the Revised Diagnostic Interview for Borderlines (DIB-R). Excellent kappa s (> .75) were found in each of these reliability substudies for the diagnosis of BPD itself. Excellent kappa s were also found in each of the three inter-rater reliability substudies for the vast majority of borderline symptoms assessed by the DIB-R. Test-retest reliability for these symptoms was somewhat lower but still very good. More specifically, one-third of the BPD symptoms assessed had a kappa in the excellent range and the remaining two-thirds had a kappa in the fair-good range (.57-.73). The dimensional reliability of BPD symptom areas was somewhat higher than for categorical measures of the subsyndromal phenomenology of BPD. More specifically, all five dimensional measures of borderline psychopathology had intraclass correlation coefficients in the excellent range for all four reliability substudies. Taken together, the results of this study suggest that both the borderline diagnosis and the symptoms of BPD can be diagnosed reliably when using the DIB-R. They also suggest that excellent reliability, once achieved, can be maintained over time for both the syndromal and subsyndromal phenomenology of BPD.

Borderline Personality Disorder↗

The stability and reliability of self-reported drinking measures.

OBJECTIVE: Estimated test-retest reliabilities of self-re- ported drinking measures are affected by the extent to which respondents provide consistent reports of their own behaviors (reliability) and the extent to which the behaviors reported are stable over time (stability). Unstable behaviors may be reliably reported but correlate poorly over time. This study tests whether an estimate of the stability of drinking patterns is related to test-retest reliabilities of drinking measures. METHOD: Data were from a general population telephone survey given twice, 1 month apart, to 307 adult drinkers. Drinking measures included age of onset, and graduated frequency measures used to estimate drinking frequencies, average quantities, and total alcohol consumption. Measures of drinking stability were estimated using a well-tested model for the analysis of drinking patterns (i.e., variances in drinking quantities and frequencies). Heteroscedastic regression models were used to partition covariances in self-reports between Times 1 and 2 into components related to stability and reliability, providing a more accurate picture of the reliability of drinking measures for different drinking groups. RESULTS: Overall test-retest reliabilities were good, ranging from a low of .65 for drinking quantities to a high of .85 for drinking frequencies. The stability of quantity measures had a large impact on estimated test- retest reliabilities. Stable drinking patterns were associated with much greater test-retest reliabilities. CONCLUSIONS: Data on alcohol use from general population telephone surveys are generally reliable. However, observed reliability is a function of the stability of drinking patterns. Ostensibly unreliable self-reports may be highly reliable but may reflect unstable drinking patterns.

Adolescent↗

Peer review of the quality of care. Reliability and sources of variability for outcome and process assessments.

CONTEXT: Peer assessments have traditionally been used to judge the quality of care, but a major drawback has been poor interrater reliability. OBJECTIVES: To compare the interrater reliability for outcome and process assessments in a population of frail older adults and to identify systematic sources of variability that contribute to poor reliability. SETTING: Eight sites participating in a managed care program that integrates acute and long-term care for frail older adults. PATIENTS: A total of 313 frail older adults. DESIGN: Retrospective review of the medical record with 180 charts randomly assigned to 2 geriatricians, 2 geriatric nurse practitioners, or 1 geriatrician and 1 geriatric nurse practitioner and 133 charts randomly assigned to either a geriatrician or a geriatric nurse practitioner. MAIN OUTCOME MEASURES: Interrater reliabilities for structured implicit judgments about process and outcomes for overall care and care for each of 8 tracer conditions (eg, arthritis). RESULTS: Outcome measures had higher interrater reliability than process measures. Five outcome measures achieved fair to good reliability (more than 0.40), while none of the process measures achieved reliabilities more than 0.40. Three factors contributed to poorer reliabilities for process measures: (1) an inability of reviewers to differentiate among cases with respect to the quality of management, (2) systematic bias from individual reviewers, and (3) systematic bias related to the professional training of the reviewer (ie, physician or nurse practitioner). CONCLUSIONS: Peer assessments can play an important role in characterizing the quality of care for complex patients with multiple interrelated chronic conditions, but reliability can be poor. Strategies to achieve adequate reliability for these assessments should be applied. These strategies include emphasizing outcomes measurement, providing more structured assessments to identify true differences in patient management, adjusting systematic bias resulting from the individual reviewer and their professional background, and averaging scores from multiple reviewers. Future research on the reliability of peer assessments should focus on improving the ability of process measures to differentiate among cases with respect to the quality of management and on identifying additional sources of systematic bias for both process and outcome measures. Explicit recognition of factors influencing reliability will strengthen efforts to develop sound measures for quality assurance.

Aged↗

The reliability of short-term measurements of heart rate variability.

Short-term assessment of heart rate variability (HRV) is a non-invasive technique to examine ANS function. Within the literature, HRV is commonly referred to as a reliable measurement technique. The aim of this review was to assess the accuracy of this description based upon a comprehensive review of the available data concerning reliability of short-term HRV measures. Reviewing only studies using appropriate statistical analyses, it was determined that reliability coefficients for HRV measures were highly varied. Coefficients of variation ranged from <1% to >100%. Similar variation was found in studies using the intraclass correlation coefficient values, and limits of agreement. Reliability coefficients reported displayed some distinct patterns. Firstly, where measurements were made during interventions such as tilt or pharmacological stimulation, reliability was poorer than when HRV was measured at rest. Secondly, clinical populations displayed poorer reliability than healthy subjects. There was little effect of test-retest duration on reliability and although no single HRV measurement appeared less reliable than another, there was evidence that optimal data collection conditions for specific frequency domain measures exist. Describing HRV as a reliable measurement technique appears to be a gross oversimplification, as results of reliability studies are heterogeneous, and dependent on a number of factors. Further studies are required, particularly in clinical populations to assess HRV reliability. Authors should refer to coefficients from similar populations measured under similar conditions when making future sample size calculations.

Heart Rate↗

Establishing the reliability of Mobility Milestones as an outcome measure for stroke.

UNLABELLED: Baer GD, Smith MT, Rowe PJ, Masterton L. Establishing the reliability of mobility milestones as an outcome measure for stroke. Arch Phys Med Rehabil 2003;84:977-81. OBJECTIVE: To establish intrarater, interrater, and test-retest reliability of a standardized measure of mobility, "mobility milestones," incorporating sitting balance, standing balance, and walking ability. DESIGN: Repeated-measures reliability study by using video data of patients with stroke. SETTING: Physiotherapy and rehabilitation departments in Scotland. PARTICIPANTS: Forty physiotherapists recruited from within the Lothian region: 20 senior physiotherapists with at least 3 years of experience working with neurologic patients and 20 staff grade physiotherapists with less than 12 months of experience working with neurologic patients. INTERVENTION: Videotape comprising 40 clips (36 original clips, 4 repeated clips) of stroke patients of differing levels of ability attempting the mobility milestones was produced. After a short training session in the interpretation and application of the mobility milestones, each physiotherapist viewed the tape separately and scored whether the milestone had been achieved or not. This was repeated at a separate test session 2 weeks later. MAIN OUTCOME MEASURE: Score for each mobility milestone. RESULTS: Kappa statistics were used to determine interrater reliability and showed good (.61-.80) to very good (.81-1.0) reliability for 3 of 4 milestones. Intraclass correlation coefficients (ICCs) were used to determine intrarater reliability of the 4 repeated clips and showed 75% of all subjects had high (ICC(2,1)=.91-1.0) reliability. The ICC(2,1) for test-retest reliability showed a similar pattern, with 70% of subjects showing good (.81-.90) or high (.91-1.0) reliability. CONCLUSIONS: The mobility milestones showed favorable levels of reliability when used by experienced or novice physiotherapists. The milestones can be adopted as a simple clinical outcome measure for use with stroke. Further research is required to establish reliability levels when the measure is used by different rehabilitation professionals.

Activities of Daily Living↗

Measurement of blood flow in the vertebral artery using colour duplex Doppler ultrasound: establishment of the reliability of selected parameters.

This study was designed to determine the reliability of the ultrasound testing procedure for evaluating vertebral artery blood flow, and to determine a robust testing protocol for future studies. Blood flow parameters were tested in ten asymptomatic subjects (mean age 33 years, standard deviation 6 years 8 months) using colour duplex Doppler imaging. Volume flow rate data at C5-6 demonstrated good reliability from a single measurement (Intraclass correlation coefficient [ICC]=0.81). Peak velocity sampled at C1-2 showed poor reliability if a single measurement was used (ICC=0.26) improving to fair levels with three measurements (ICC=0.77). Reliability for this parameter was good if five measurements were taken (ICC=0.83-0.84). Systolic/diastolic ratio measured at C5-6 showed poor reliability (ICC=0.57) if a single measurement was taken in the manner of Thiel et al. (1994). This improved to fair reliability (ICC=0.75) if the mean of three measurements was used. There was no further improvement if five measures were sampled. Sampling at C2-3 in the manner of Refshauge (1994) was found to be technically difficult and it was not possible to detect a Doppler shift in three of the ten subjects at this level. Reliability of peak velocity at C2-3 was found to be poor, regardless of whether single or multiple averaged measurements were taken (ICC=0.37-0.63). Mean (time averaged) velocity measurements at C2-3 showed poor reliability if a single measurement was taken (ICC=0.39), fair reliability if the first three measurements were averaged (ICC=0.73) and good to high reliability levels if five measurements were sampled (ICC=0.88-0.91). A review of the literature suggests that sampling volume flow rate at C5-6 and peak velocity at C1-2 represents a clinically meaningful combination of parameters to detect narrowing in the VA. The results of this current study indicate the desirability of taking a single measurement of volume flow rate at C5-6 and the mean of three measurements of peak velocity at C1-2, with the additional calculation of the standard error of measurement, if reliable results are to be achieved.

Adult↗

A study of the test-retest reliability of ten olfactory tests.

Ten tests of olfactory function (including tests of odor identification, detection, discrimination, memory, and suprathreshold odor intensity and pleasantness perception) were administered on two test occasions to 57 subjects ranging in age from 18 to 83 years. The stability of the average test scores was determined across the two test sessions for 14 measures derived from these 10 tests and for subcomponents of the Japanese T&T olfactometer threshold test. In addition, the test-retest reliability (Pearson r) of each test measure was established. With the exception of a response bias measure, the average test scores did not differ significantly across the two test sessions. Statistically, the reliability coefficients of the primary test measures fell into three general classes bound by the following r values: 0.43-0.53; 0.67-0.71; 0.76-0.90. Detection threshold values were more reliable than recognition threshold values; those based upon a single ascending presentation series were much less reliable than those based upon a staircase procedure. The relationship between test length and reliability was examined for several of the tests and mathematically modeled. For example, within the staircase series incorporating the odorant phenyl ethyl alcohol, reliability was related (R2 = 0.984) to the number of reversals included in the threshold estimate by a function derived from the Spearman-Brown formula; namely, reliability = 0.455* # reversals/[1 + 0.455 (# reversals - 1)]. Reversal location, per se, had little influence on reliability. Overall, this study suggests that (i) considerable variation is present in the reliability of olfactory tests, (ii) reliability is a function of test length, and (iii) caution is warranted in comparing results from nominally different olfactory tests in applied settings since the findings may, in some instances, simply reflect the differential reliability of the tests.

Adolescent↗

Interrater reliability of auscultation of breath sounds among physical therapists.

BACKGROUND AND PURPOSE: Although auscultation is routinely used in the assessment of respiratory status, the ability of the rater to accurately and consistently identify lung sounds has been questioned. The literature on this issue is sparse and has focused on reliability of auscultation of tape-recorded rather than in vivo lung sounds. The purposes of this study were to determine the interrater reliability of physical therapists in the direct auscultation of lung sounds based on their clinical experience in chest physical therapy and to determine whether the adoption of standardized nomenclature and education on proper technique and interpretation affects reliability. SUBJECTS AND METHODS: A group of 57 registered physical therapists were stratified by clinical experience into four groups. Sixteen therapists (ie, 4 in each stratum) were randomly chosen using a random number table. The following criteria were developed to delineate clinical experience. Group 1 subjects were senior chest physical therapists with at least 5 years of experience in this area of practice. Group 2 subjects were experienced therapists who had a minimum of 2 years of experience in chest physical therapy and were currently practicing in this area. Group 3 subjects were experienced physical therapists in other areas who were also practicing in chest physical therapy on occasional weekend service. Group 4 subjects were new graduates. Ten patients were evaluated by each group of 4 physical therapists using a teaching stethoscope with one diaphragm/bell and four pairs of earpieces. The education session consisted of discussion of the adoption of standardized nomenclature and education on proper technique and interpretation of auscultation. Interrater reliability was assessed before and after the education session using kappa (kappa) values. Comparisons were made between kappa values before and after the education session to determine the effect of education and between groups to determine the effect of clinical experience. RESULTS: The kappa values before the education session were low, indicating poor reliability in detecting specific abnormal sounds (kappa = -.02-.59). Group 1 (seniors in respiratory therapy) and group 4 (new graduates) demonstrated the greatest reliability levels. The lowest kappa values were observed for detecting and categorizing the quality of breath sounds (normal, absent, bronchial, or decreased) (kappa = -.02-.25). Following the education session, there was a general improvement in reliability (kappa = -.30-.77), especially for group 3 (specialists in other areas). The most improvement was noted for the detection of the quality of breath sounds (kappa = .08-.50). CONCLUSION AND DISCUSSION: Reliability of auscultation was poor to fair, in general, before the education session. There was a definite improvement in reliability after the education session. There was no clear effect of clinical experience on reliability, and the agreement among observers appeared to depend on the abnormal lung sound present. Limitations of this study and recommendations for future research are discussed. [Brooks D, Thomas J. Interrater reliability of auscultation of breath sounds among physical therapists.

Auscultation↗

Reliability of scoring respiratory disturbance indices and sleep staging.

STUDY OBJECTIVES: Unattended, home-based polysomnography (PSG) is increasingly used in both research and clinical settings as an alternative to traditional laboratory-based studies, although the reliability of the scoring of these studies has not been described. The purpose of this study is to describe the reliability of the PSG scoring in the Sleep Heart Health Study (SHHS), a multicenter study of the relation between sleep-disordered breathing measured by unattended, in-home PSG using a portable sleep monitor, and cardiovascular outcomes. DESIGN: The reliability of SHHS scorers was evaluated based on 20 randomly selected studies per scorer, assessing both interscorer and intrascorer reliability. RESULTS: Both inter- and intrascorer comparisons on epoch-by-epoch sleep staging showed excellent reliability (kappa statistics >0.80), with stage 1 having the greatest discrepancies in scoring and stage 3/4 being the most reliably discriminated. The arousal index (number of arousals per hour of sleep) was moderately reliable, with an intraclass correlation (ICC) of 0.54. The scorers were highly reliable on various respiratory disturbance indices (RDIs), which incorporate an associated oxygen desaturation in the definition of respiratory events (2% to 5%) with or without the additional use of associated EEG arousal in the definition of respiratory events (ICC>0.90). When RDI was defined without considering oxygen desaturation or arousals to define respiratory events, the RDI was moderately reliable (ICC=0.74). The additional use of associated EEG arousals, but not oxygen desaturation, in defining respiratory events did little to increase the reliability of the RDI measure (ICC=0.77). CONCLUSIONS: The SHHS achieved a high degree of intrascorer and interscorer reliability for the scoring of sleep stage and RDI in unattended in-home PSG studies.

Humans↗

Reliability of electrocardiogram interpretation in critically ill patients.

OBJECTIVE: To assess the intrarater and interrater reliability of electrocardiogram (ECG) interpretation in critically ill patients and to assess the effect of knowledge of cardiac troponin values on these reliability estimates. DESIGN: Prospective cohort study. SETTING: Fifteen-bed medical-surgical intensive care unit. PATIENTS: Consecutive adults admitted over a 2-month period. MEASUREMENTS AND RESULTS: All consecutive 12-lead ECGs were interpreted independently by two raters for the presence of myocardial ischemia or infarction and secondarily for specific ischemic ECG abnormalities. The ECGs were first interpreted blinded to the patient's troponin levels and reinterpreted on two separate occasions, blinded and unblinded to the troponin values. Results are reported using chance-independent agreement (phi) with associated 95% confidence intervals. For the presence of ischemia or infarction, the intrarater reliability ranged from fair to moderate (phi = 0.35 [95% confidence interval = 0.16, 0.52] and 0.59 [0.33, 0.77] for the two raters, respectively); interrater reliability was slight when blinded to troponin levels (phi = 0.18 [0.03, 0.32]) and increased to moderate when the raters were unblinded to troponin values (phi = 0.52 [0.33, 0.66], p value for the difference = .004). For specific ECG changes, the intrarater and interrater reliability were low for T-wave flattening, whereas detection of a left bundle branch block showed high reliability. CONCLUSIONS: ECG interpretation in critically ill patients for the presence of myocardial ischemia or infarction showed moderate reliability at best; however, there was high reliability for specific ECG changes. Knowledge of the patient's troponin values increased the reliability for all studied ECG changes and resulted in a statistically significant increase in the interrater reliability for diagnosing myocardial ischemia or infarction. Additional studies assessing the appropriate methods of diagnosing myocardial ischemia and infarction and assessing the reliability of these diagnostic tests in critically ill patients are required.

Critical Illness↗

Reliability analysis based on the losses from failures.

The conventional reliability analysis is based on the premise that increasing the reliability of a system will decrease the losses from failures. On the basis of counterexamples, it is demonstrated that this is valid only if all failures are associated with the same losses. In case of failures associated with different losses, a system with larger reliability is not necessarily characterized by smaller losses from failures. Consequently, a theoretical framework and models are proposed for a reliability analysis, linking reliability and the losses from failures. Equations related to the distributions of the potential losses from failure have been derived. It is argued that the classical risk equation only estimates the average value of the potential losses from failure and does not provide insight into the variability associated with the potential losses. Equations have also been derived for determining the potential and the expected losses from failures for nonrepairable and repairable systems with components arranged in series, with arbitrary life distributions. The equations are also valid for systems/components with multiple mutually exclusive failure modes. The expected losses given failure is a linear combination of the expected losses from failure associated with the separate failure modes scaled by the conditional probabilities with which the failure modes initiate failure. On this basis, an efficient method for simplifying complex reliability block diagrams has been developed. Branches of components arranged in series whose failures are mutually exclusive can be reduced to single components with equivalent hazard rate, downtime, and expected costs associated with intervention and repair. A model for estimating the expected losses from early-life failures has also been developed. For a specified time interval, the expected losses from early-life failures are a sum of the products of the expected number of failures in the specified time intervals covering the early-life failures region and the expected losses given failure characterizing the corresponding time intervals. For complex systems whose components are not logically arranged in series, discrete simulation algorithms and software have been created for determining the losses from failures in terms of expected lost production time, cost of intervention, and cost of replacement. Different system topologies are assessed to determine the effect of modifications of the system topology on the expected losses from failures. It is argued that the reliability allocation in a production system should be done to maximize the profit/value associated with the system. Consequently, a method for setting reliability requirements and reliability allocation maximizing the profit by minimizing the total cost has been developed. Reliability allocation that maximizes the profit in case of a system consisting of blocks arranged in series is achieved by determining for each block individually the reliabilities of the components in the block that minimize the sum of the capital, operation costs, and the expected losses from failures. A Monte Carlo simulation based net present value (NPV) cash-flow model has also been proposed, which has significant advantages to cash-flow models based on the expected value of the losses from failures per time interval. Unlike these models, the proposed model has the capability to reveal the variation of the NPV due to different number of failures occurring during a specified time interval (e.g., during one year). The model also permits tracking the impact of the distribution pattern of failure occurrences and the time dependence of the losses from failures.

Journal Article↗