PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Getting the story straight: evaluating the test-retest reliability of a university health history questionnaire.

This study was designed to establish the reliability of a health history questionnaire used as a screening tool for incoming university students. The authors used a test-retest design, with a test interval of 6 months, on a sample of medical and nursing students. The analysis focused on overall reliability of the questionnaire and reproducibility of specific items, based on question format. Questionnaire items of specific interest were those with dichotomous yes/no response options versus open-ended format questions, those using the words frequently or recently, or those that asked multiple questions. Demographic characteristics of the subjects were considered in the evaluation of reliability. Overall reliability of the questionnaire (93.6%) was above the anticipated level of 90%, and subject sex or program of study did not show any significant differences in reproducibility of responses. Although wording of questions did not affect item reliability, dichotomous format questions demonstrated a higher degree of reliability (96.4%) than the overall reliability of the questionnaire. Recommendations for enhancing the reliability of the questionnaire are based on item analysis and information gathered from interviews with subjects.

Adult↗

Inter-observer reliability of clinical outcome measures in a lower limb amputee population.

PURPOSE: In an attempt to find a more clinically useful functional outcome measure specifically tailored for lower limb amputees undergoing inpatient prosthetic rehabilitation, a 6-month prospective assessment of inter-rater reliability for Harold Wood-Stanmore Mobility Scale Data, including two handicap scales, was undertaken. An analysis of the data is presented in this paper. METHODS: An inter-rater reliability study was undertaken using four observers to complete admission and discharge scores for the three disability/handicap scales on 14 consecutive patients over 6 months. RESULTS: The disability mobility scale demonstrated perfect observer agreement on admission and at discharge the inter-rater reliability for this measure was high (0.83). By contrast, reliability between observers for admission scores on the handicap mobility scale was poor at 0.49 but reasonably high on discharge (0.83). On admission, inter-rater reliability for handicap physical independence was very low (0.15). At discharge, reliability improved to 0.69 being more consistent with results achieved for the other axes. CONCLUSIONS: This study confirms the good inter-rater reliability demonstrated previously in the literature but reveals poor inter-rater reliability for the two handicap scales. The latter will require modification before they can be used with confidence in conjunction with the disability scale.

Adult↗

Reliability of a questionnaire and an ergonomic checklist for assessing working conditions and health at call centres.

BACKGROUND: The purpose was to study the test-retest reliability and internal consistency of questions in a questionnaire concerning working conditions and health and the inter-rater reliability of observations and measurements according to an ergonomic checklist. METHOD: Fifty-seven operators participated in a retest questionnaire and 58 operators participated in an inter-observer test. RESULTS: The questions had fair to good or higher reliability in 142 of the total of 312. Twenty-seven of the total of 44 variables in the ergonomic checklist were classified as having fair to good or higher reliability. CONCLUSIONS: About half of the questions had fair to good or higher reliability and can be recommended for further analyses. The majority of variables in the ergonomic checklist were classified as having fair to good or higher reliability. Low reliability does not necessarily indicate that the reliability of the test, per se, is low but may signify that the conditions measured vary over time or that the answers are aggregated in one part of the scale.

Cross-Sectional Studies↗

Clinical tests on impairment level related to low back pain: a study of test reliability.

The objectives of the study were, in a working population, to standardize and evaluate a set of clinical tests on impairment level related to the low back with reference to intra- and inter-rater reliability. The study was undertaken in two steps. In step 1, 15 tests were examined for inter-rater reliability by three pairs of physiotherapists and for intra-rater reliability by one physiotherapist. Intra-rater reliability was acceptable (kappa > 0.40) for 14 of the 15 tests. Inter-rater reliability was acceptable for 7 of the 15 tests. In step 2, the tests, indicating a non-acceptable inter-rater reliability (kappa < 0.40) were further standardized and retested by two of the physiotherapists. This further standardization procedure resulted in an acceptable inter-rater reliability for all of these tests. Clinical tests of a working population should preferably be performed by the same rater. However, when tests are performed by different raters, it is suggested that test procedures should be regularly standardized, and in pain provocation tests, the magnitude of the applied pressure should be checked regularly and compared with co-raters, in order to improve inter-rater reliability.

Adult↗

Accuracy and reliability of a dynamic biomechanical skin measurement probe for the analysis of stiffness and viscoelasticity.

A novel instrument has been devised for the in vivo examination of the dynamic biomechanical properties of skin. These properties include stiffness and viscoelasticity. The advantage of the device is its ability to examine the skin dynamically, thereby eliminating preconditioning effects. Furthermore, it is portable, hand-held and easy to operate in the clinical environment. The objective of this study was to determine the accuracy and reliability of the dynamic biomechanical skin measurement (DBSM) probe. The accuracy was determined by examining a series of silicone elastomer specimens. A comparison of the shear modulus (G*), obtained from a static indentation system, with stiffness, obtained from the DBSM probe, was performed. The reliability was determined by examining both silicone elastomers and forearm volar skin in vivo. In both cases assessment was by six different operators (inter-reliability) and also by an individual operator (intra-reliability). Statistical analysis was performed using Levene's test of homogeneity and analysis of variance to ascertain if there were significant differences between operators (inter-reliability) and with one individual operator (intra-reliability). It can be concluded, from this study, that the DBSM probe is accurate (R2 = 0.96, p = 0.01). It is also inter- and intra-reliable when assessing elastomer stiffness and skin stiffness. However, phase lag was not found to be a useful indicator of device reliability. It is anticipated that this device will be used to examine dermatological conditions and the benefits, or otherwise, of treatment. The DBSM probe promises to contribute to the objective measurement of physical properties of the skin in future investigative studies.

Biomechanical Phenomena↗

The validity and reliability of the affective competency score to evaluate death disclosure using standardized patients.

OBJECTIVE: To explore the validity and reliability of the affective competency score (ACS), compared to a global rating measure to predict overall competency to perform a death disclosure in a standardized patient exercise and to investigate useful thresholds of the ACS. METHODS: Thirty-seven fourth-year students underwent standardized patient training in death disclosure during a fourth-year emergency medicine clerkship. Students were evaluated using a checklist, an ACS, and a global rating assessment. ACS interrater reliability, interitem reliability, item-total reliability, and split-half reliability were calculated. Area under the curve (AUC) measurements were used to establish criterion validity. RESULTS: For the ACS, item-total correlations ranged from 0.76 to 0.85, 0.76 to 0.93, and 0.42 to 0.87; the split-half reliability was 0.82 (p = 0.0001), 0.86 (p = 0.0001) and 0.55 (p = 0.0007) for the standardized patient (SP), the faculty and the medical students, respectively. Interitem correlations were adequate. A moderate interrater correlation of the ACS was observed between the faculty observer and the SP (r = 0.47; p = 0.04); however, the medical students' self evaluation did not correlate significantly with either the SP (r = -0.04; p = 0.79), or the faculty observer (r = 0.00; p = 0.99). The AUC for was 0.98 (95% confidence interval [CI] 0.94 to 1.00), 0.87 (95% CI 0.73 to 0.99), and 0.74 (95% CI 0.53 to 0.95) for the faculty, SP, and medical student, respectively. CONCLUSIONS: The ACS may be a valid, reliable, and useful measure to assess communication skills by faculty or SPs in this setting. At an ACS score of 16, 19, and 21 points for faculty, SPs, and medical students, respectively, there is 100% specificity for the detection of competency assessed on a global rating. However, the ACS appears to have limited reliability and validity when used by medical students.

Adult↗

A reliability study of an instrument for measuring general practitioner consultation skills: the LIV-MAAS scale.

OBJECTIVE: To evaluate the reliability of a new tool, the LIV-MAAS, in assessing consultation competence in UK general practice. DESIGN: These were pilot studies, with small numbers of participants. Videoed general practitioner (GP) consultations were analysed by trained lay and professional raters, using the LIV-MAAS. The inter-rater reliabilities were assessed. Four videos were assessed by five raters in a pilot study. After this, 71 consultations from eight doctors were assessed by sets of three raters. MAIN MEASURES: Inter-rater reliabilities and inter-consultation reliabilities. RESULTS: For the pilot study, the estimated inter-rater reliability ranged from 0.69 (one rater) to 0.91 (five raters). For the main study, the estimated inter-rater reliability for the LIV-MAAS checklist using two raters was 0.71, and using three raters it was 0.78. Mean differences in reliability within each series of nine consultations were 0.20 (three raters) and 0.42 (two raters). CONCLUSIONS: As a measure of 'consultation competence', administered by trained raters (medical or lay) to real GP consultations, the LIV-MAAS instrument shows adequate reliability and stability but would benefit from considerable shortening. Further development of the LIV-MAAS and testing with larger samples are required.

Adult↗

Goniometric reliability in a clinical setting. Subtalar and ankle joint measurements.

Measurements of the subtalar joint neutral (STJN) position and passive range of motion (PROM) of the ankle joint and the subtalar joint (STJ) are often part of a physical therapy evaluation. These measurements may be used in treatment planning, such as in the prescription of specialized shoes or orthoses. Therefore, reliability of these measurements, as they are obtained clinically, must be determined. The purpose of this study was to examine the reliability of measurements of the STJN position and of ankle and STJ PROM. To determine reliability, repeated measurements of the STJN position and of STJ PROM were taken on the involved feet of 43 patients with neurologic orthopedic disorders (including both feet of 7 patients), and measurements of ankle PROM (dorsiflexion and plantar flexion) were taken on 42 of these patients (including both feet of 7 patients). Intraclass correlation coefficients (ICCs) for intratester reliability ranged from .74 to .90 for ankle and STJ measurements. The ICCs for intertester reliability were .25 for measuring the STJN position, .32 for STJ inversion, and .17 for SJJ eversion. The ICCs for intertester reliability were .50 for ankle dorsiflexion and .72 for ankle plantar flexion. Goniometric measurements of the STJN position and of PROM of the ankle and STJ appear to be moderately reliable if taken by the same therapist over a short period of time. With the exception of ankle plantar flexion, these measurements cannot be considered to be reliable between therapists.

Adolescent↗

The reliability of the three-dimensional FASTRAK measurement system in measuring cervical spine and shoulder range of motion in healthy subjects.

OBJECTIVES: To assess the inter-observer and intra-observer reliability of a new three-dimensional measurement system, the FASTRAK, in measuring cervical spine flexion/extension, lateral flexion and rotation and shoulder flexion/extension, abduction and external rotation in healthy subjects. METHODS: The study was conducted in two parts. One part assessed inter-observer reliability with two observers measuring 40 subjects. The other part assessed intra-observer reliability with one observer measuring 32 subjects on three occasions. All subjects had unrestricted, pain-free cervical spine and shoulder movement. Reliability was measured by the intraclass correlation coefficient [ICC(2,1)]. RESULTS: The inter-observer ICCs for the cervical spine ranged from 0.61 to 0.89 and for the shoulder from 0.68 to 0.75. After removal of outliers, all ICCs were above 0.70. Intra-observer ICCs for the cervical spine ranged from 0.54 to 0.82 and for the shoulder from 0.62 to 0.81. After removal of outliers, all ICCs were above 0.70 except for shoulder abduction (0.62). CONCLUSIONS: Whilst all movements measured by the FASTRAK showed good reliability, the reliability of the whole movement in a plane (e.g. left plus right lateral flexion) was better than for the separate movements (e.g. left and right lateral flexion taken separately). Inter-observer reliability was generally better than intra-observer reliability for most cervical spine movements, suggesting that variability of movement within subjects (e.g. over a period of days) for these movements was greater than variability between measures on the same occasion.

Adult↗

The reliability and functional validity of visual and semiautomatic sleep/wake scoring in the Møll-Wistar rat.

The present paper has three major objectives: first, to document the reliability of a published criteria set for sleep/wake scoring in the rat; second, to develop a computer algorithm implementation of the criteria set; and third, to document the reliability and functional validity of the computer algorithm for sleep/wake scoring. The reliability of the visual criteria was assessed by letting two raters separately score 8 hours of polygraph records from the light period from five rats (14,040 10-second scoring epochs). Scored stages were waking, slow-wave sleep-1, slow-wave sleep-2, transition type sleep and rapid eye movement (REM) sleep. The visual criteria had good interrater reliability [Cohen's kappa (kappa) = 0.68], with 92.6% agreement on the waking/nonrapid eye movement (NREM) sleep/REM sleep distinction (kappa = 0.89). This indicated that the criteria allow separate raters to independently classify sleep/wake stages with very good agreement. An independent group of 10 rats was used for development of an algorithm for semiautomatic computer scoring. A close implementation of the visual criteria was chosen. The algorithm was based on power spectral densities from two electroencephalogram (EEG) leads and on electromyogram (EMG) activity. Five 2-second fast Fourier transform (FFT) epochs from each EEG/EMG lead per 10-second sleep/wake scoring epoch were used to take the spatial and temporal context into account. The same group of five rats used in visual scoring was used to appraise reliability of computerized scoring. The computer score was compared with the visual score for each rater. There was a lower agreement (kappa = 0.57 and 0.62 for the two raters) than in interrater visual scoring [percent agreement 87.7 and 89.1% (kappa = 0.82 and 0.84) in the waking/NREM sleep/REM sleep distinction]. Subsequently, the computer scores of the raters were compared. The interrater reliability was better than the interrater reliability for visual scoring (kappa = 0.75), with 92.4% agreement for the waking/NREM sleep/REM sleep distinction (kappa = 0.89). The computer scoring algorithm was applied to data from a third independent group of rats (n = 6) from an acoustical stimulus arousal threshold experiment, to assess the functional validity of the scoring directly with respect to arousal threshold. The computer algorithm scoring performed as well as the original visual sleep/wake stage scoring. This indicated that the lower intrarater reliability did not have a significant negative influence on the functional validity of the sleep/wake score.

Algorithms↗

Night-to-night arousal variability and interscorer reliability of arousal measurements.

STUDY OBJECTIVES: Measurement of arousals from sleep is clinically important, however, their definition is not well standardized, and little data exist on reliability. The purpose of this study is to determine factors that affect arousal scoring reliability and night-to-night arousal variability. DESIGN: The night-to-night arousal variability and interscorer reliability was assessed in 20 subjects with and without obstructive sleep apnea undergoing attended polysomnography during two consecutive nights. Five definitions of arousal were studied, assessing duration of electroencephalographic (EEG) frequency changes, increases in electromyographic (EMG) activity and leg movement, association with respiratory events, as well as the American Sleep Disorders Association (ASDA) definition of arousals. SETTING: NA. PATIENTS: NA. INTERVENTIONS: NA. RESULTS: Interscorer reliability varied with the definition of arousal and ranged from an Intraclass correlation (ICC) of 0.19 to 0.92. Arousals that included increases in EMG activity or leg movement had the greatest reliability, especially when associated with respiratory events (ICC 0.76 to 0.92). The ASDA arousal definition had high interscorer reliability (ICC 0.84). Reliability was lowest for arousals consisting of EEG changes lasting <3 seconds (ICC 0.19 to 0.37). The within subjects night-to-night arousal variability was low for all arousal definitions CONCLUSION: In a heterogeneous population, interscorer arousal reliability is enhanced by increases in EMG activity, leg movements, and respiratory events and decreased by short duration EEG arousals. The arousal index night-to-night variability was low for all definitions.

Arousal↗

Improving the reliability of a combined phenological time series by analyzing observation quality.

Collecting phenological data is a slow process. Although such data have been collected by a number of organizations, the reliability of these data is not known because the data-generating process cannot be repeated. No further observations to improve the reliability can be obtained. However, the data usually consist of several overlapping observation series and this overlap can be utilized to construct a combined phenological time series and to improve its reliability. We have developed two techniques for selecting the most reliable observations or observation series and thereby improve the reliability of the combined time series. Both techniques require that the method used to combine the separate phenological time series adjusts the individual series to eliminate possible systematic differences between them. A data set of bud burst in Betula pendula Roth collected in Central Finland during 1896-1955 was adjusted and used to test both techiques. Both techniques considerably improved the reliability of the combined time series; the mean of the confidence intervals of the annual means decreased by 12%. Despite the improvement in reliability, the resulting changes in the annual values of the combined time series were small, the largest change being 2.5 days. Removing outliers was the most effective method of improving reliability, i.e., it resulted in the greatest improvement with the smallest number of discarded observations.

Journal Article↗

Assessment of the intrarater and interrater reliability of an established clinical task analysis methodology.

BACKGROUND: Task analysis may be useful for assessing how anesthesiologists alter their behavior in response to different clinical situations. In this study, the authors examined the intraobserver and interobserver reliability of an established task analysis methodology. METHODS: During 20 routine anesthetic procedures, a trained observer sat in the operating room and categorized in real-time the anesthetist's activities into 38 task categories. Two weeks later, the same observer performed task analysis from videotapes obtained intraoperatively. A different observer performed task analysis from the videotapes on two separate occasions. Data were analyzed for percent of time spent on each task category, average task duration, and number of task occurrences. Rater reliability and agreement were assessed using intraclass correlation coefficients. RESULTS: Intrarater reliability was generally good for categorization of percent time on task and task occurrence (mean intraclass correlation coefficients of 0.84-0.97). There was a comparably high concordance between real-time and video analyses. Interrater reliability was generally good for percent time and task occurrence measurements. However, the interrater reliability of the task duration metric was unsatisfactory, primarily because of the technique used to capture multitasking. CONCLUSIONS: A task analysis technique used in anesthesia research for several decades showed good intrarater reliability. Off-line analysis of videotapes is a viable alternative to real-time data collection. Acceptable interrater reliability requires the use of strict task definitions, sophisticated software, and rigorous observer training. New techniques must be developed to more accurately capture multitasking. Substantial effort is required to conduct task analyses that will have sufficient reliability for purposes of research or clinical evaluation.

Adult↗

Reliability of proxy-reported and self-reported household appliance use.

Exposure assessment presents a major challenge for studies evaluating the association between household exposure to electric and magnetic fields and adverse health outcomes, especially the reliance on proxy respondents when study subjects themselves have died. We evaluated the reliability of proxy- and self-reported household appliance exposure. We recruited 92 healthy couples through either random-digit dialing or newspaper advertisements. Trained interviewers administered questionnaires to each member of a couple independently to assess the reliability of proxy-reported household appliance use. Eighty-five couples completed a second interview 2 months later to assess the reliability of self-reported appliance use. Reliability of proxy-reported appliance exposure was good when we inquired about having any exposure to each of the eight indicator appliances during the past year (range of kappa coefficients = 0.63-0.85; median = 0.76) but was lower with increased time to recall or increased detail. Reliability of self respondents reporting 2 months apart was excellent (range of kappa coefficients = 0.75-0.94; median = 0.87) for having any exposure to the eight indicator appliances during the past year, but reliability was again lower with increased detail. When we used self reports at the first interview as the standard, little systematic over- or underreporting occurred for proxy respondents or for self respondents reporting 2 months later. Because this study did not include cases of specific disease, these findings of no systematic differences in reporting do not refer to case or control status. In summary, reliability of self respondents' reports of appliance use is very good for recent time periods and good for broad aspects of exposure in distant time periods. Proxy respondents can provide information regarding broad aspects of appliance exposure in the past year, but detailed aspects of exposure or exposure in more distant time periods is not reliable.

Adult↗

Improvement of reliability of an oral examination by a structured evaluation instrument.

The main purposes of this study were to estimate the reliability of oral examinations administered to medical students during a clinical clerkship and to improve the reliability of this evaluation technique. In the first part of the study, the reliability of oral examinations as traditionally administered was estimated. The average intraclass reliability coefficient for these examinations was .48. Cassette recordings of these oral examinations were also rated by the faculty members. The average intraclass reliability coefficient of the ratings of the taped performances was .82. In the second part of the study, the reliability of oral examinations was investigated with the raters using a newly developed evaluation form. The average intraclass reliability of the oral examination using the evaluation form was .67, a noticeable increase over the .48 obtained without the form. The average intraclass reliability of ratings made from tape recordings of these oral examinations was .62.

Clinical Clerkship↗

Gunshot femoral shaft fractures: is the current classification system reliable?

The reliability of the AO/Orthopaedic Trauma Association classification system has not been evaluated for diaphyseal fractures or fractures attributable to gunshot injuries. Therefore, the current authors assessed its reliability for diaphyseal femur fractures and investigated the effect of a gunshot mechanism of injury. Forty-seven diaphyseal femur fractures, 23 caused by gunshots and 24 caused by blunt trauma, were classified by four observers on two occasions. The interobserver and intraobserver reliability of each level of the AO/Orthopaedic Trauma Association classification was assessed with kappa statistics. Determination of fracture type had substantial interobserver and intraobserver reliability for gunshot and blunt injuries. Reliability decreased at the subsequent levels of the classification. Fractures caused by gunshots compared with those caused by blunt trauma were characterized by significantly lower interobserver agreement on fracture group (k = 0.26 versus 0.45) and subgroup (k = 0.21 versus 0.38). The AO/Orthopaedic Trauma Association classification system has substantial interobserver and intraobserver reliability when evaluating the type of diaphyseal femur fractures. Determination of fracture group and subgroup, however, progressively reduces the reliability of the classification, especially for fractures caused by a gunshot. Diaphyseal femur fractures caused by gunshots, by means of their fracture patterns, cannot be classified reliably with the AO/Orthopaedic Trauma Association classification system.

Femoral Fractures↗

Reliability of the Glasgow Coma Scale when used by emergency physicians and paramedics.

We sought to determine the reliability of the Glasgow Coma Scale (GCS) when used by emergency physicians and paramedics. We performed a prospective sequential trial in a classroom setting, with subjects blinded to others' scoring. Nineteen university-affiliated emergency physicians and 41 professional paramedics from an urban EMS system voluntarily participated. Participants viewed four videotaped scenes in which a patient is assessed by a paramedic. The first three scenes represented severe, intermediate, and no/mild alteration in level of consciousness (LOC). The findings in the fourth scene were identical to the first, allowing determination of intrarater reliability. The Kappa statistic was used to determine interrater reliability; the reliability coefficient determined intrarater reliability. Kappa was significant (p < 0.0001) for severe (kappa = 0.48), intermediate (kappa = 0.34), and no/mild (kappa = 0.85) conditions. Intrarater reliability (r1,2) for emergency physicians was 0.66 (p < 0.01) and for paramedics was 0.63 (p < 0.01). The GCS shows statistically significant reliability (i.e., significant agreement) between emergency physicians and emergency medical technician-paramedics. It also has a significant level of intrarater reliability.

Allied Health Personnel↗

Reliability in adolescent reporting of clinician counseling, health care use, and health behaviors.

BACKGROUND: Accurate measures of health-care use by adolescents would be useful in managed care quality assurance, public health surveillance, and health-care research. OBJECTIVE: To assess test-retest reliability and factors associated with reliability of adolescent reports of clinician counseling, preventive health services, and health behaviors. RESEARCH DESIGN: A convenience sample of high school students (N = 253) completed identical paper-and-pencil surveys in school and 2 weeks apart. Multiple linear regression was used to evaluate the influence on response reliability of individual factors and question item characteristics. Reliability was assessed using Cohen kappa. RESULTS: Kappa values for specific questions varied widely (0.94-0.33). Median kappa values for behavioral, counseling, and health-service questions were 0.74, 0.63, 0.56, respectively. Lower sentence complexity, certain time frames (ever, age at first occurrence, last time), and behavioral question type were associated with greater reliability in adolescent reporting (final model R2 = 0.54). Adolescents' age and ethnicity were not predictive of reliability, though girls were slightly more reliable reporters than boys. Overall, the prevalence of responses at times 1 and 2 were similar; 95% of responses at time 2 were within 5 percentage points of time-1 estimates (SD = 2.4). CONCLUSIONS: The reliability of adolescent reporting was strongly influenced by question characteristics such as sentence complexity and time frame; these should be carefully considered in the construction of questionnaires for adolescents. Adolescents can be an accurate source of health-care service data.

Adolescent↗