PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Intraobserver and interobserver reliability of assessments of impairments and disabilities.

BACKGROUND AND PURPOSE: The purpose of this study was to evaluate the interobserver and intraobserver reliability of assessments of impairments and disabilities. SUBJECTS AND METHODS: One physical therapist's assessments were examined for intraobserver reliability. Judgments of two pairs of therapists were used to examine interobserver reliability. Reliability was assessed by Cohen's kappa. RESULTS: Of the 42 impairments and disabilities assessed by the physical therapist in the intraobserver reliability study, kappa values could be calculated for 33 items. For 31 items (94%), kappa values ranged from .40 to .91, and 2 items (6%) had kappa values of less than .40. To determine interobserver reliability, 37 items were assessed in one practice. Kappa values could be calculated for 34 items, with 30 items (88%) having kappa values ranging from .41 to .80 and 4 items (12%) showing "poor" agreement. In the second practice, 47 items were assessed for interobserver reliability. Kappa values could be calculated for 40 items, with 11 items (27.5%) having kappa values ranging from .41 to .84. Poor agreement was shown for the remaining 29 items (72.5%). CONCLUSION AND DISCUSSION: Assessments of impairments and disabilities are potentially reliable. The differences between practices of the interobserver reliability study can be explained by the fact that one of the therapists did not receive training in the use of the assessment form. More generalizable conclusions will require further study with more subjects and therapists.

Adolescent↗

Reliability and construct validity studies of an obstacle course assessment of wheelchair user performance.

The objective of this study is to evaluate the reliability and construct validity of an obstacle course assessment of wheelchair user performance (OCAWUP). Seventeen experienced wheelchair users using three different propulsion methods (two hands, one hand and one foot or motorized wheelchair) were assessed twice on the 10 obstacles of the OCAWUP. To evaluate reliability, time (in seconds) and degree of ease (DE) in overcoming obstacles (four-level scale) were assessed by three occupational therapists. Construct validity was assessed by verifying whether the OCAWUP's global score of ease (GSE) varied with wheelchair propulsion methods and the functional independence measure (FIM). Intraclass correlation coefficients calculated for reliability of time and GSE were up to 0.74 for test-retest reliability and up to 0.97 for interrater reliability. Cohen's kappa coefficients calculated for DE reliabilities varied from 0.09 to 1.0 with degrees of association up to 65%. A significant difference (P<or =0.01) in GSE scores is shown between propulsion methods. A high significant correlation (rs=0.84, P< or =0.05) was found between GSE and scores of FIM items associated with mobility. Reliability of time and GSE is substantial or better. Reliability of DE is not as good, with a few items indicating poor reliability. OCAWUP's construct validity is good.

Acceleration↗

Infant polysomnography: reliability and validity of infant arousal assessment.

Infant arousal scoring based on the Atlas Task Force definition of transient EEG arousal was evaluated to determine (1). whether transient arousals can be identified and assessed reliably in infants and (2). whether arousal and no-arousal epochs scored previously by trained raters can be validated reliably by independent sleep experts. Phase I for inter- and intrarater reliability scoring was based on two datasets of sleep epochs selected randomly from nocturnal polysomnograms of healthy full-term, preterm, idiopathic apparent life-threatening event cases, and siblings of Sudden Infant Death Syndrome infants of 35 to 64 weeks postconceptional age. After training, test set 1 reliability was assessed and discrepancies identified. After retraining, test set 2 was scored by the same raters to determine interrater reliability. Later, three raters from the trained group rescored test set 2 to assess inter- and intrarater reliabilities. Interrater and intrarater reliability kappa's, with 95% confidence intervals, ranged from substantial to almost perfect levels of agreement. Interrater reliabilities for spontaneous arousals were initially moderate and then substantial. During the validation phase, 315 previously scored epochs were presented to four sleep experts to rate as containing arousal or no-arousal events. Interrater expert agreements were diverse and considered as noninterpretable. Concordance in sleep experts' agreements, based on identification of the previously sampled arousal and no-arousal epochs, was used as a secondary evaluative technique. Results showed agreement by two or more experts on 86% of the Collaborative Home Infant Monitoring Evaluation Study arousal scored events. Conversely, only 1% of the Collaborative Home Infant Monitoring Evaluation Study-scored no-arousal epochs were rated as an arousal. In summary, this study presents an empirically tested model with procedures and criteria for attaining improved reliability in transient EEG arousal assessments in infants using the modified Atlas Task Force standards. With training based on specific criteria, substantial inter- and intrarater agreement in identifying infant arousals was demonstrated. Corroborative validation results were too disparate for meaningful interpretation. Alternate evaluation based on concordance agreements supports reliance on infant EEG criteria for assessment. Results mandate additional confirmatory validation studies with specific training on infant EEG arousal assessment criteria.

Arousal↗

Effects of interrater reliability of psychopathologic assessment on power and sample size calculations in clinical trials.

Although rater training is increasingly used to improve the quality of the investigated outcome parameters, the reliability of assessments is not perfect. Thus, empirical reliability estimates should be used instead of theoretically assumed perfect reliability. Implications of the reliability of psychiatric assessments for sample size and power calculations in clinical trials are presented. The theoretical basis of sample size and power calculations using empirical reliability scores is delineated. Examples from contemporary research on schizophrenia and depression are used to illustrate several implications for study design and interpretation of results. The tremendous impact of the lack of reliability of psychopathologic assessments on sample size, power, and detectable true score differences in clinical trials is shown. The problem of multiple outcome variables with different reliabilities is addressed. Studies lacking power because of unreliable assessments carry the risk of false-negative findings and raise ethical questions. Rater training is strongly recommended to assess and improve interrater reliability whenever necessary and possible before trials are started. Sample size calculations and power analysis should be based on empirical reliability values of outcome parameters as part of quality assurance and cost savings.

Clinical Trials as Topic↗

The reliability of monopolar and bipolar fine-wire electromyographic measurement of muscle fatigue.

Bipolar intramuscular wire electrodes and spectral analysis of the electromyographic signal have been used to measure fatigue in muscles that cannot be studied with surface electrodes. Intramuscular electrodes can detect a greater range of frequencies from muscle, obtain a less distorted signal, and are therefore felt to be more sensitive to detecting fatigue. To determine the reliability and sensitivity of electrode placement (with a fixed distance) for assessing muscle fatigue, we placed three intramuscular electrodes in and two surface electrodes on the biceps brachii of 30 healthy male subjects. With these electrodes, we devised eight configurations that were analyzed separately for reliability. Subjects performed four, 30-s isometric fatiguing contractions divided between two testing sessions. Mean and median frequency of the power density spectrum were plotted against time. Linear regression was performed to obtain slopes, which were used as indicators of fatigue. The bipolar surface electrode configuration displayed mean and median frequency intrasession and mean frequency intersession reliability for slope. All four bipolar fine-wire configurations had mean and median frequency intrasession reliability (P < or 0.05). Only three of the four bipolar fine-wire configurations approached mean frequency intersession reliability, and none fo the four displayed median frequency intersession reliability. the configuration with distal bipolar intramuscular electrodes placed 1 cm apart was the most reliable intramuscular technique. The bipolar fine-wire configuration studied showed a trend toward better reliability than monopolar fine-wire configurations. No intramuscular technique, however, was reliable enough for clinical use in the study of fatigue.

Adolescent↗

Interobserver and intraobserver reliability of the japanese orthopaedic association scoring system for evaluation of cervical compression myelopathy.

STUDY DESIGN: The inter- and intraobserver reliabilities of an assessment scale for cervical compression myelopathy were examined statistically. This scoring system consists of seven categories: motor function of fingers, shoulder and elbow, and lower extremity; sensory function of upper extremity, trunk and lower extremity; and function of the bladder. It evaluates the severity of myelopathy by allocating points based on degree of dysfunction in each category. OBJECTIVES: To determine the inter- and intraobserver reliabilities of the revised scoring system (17 - 2 points) for cervical compression myelopathy proposed by the Japanese Orthopedic Association. SUMMARY OF BACKGROUND DATA: Several scales to assess clinical outcome from treatment of cervical compression myelopathy have been proposed. Most of these scales include items evaluated by observers. However, no system, including the Japanese Orthopedic Association scoring system, has yet been validated in terms of interobserver reliability. METHODS: From five different university hospitals, 10 spine surgery specialists, 10 orthopedic surgeons who had just passed the board examination of the Japanese Orthopedic Association, and 13 residents in the first or second year of orthopedic residency programs were chosen. The participants in this study were 29 patients with myelopathy secondary to ossification of the posterior longitudinal ligament selected from five participating university hospitals. Several surgeons interviewed each patient twice at intervals of 1 to 6 weeks. Inter- and intraobserver reliabilities of the total score for all categories were evaluated by the intraclass correlation coefficient. The extension of the kappa coefficient of Kraemer also was calculated for each category to assess reliability of multivariate categorical data. RESULTS: The interobserver reliability of the total score for the first interview (intraclass correlation coefficient = 0.813) and the intra- and interobserver reliabilities of the total score (intraclass correlation coefficient = 0.826) were high. The level of experience and the hospital slightly affected the reliability of the Japanese Orthopedic Association scoring system. The kappa values for intraobserver data generally were high in each category, whereas the kappa values for interobserver data were relatively low for the categories of shoulder-elbow motor function and lower extremity sensory function. CONCLUSIONS: The inter- and intraobserver reliabilities of the Japanese Orthopedic Association scoring system for cervical myelopathy were high, suggesting that this system is useful for assessment of cervical myelopathy in comparative studies of treatment.

Activities of Daily Living↗

Reliability of the lumbar flexion, lumbar extension, and passive straight leg raise test in normal populations embedded within a complete physical examination.

STUDY DESIGN: The study measured the reliability of the passive straight leg raise (SLR) test and lumbar range of motion (LROM) tests measured as continuous variables embedded within a comprehensive physical examination. OBJECTIVES: To determine the reliability of the SLR and LROM test scores when they are measured with a Cybex electronic inclinometer (Lumex, Inc., New York, NY) within a physical examination. SUMMARY OF BACKGROUND DATA: Good published empirical reliability exists for the Cybex and for SLR and LROM tests when the measurements are taken in isolation from other physical examination procedures. Reliability of the Cybex for continuous SLR and LROM measurement within a physical examination has not been assessed, however. METHODS: Forty-five participants were seen by one of two physician/physiotherapist teams. Participants were examined by both team members. The first examiner conducted the first tests and retested 1 week later (intrarater reliability). The second examined the participants the day after their first appointment (inter-rater reliability). RESULTS: Only two scores showed substantial reliability (defined as r > or = 0.60). These scores were left (r = 0.81) and right (r = 0.79) SLR intrarater reliability. All other scores fell below the specified cutoff. CONCLUSIONS: SLR and LROM scores used clinically are collected during comprehensive physical examinations. Most scores gathered under these conditions were not reliable. These findings have implications for the use of clinically derived SLR and LROM scores.

Adult↗

Intratester and intertester reliability of clinical measures of lower extremity anatomic characteristics: implications for multicenter studies.

OBJECTIVE: To determine whether multiple examiners could be trained to measure lower extremity anatomic characteristics with acceptable reliability and precision, both within (intratester) and between (intertester) testers. We also determined whether testers trained 18 months apart could perform these measurements with good agreement. SETTING: University's Applied Neuromechanics Research Laboratory. PARTICIPANTS: Sixteen, healthy participants (7 men, 9 women). ASSESSMENT OF RISK FACTORS: Six investigators measured 12 anatomic characteristics on the right lower extremity in the Fall of 2004. Four testers underwent training immediately preceding the study, and measured subjects on 2 separate days to examine intratester reliability. Two testers trained 18 months before the study (Spring 2002) measured each subject on day 1 to examine the consistency of intertester reliability when testers are trained at different times. MAIN OUTCOME MEASUREMENTS: Knee laxity, genu recurvatum, quadriceps angle, tibial torsion, tibiofemoral angle, hamstring extensibility, pelvic angle, navicular drop, femur length, tibial length, and hip anteversion. RESULTS: With few exceptions, all testers consistently measured each variable between test days (intraclass correlation coefficient>or=0.80). Intraclass correlation coefficient values were lower for intertester reliability (0.48 to 0.97), and improved from day 1 to day 2. Intertester reliability was similar when comparing testers trained 18 months before those trained immediately before the study. Absolute measurement error varied considerably across individual testers. CONCLUSIONS: Multiple investigators can be trained at different times to measure anatomic characteristics with good to excellent intratester reliability. Intratester reliability did not always ensure acceptable intertester reliability or measurement precision, suggesting more training (or more experience) may be required to achieve acceptable measurement reliability and precision between multiple testers.

Adult↗

Test-retest reliability of the alcohol use disorder identification test in a general population sample.

BACKGROUND: A number of different screening tests are frequently used in alcohol research, but our knowledge about the reliability of many of them is quite limited. Recently, this problem has received more attention. This article examines the test-retest reliability of one of these instruments-the Alcohol Use Disorder Identification Test (AUDIT)-in a general population sample. METHODS: A general population sample (n = 457) was tested and, after approximately 1 month, was retested by using the AUDIT. Correlation between the two tests has been examined with the intraclass correlation coefficient and the kappa coefficient in analysis of dichotomous variables. Specificity and sensitivity at a number of different cutoff scores have also been analyzed by using the first test as a criterion. RESULTS: On the item level, the correlations ranged between 0.6 and 0.8. The overall reliability of total AUDIT scores was 0.84. When stratified by gender, age, and consumer status, the total score reliability approximated 0.80 for all the categories except low alcohol consumers (0.51). Agreement using the recommended cutoff score of 8+ was also examined. The reliability (kappa) observed in the whole sample was 0.691, which was interpreted as a substantial agreement. By this cutoff, 91% were correctly classified at retest compared with the first test. AUDIT 8+ showed higher reliability for males, young people, and moderate consumers and low reliability among low consumers. In terms of reliability, the most optimal cutoff for women turned out to be 6 or more. CONCLUSIONS: According to these results, the test-retest reliability of AUDIT is high. The next step might be to examine to what extent the findings apply within health-care settings, which is what the test originally was designed for.

Adolescent↗

Reliability of grading lissamine green conjunctival staining.

PURPOSE: To assess the reliability of a lissamine green grading scale for conjunctival images. METHODS: A 20-second video clip of the right eye of 288 contact lens-wearing individuals was recorded using a digital slip-lamp camera after instilling liquid lissamine green. A single nasal and temporal still image were selected. A masked reader used the Oxford grading scale to grade the images on two occasions whereas a second masked reader graded each image on 1 occasion. kappa statistics and 95% confidence intervals (CIs) were used to determine the within- and between-grader reliability overall and when the sample was stratified by age, sex, contact lens type, and disease severity. RESULTS: There was substantial within-grader reliability for both the nasal (kappasimple = 0.69, 95% CI, 0.63-0.75) and temporal (kappasimple = 0.73, 95% CI, 0.67-0.79) images. There was moderate between-grader reliability for both the nasal (kappasimple = 0.51, 95% CI, 0.44-0.58) and temporal (kappasimple = 0.51, 95% CI, 0.44-0.58) images. Age, sex, and contact lens type did not affect within- or between-examiner reliability. There may have been an influence of disease severity on within-examiner reliability, because grading of the temporal images was significantly less reliable in the images with more significant staining. CONCLUSION: Within- and between-grader reliability of lissamine green staining seems to be at least substantial to moderate. Because the extent of conjunctival staining may influence reliability, this should be considered when studies may include patients with significant staining.

Adult↗

Reliability of an image analysis system for quantifying the radiographic trabecular pattern.

A reliability evaluation technique was used to examine the reliability of an image analysis system of the trabecular pattern and to determine the contribution of three possible sources of error variance. Two series of radiographs were taken of 14 lumbar vertebral slices (28 radiographs). Every radiograph was placed on a viewing box for digitization four times by a single operator (112 positions of radiographs) and from every position of a radiograph an area of 15 mm x 15 mm was digitized twice (224 samples for analysis). Ten geometrical characteristics of the trabecular pattern were studied and its orientation was analyzed in 12 directions. Reliability was determined by calculating Cronbach's alpha. This design enabled dividing the measurement error (1-alpha) into fractions associated with the X-ray procedure, the operator and the system. Using this reliability evaluation technique, it was found that the orientation variables are more reliable than the geometric variables. It was found that effort to increase the reliability should be directed toward improving the technical procedure of this image analysis system. Also, repeated measurements will increase the reliability. The number of repeated measurements based on a desired reliability can be calculated. This procedure of evaluation gives the opportunity to select a source of error variance which have to be reduced to increase reliability most effectively.

Adult↗

Appraising the quality of randomized controlled trials: inter-rater reliability for the OTseeker evidence database.

RATIONALE AND AIMS: 'OTseeker' is an online database of randomized controlled trials (RCTs) and systematic reviews relevant to occupational therapy. RCTs are critically appraised and rated for quality using the 'PEDro' scale. We aimed to investigate the inter-rater reliability of the PEDro scale before and after revising rating guidelines. METHODS: In study 1, five raters scored 100 RCTs using the original PEDro scale guidelines. In study 2, two raters scored 40 different RCTs using revised guidelines. All RCTs were randomly selected from the OTseeker database. Reliability was calculated using Kappa and intraclass correlation coefficients [ICC (model 2,1)]. RESULTS: Inter-rater reliability was 'good to excellent' in the first study (Kappas >or= 0.53; ICCs >or= 0.71). After revising the rating guidelines, the reliability levels were equivalent or higher to those previously obtained (Kappas >or= 0.53; ICCs >or= 0.89), except for the item, 'groups similar at baseline', which still had moderate reliability (Kappa = 0.53). In study 2, two PEDro scale items, which had their definitions revised, 'less than 15% dropout' and 'point measures and variability', showed higher reliability. In both studies, the PEDro items with the lowest reliability were 'groups similar at baseline' (Kappas = 0.53), 'less than 15% dropout' (Kappas <or= 0.68) and 'point measures and variability data' (Kappas <or= 0.68). CONCLUSION: The PEDro scale is a reliable instrument for rating the quality of RCTs. Revised rating guidelines are provided for scale items that are difficult to rate, and helped to improve inter-rater reliability.

Australia↗

Validity and reliability of the Turkish Migraine Disability Assessment (MIDAS) questionnaire.

OBJECTIVES: The aim of this study is to assess the comprehensibility, internal consistency, patient-physician reliability, test-retest reliability, and validity of Turkish version of Migraine Disability Assessment (MIDAS) questionnaire in patients with headache. BACKGROUND: MIDAS questionnaire has been developed by Stewart et al and shown to be reliable and valid to determine the degree of disability caused by migraine. DESIGN AND METHODS: This study was designed as a national multicenter study to demonstrate the reliability and validity of Turkish version of MIDAS questionnaire. Patients applying to 17 Neurology Clinics in Turkey were evaluated at the baseline (visit 1), week 4 (visit 2), and week 12 (visit 3) visits in terms of disease severity and comprehensibility, internal consistency, test-retest reliability, and validity of MIDAS. Since the severity of the disease has been found to change significantly at visit 2 compared to visit 1, test-retest reliability was assessed using the MIDAS scores of a subgroup of patients whose disease severity remained unchanged (up to +/-3 days difference in the number of days with headache between visits 1 and 2). RESULTS: A total of 306 patients (86.2% female, mean age: 35.0 +/- 9.8 years) were enrolled into the study. A total of 65.7%, 77.5%, 82.0% of patients reported that "they had fully understood the MIDAS questionnaire" in visits 1, 2, and 3, respectively. A highly positive correlation was found between physician and patient and the applied total MIDAS scores in all three visits (Spearman correlation coefficients were R= 0.87, 0.83, and 0.90, respectively, P <.001). Internal consistency of MIDAS was assessed using Cronbach's alpha and was found at acceptable (>0.7) or excellent (>0.8) levels in both patient and physician applied MIDAS scores, respectively. Total MIDAS score showed good test-retest reliability (R= 0.68). Both the number of days with headache and the total MIDAS scores were positively correlated at all visits with correlation coefficients between 0.47 and 0.63. There was also a moderate degree of correlation (R= 0.54) between the total MIDAS score at week 12 and the number of days with headache at visit 2 + visit 3, which quantify headache-related disability over a 3-month period similar to MIDAS questionnaire. CONCLUSION: These findings demonstrated that the Turkish translation is equivalent to the English version of MIDAS in terms of internal consistency, test-retest reliability, and validity. Physicians can reliably use the Turkish translation of the MIDAS questionnaire in defining the severity of illness and its treatment strategy when applied as a self-administered report by migraine patients themselves.

Adolescent↗

Interobserver and intraobserver reliability of venous transcranial color-coded flow velocity measurements.

BACKGROUND AND PURPOSE: Venous transcranial color-coded duplex sonography is a new technique for noninvasive evaluation of the intracranial venous system. However, the interobserver and intraobserver reliability of this method is unclear. METHODS: In 23 healthy volunteers (30 +/- 7.3 years of age), the deep middle cerebral vein (dMCV), basal vein (BV), vein of Galen (VG), and straight (SRS), transverse (TS), and superior sagittal (SSS) sinuses in addition to the arterial segments of the circle of Willis were insonated through the temporal bone window on 2 consecutive days by 2 experienced examiners. The examiners were blinded to each other's results. The interobserver and intraobserver reliability was calculated using a method described by Bland and Altman, resulting in 2-SD confidence intervals. RESULTS: Non-angle-corrected and angle-corrected systolic and end diastolic venous flow velocities (FV) were in good accordance with published normal values, ranging between 8.6 and 19.2 cm/s. The interobserver reliabilities for non-angle-corrected systolic FVs in the dMCV, BV, VG, SRS, and TS were +/- 1.8, 2.4, 2.6, 3.3, and 4.6 cm/s; for angle-corrected systolic FVs, the interobserver reliabilities were +/- 2.5, 3.1, 13.9, 11.6, and 7.7 cm/s. The intraobserver reliabilities for non-angle-corrected systolic FVs in the dMCV, BV, VG, SRS, and TS were +/- 2.9, 3.2, 2.6, 3.2, and 6.1 cm/s; for angle-corrected systolic FVs, the intraobserver reliabilities were 3.2, 3.7, 13.9, 11.6, and 7.5 cm/s. Angle correction was not attempted for the SSS. The interobserver and intraobserver reliabilities for systolic FVs in the SSS were +/- 3.3 and +/- 3.3 cm/s, respectively. CONCLUSIONS: Intracranial venous FVs can be measured with a high interobserver and intraobserver reliability in healthy human subjects. Intraobserver reliability was higher for cerebral veins than for dural sinuses, predisposing them for follow-up examinations; however, angle correction for venous FVs in the VG and the SRS is not advisable.

Adult↗

Influence of ionic conductances on spike timing reliability of cortical neurons for suprathreshold rhythmic inputs.

Spike timing reliability of neuronal responses depends on the frequency content of the input. We investigate how intrinsic properties of cortical neurons affect spike timing reliability in response to rhythmic inputs of suprathreshold mean. Analyzing reliability of conductance-based cortical model neurons on the basis of a correlation measure, we show two aspects of how ionic conductances influence spike timing reliability. First, they set the preferred frequency for spike timing reliability, which in accordance with the resonance effect of spike timing reliability is well approximated by the firing rate of a neuron in response to the DC component in the input. We demonstrate that a slow potassium current can modulate the spike timing frequency preference over a broad range of frequencies. This result is confirmed experimentally by dynamic-clamp recordings from rat prefrontal cortical neurons in vitro. Second, we provide evidence that ionic conductances also influence spike timing beyond changes in preferred frequency. Cells with the same DC firing rate exhibit more reliable spike timing at the preferred frequency and its harmonics if the slow potassium current is larger and its kinetics are faster, whereas a larger persistent sodium current impairs reliability. We predict that potassium channels are an efficient target for neuromodulators that can tune spike timing reliability to a given rhythmic input.

Animals↗

Assessment of reliability of lung function screening programs or longitudinal studies.

The aim was to determine reliability of lung function measurements performed according to recommendations of the American Thoracic Society (ATS) at a screening program in a large South African gold mine and to determine the usefulness of the reliability coefficient G for monitoring the reliability of lung function measurements in a mass screening program. The reliability coefficient G estimates the amount of random error of measurement, relative to the total variation in a measurement. The coefficient G was calculated as a correlation coefficient between two consecutive lung function tests performed within 6 mo, over a period of 43 mo on 3,378 miners. There was significant temporal variability in the reliability. For FEV(1), the coefficient G showed increased variability over the first 5 mo and stabilized at a value of 0.93 for the next 23 mo, after which it systematically declined over the next 15 mo. We estimated that in a large screening program, an optimal sample size of around 900 miners, examined randomly throughout the year, on a yearly basis, would provide a sufficient sample to examine monthly or quarterly fluctuation in the reliability. The value of the reliability coefficient G did not change when the time between two consecutive tests increased up to 15 mo. In conclusion, monitoring of lung function reliability in a screening program by the reliability coefficient G should improve data quality, and provide a measure on which the confidence in a decision-making process could be based when examining temporal changes in lung function for individual subjects.

Adult↗

Assessment of the perceptual threshold of touch (PTT) with high-frequency transcutaneous electric nerve stimulation (Hf/TENS) in elderly patients with stroke: a reliability study.

OBJECTIVE: To evaluate the inter-rater reliability and reliability between occasions of assessing the perceptual threshold of touch (PTT) with high-frequency transcutaneous electric nerve stimulation (Hf/TENS) in elderly patients with stroke. DESIGN: A test-retest study of reliability using intraclass correlation coefficient (ICC) and limits of agreement. SETTING: Geriatric rehabilitation unit. SUBJECTS: Thirty-two consecutive patients with stroke > or = 65 years of age. MAIN OUTCOME MEASURES: Two-channel current stimulator TENS CEFAR Tempo with four self-adhesive skin electrodes. The stimulator delivered a high-frequency constant current of 40 Hz. The strength of the stimulation was quantifiable and assessed in milliampere (mA). INTERVENTIONS: The assessments were performed on the hands and feet by two raters. The PTT was identified as the level registered in milliampere (mA) at which the patients perceived a tingling sensation. RESULTS: The ICC values (0.94-0.99) were shown to be good for inter-rater reliability, as well as reliability between occasions. However an additional analysis with limits of agreement showed a high level of agreement for assessment of the hand but a moderate to low agreement for assessment of the foot where some bias was also identified. Clinical acceptable reliability: > or = 1 mA for the hand and > or = 5 mA for the foot are so far recommended for establishing real differences in clinical measures. CONCLUSION: Hf/TENS shows an overall high reliability for assessing the PTT of the hand and moderate to low reliability for the foot. Additional research with exclusion of bias is needed to determine the reliability of assessing the foot.

Aged↗

Validity and reliability of the neonatal skin condition score.

OBJECTIVE: To demonstrate the validity and reliability of the Neonatal Skin Condition Scale (NSCS) used in the Association of Women's Health, Obstetric and Neonatal Nurses (AWHONN) and the National Association of Neonatal Nurses (NANN) neonatal skin care evidence-based practice project. SETTING: NICU and well-baby units in 27 hospitals located throughout the United States. PARTICIPANTS: Site coordinators (N = 27) and neonates (N = 1,006) observed during both the pre and postimplementation phases of the original neonatal skin care project. METHOD: To assess reliability, two consecutive NSCS assessments on a single infant were analyzed. Site coordinators were contacted after the original project was concluded. Sites indicating that a single nurse scored all infant skin observations provided data that were used to evaluate intrarater reliability. Sites using more than one nurse to score skin observations provided data that were used to assess interrater reliability. To assess validity, the following variables were used from the original data set: the Neonatal Skin Condition Scale (NSCS), with three subscales for dryness, erythema, and breakdown; birth weight in grams; number of skin score observations for each infant; and the prevalence of infection, defined as a positive blood culture. RESULTS: For intrarater reliability, 16 sites used a single nurse for all NSCS assessments; total NSCS assessments 475. For interrater reliability, 11 sites used multiple raters; total assessments 531. The NSCS demonstrated adequate reliability for each of the three subscales and for the total score, with the percent agreement between scores ranging from 68.7% to 85.4% (intrarater) and 65.9% to 89% (interrater); all Kappas were significant at p < .001 and were in the moderate range for reliability. The validity of the NSCS was demonstrated by the findings that smaller infants were 6 times more likely to have erythema (chi2(6) = 109.55, p < .0001), and approximately twice as likely to have the most severe breakdown (chi2(6) = 108.01, p < .0001). Infants with more observations (longer length of stay) had higher skin scores (odds ratio = 1.21, p < .0001), and an increased probability of infection was noted for infants with higher skin scores (odds ratio = 2.25, p < .0001). CONCLUSIONS: The Neonatal Skin Condition Score (NSCS) is reliable when used by single and multiple raters to assess neonatal skin condition, even across weight groups and racial groups. Validity of the NSCS was demonstrated by confirmation of the relationship of the skin condition scores with birth weight, number of observations, and prevalence of infection.

Evidence-Based Medicine↗