PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Reliability of the knee examination in osteoarthritis: effect of standardization.

OBJECTIVE: To assess the reliability of physical examination of the osteoarthritic (OA) knee by rheumatologists, and to evaluate the benefits of standardization. METHODS: Forty-two physical signs and techniques were evaluated using a 6 x 6 Latin square design. Patients with mild to severe knee OA, based on physical and radiographic signs, were examined in random order prior to and following standardization of techniques. For those signs with dichotomous scales, agreement among the rheumatologists was calculated as the prevalence-adjusted bias-adjusted kappa (PABAK), while for the signs with continuous and ordinal scales, a reliability coefficient (R(c)) was calculated using analysis of variance. A PABAK of >0.60 and an R(c) of >0.80 were considered to indicate adequate reliability. RESULTS: Adequate poststandardization reliability was achieved for 30 of 42 physical signs/techniques (71%). The most highly reliable signs identified by physical examination of the OA knee included alignment by goniometer (R(c) = 0.99), bony swelling (R(c) = 0.97), general passive crepitus (R(c) = 0.96), gait by inspection (PABAK = 0.78), effusion bulge sign (R(c) = 0.97), quadriceps atrophy (R(c) = 0.97), medial tibiofemoral tenderness (R(c) = 0.94), lateral tibiofemoral tenderness (R(c) = 0.85), patellofemoral tenderness by grind test (R(c) = 0.94), and flexion contracture (R(c) = 0.95). The standardization process resulted in substantial improvements in reliability for evaluation of a number of physical signs, although for some signs, minimal or no effect of standardization was noted. After standardization, warmth (PABAK = 0.14), medial instability at 30 degrees flexion (PABAK = 0.02), and lateral instability at 30 degrees flexion (PABAK = 0.34) were the only 3 signs that were highly unreliable. CONCLUSION: With the exception of physical examinations for instability, a comprehensive knee examination can be performed with adequate reliability. Standardization further improves the reliability for some physical signs and techniques. The application of these findings to future OA studies will contribute to improved outcome assessments in OA.

Adult↗

Test-retest reliability of the unified Parkinson's disease rating scale in patients with early Parkinson's disease: results from a multicenter clinical trial.

Our objective was to assess the test-retest reliability of the Unified Parkinson's Disease Rating Scale (UPDRS). The UPDRS is the most widely used instrument for measuring severity of parkinsonian symptoms in clinical research and in practice. The validity and inter-rater reliability of this scale have been previously studied. We examined the test-retest (intrarater) reliability of the UPDRS and derived subscales. Four hundred patients with early-stage Parkinson's disease (PD) who were participating in a multicenter clinical trial were evaluated using the UPDRS on two separate occasions (screening and baseline visits) prior to receiving treatment. The same neurologist at each center rated the subjects at both examinations that were, on average, 14.6 +/- 7.6 days apart (range 3-36 days). Test-retest reliability was estimated using the intraclass correlation coefficient (ICC) for the total UPDRS score, the mental, ADL, and motor subscale scores, and other derived subscale scores. Weighted kappa statistics were calculated for individual UPDRS items. The ICCs for the UPDRS scores were as follows: total score, 0.92; mental, 0.74; ADL, 0.85; motor, 0.90. ICCs for derived symptom-based scales ranged from 0.69-0.88. Reliability of specific items was generally lower than for summary scales. Reliability was slightly better in patients for whom the testing interval was within 14 days. Based on conventional standards, the UPDRS scores were found to have excellent test-retest reliability in this sample of patients with early PD rated by academic movement disorder specialists. The findings are in agreement with previous reports on interrater reliability.

Activities of Daily Living↗

Inter- and intrajudge reliability for videofluoroscopic swallowing evaluation measures.

Interjudge reliability for videofluoroscopic (VFS) swallowing evaluations has been investigated, and results have, for the most part, indicated that reliability is poor. While previous studies are well-designed investigations of interjudge reliability, few reports of intrajudge reliability are available for VFS measures derived from frame-by-frame analysis that clinicians typically employ. The purpose of this study was to examine the inter- and intrajudge reliability of VFS examination measures commonly used to assess swallowing functions. No training to criteria occurred. VFS examinations were conducted on 20 patients who had suffered a stroke within six weeks and had no structural abnormalities or tracheostomies. Three clinical judges served as subjects and rated the VFS examinations from videotape using frame-by-frame analysis. A clinician's repeated review of measures employed in the 20 examinations indicated high intrajudge reliability for a number of measures, suggesting that an experienced clinician may employ consistent standards for rating certain VFS measures across patients and time. These standards appear to vary among clinicians and yield unacceptable interjudge reliability. The need to train clinicians to criteria to improve interjudge reliability is discussed.

Adult↗

The comparison of reliabilities in dental imaging methods.

OBJECTIVES: Common practice in the statistical comparison of imaging instruments with limited reproducibility consists in the separate estimation of the instrument's reliabilities. However, as soon as one of the imaging methods is subject to item-specific bias (which has to be expected in many dentomaxillofacial imaging procedures), this approach will end in severe errors in reliability computation and in corresponding erroneous clinical conclusions. This paper seeks to point out these effects and to illustrate a more appropriate model for the comparison of instrumental reliabilities. METHODS: A standard reliability model was adjusted for item-specific bias and illustrated by the comparison of twice repeated planimetric cephalometry versus twice repeated noninvasive orthodontic video imaging (based on the Digigraph 100 device) in 50 children; the anterior cranial base length was used for illustration. RESULTS: The proposed model revealed pronounced inferiority of the video-based imaging system concerning its reliability compared with the X-ray based standard. Analysis using separate estimation of the two reliabilities would result in the reverse conclusion and thus falsely establish video imaging, which is in fact less reliable, as a superior diagnostic method. CONCLUSION: The reliabilities of dentomaxillofacial imaging methods have to be adjusted for potential item-specific bias to avoid the erroneous conclusion of the superiority of a diagnostic innovation.

Adolescent↗

The reliability of isokinetic testing of the ankle joint and a heel-raise test for endurance.

The aim of the present study was to investigate the reliability of different methods used for isokinetic testing of calf muscle strength and endurance. The detailed evaluation of test-retest reliability serves the purpose of establishing reliable research tools when evaluating patients who have sustained an Achilles tendon rupture. The test-retest reliability of isokinetic measurements at the ankle for eccentric and concentric muscle action was calculated in ten healthy male volunteers using intra-class correlation (ICC) and coefficient of variation (CV). Three different positions were compared at the angular velocities of 30 degrees /s and 180 degrees /s for right and left ankles. The ICC for plantar flexion was 0.37-0.95, whilst it was 0.00-0.96 for dorsiflexion. The corresponding CVs were 4.0-19.9 and 2.4-19.8 respectively. The test-retest reliability of standardised heel-raises, Achilles tendon width, calf circumference and ankle range of motion revealed ICC values of 0.71-0.98 and CVs of 0.67-19.1. The test-retest interval was 5 to 7 days. We conclude that all three positions studied for the isokinetic evaluation of calf muscle function are equally reliable concerning plantar flexion at the ankle joint. The same level of reliability was also found in the evaluation of the standing heel-raise test and the isokinetic dorsiflexion test, except for dorsiflexion in the supine position. The reliability of the investigated methods was only fair despite the use of a detailed and standardised test protocol.

Achilles Tendon↗

Physical exposure of sign language interpreters: baseline measures and reliability analysis.

Measurement of physical exposure to musculoskeletal disorder risk factors must generally be performed directly in the field to assess the effectiveness of ergonomic interventions. To perform such an evaluation, the reliability of physical exposure measures under similar field conditions must be known. The objectives of this study were to estimate the reliability of physical exposure measures performed in the field and to establish the baseline values of physical exposure in sign language interpreters (SLI) before the implementation of an intervention. The electromyography (EMG) of the trapezius muscles as well as the wrist motions of the dominant arm were measured using goniometry on nine SLI on four different days. Several exposure parameters, proposed in the literature, were computed and the generalizability theory was used as a framework to assess reliability. Overall, SLI showed a relatively low level of trapezius muscle activity, but with little time at rest, and highly dynamic wrist motions. Electromyography exposure parameters showed poor to moderate reliability, while goniometry parameter reliability was moderate to excellent. For EMG parameters, performing repeated measurements on different days was more effective in increasing reliability than extending the duration of the measurement over one day. For goniometry, repeating measurements on different days was also effective in improving reliability, although good reliability could be obtained with a single sufficiently long measurement period.

Adult↗

Interrater reliability of videofluoroscopic swallow evaluation.

The past two decades have brought an enormous widening of interest in and knowledge about swallowing disorders. The most frequently used technique for swallow evaluation is X-ray videofluoroscopy. Most interventions are based on this examination. Only a few studies assessing interobserver reliability of videofluoroscopy have been published. The aim of our study was to assess the interobserver reliability of videofluoroscopy for swallow evaluation. Fifty-one consecutive dysphagic patients referred for videofluoroscopy were entered into the study regardless of their underlying disorder. The first swallow (5 ml of a semisolid radio-opague contrast media) of each patient was assessed in the lateral projection by 9 independent, experienced observers from different international swallow centers. All studies were evaluated according to a standardized protocol sheet and the interobserver reliability was calculated. The interobserver reliabilities assessed as kappa coefficient for parameters of the oral and pharyngeal phase, for the temporal occurrence of penetration/aspiration, and for the location of bolus residue ranged from 0.01 to 0.56. High reliability with an intraclass coefficient of 0.80 was achieved only with the well defined penetration/aspiration score. Our study underlines the need for exact definitions of the parameters assessed by videofluoroscopy, in order to raise interobserver reliability. To date, only aspiration is evaluated with high reliability by videofluoroscopy, whereas the reliability of all other parameters of oropharyngeal swallow is poor.

Adult↗

A new skin-surface device for measuring the curvature and global and segmental ranges of motion of the spine: reliability of measurements and comparison with data reviewed from the literature.

There is an increasing awareness of the risks and dangers of exposure to radiation associated with repeated radiographic assessment of spinal curvature and spinal movements. As such, attempts are continuously being made to develop skin-surface devices for use in examining the progression and response to treatment of various spinal disorders. However, the reliability and validity of measurements recorded with such devices must be established before they can be recommended for use in the research or clinical environment. The aim of this study was to examine the reliability of measurements using a newly developed skin-surface device, the Spinal Mouse. Twenty healthy volunteers (mean age 41 +/- 12 years, nine males, 11 females) took part. On 2 separate days, spinal curvature was measured with the Spinal Mouse during standing, full flexion, and full extension (each three times by each of two examiners). Paired t-tests, intraclass correlation coefficients (ICC), and standard errors of measurement (SEM) with 95% confidence intervals were used to characterise between-day and interexaminer reliability for: standing sacral angle, lumbar lordosis, thoracic kyphosis, and ranges of motion (flexion, extension) of the thoracic spine, lumbar spine, hips, and trunk. The between-day reliability for segmental ranges of flexion was also determined for each motion segment from T1-2 to L5-S1. The majority of parameters measured for the 'global regions' (thoracic, lumbar, or hips) showed good between-day reliability. Depending on the parameter of interest, between-day ICCs ranged from 0.67 to 0.92 for examiner 1 (average 0.82) and 0.57 to 0.95 for examiner 2 (average 0.83); for 70% of the parameters measured, the ICCs were greater than 0.8 and generally highest for the lumbar spine and whole trunk measures. For lumbar spine range of flexion, the SEM was approximately 3 degrees. The ICCs were also good for the interexaminer comparisons, ranging from 0.62 to 0.93 on day 1 (average 0.81) and 0.70 to 0.94 on day 2 (average 0.86), although small systematic differences were sometimes observed in their mean values. The latter were still evident even if both examiners used the same skin markings. For segmental ranges of flexion, the ICCs varied between vertebral levels but overall were lower than for the global measures (average for all levels in all analyses, ICC 0.6). For each examiner, the average between-day SEM over all vertebral levels was approximately 2 degrees. For 'global' regions of the spine, the Spinal Mouse delivered consistently reliable values for standing curvatures and ranges of motion which compared well with those reported in the literature. This suggests that the device can be reliably implemented for in vivo studies of the sagittal profile and range of motion of the spine. As might be expected for the smaller angles being measured, the segmental ranges of flexion showed lower reliability. Their usefulness with regard to the interpretation of individual results and the detection of 'real change' on an individual basis thus remains questionable. Nonetheless, the group mean values showed few between-day differences, suggesting that the device may still be of use in providing clinically interesting data on segmental motion when examining groups of individuals with a given spinal pathology or undergoing some type of intervention.

Adult↗

Reliability of traditional and fractal dimension measures of quiet stance center of pressure in young, healthy people.

OBJECTIVES: To assess reliability of traditional and fractal dimension measures of quiet stance center of pressure (COP). DESIGN: Cross-sectional study. SETTING: University laboratory. PARTICIPANTS: Thirty young healthy men (n=20) and women (n=10) (mean age, 23 y). INTERVENTIONS: Not applicable. MAIN OUTCOME MEASURES: COP was recorded for 3 trials across 4 conditions: eyes open and eyes closed standing on firm and foam surfaces. Traditional COP variables--peak sway velocity and range of sway, both in the anteroposterior (AP) and mediolateral (ML) directions, and total excursion area, and fractal dimension of the COP in the AP and ML directions--were calculated. Reliability statistics were calculated. RESULTS: Range of sway (AP) was the most reliable traditional variable (intraclass correlation coefficient model 2,1 [ICC(2,1)] range -.28 to .72.). Peak sway velocity (AP) had poorest reliability (ICC(2,1) range, .05-.29). Only 1 of the traditional variables had excellent reliability; total excursion area (firm, eyes closed) (ICC(2,1)=.95). All bar 1 fractal dimension measures had excellent ICCs. Relative technical error of measurement ranged from 4% to 7% for the fractal dimension measures. Coefficients of variation were also very good, ranging from 1.8% to 6.7%. CONCLUSIONS: Fractal dimension measures were more reliable than traditional measures of COP. Although traditional measures are used extensively to assess COP, their reliability is questionable. Fractal dimension measures show promise to reliably quantify COP and warrant further investigation.

Adult↗

Interrater reliability of the history and physical examination in patients with mechanical neck pain.

OBJECTIVE: To examine the interrater reliability of the history and physical examination in patients with mechanical neck pain. DESIGN: Single-group repeated measures for interrater reliability. SETTING: Outpatient physical therapy clinic. PARTICIPANTS: Twenty-two patients with mechanical neck pain underwent a standardized history and physical examination by a physical therapist. INTERVENTION: Following a 5-minute break, a second therapist who was blind to the findings of examiner 1 performed the second standardized history and physical examination. MAIN OUTCOME MEASURES: The Cohen kappa and weighted kappa were used to calculate the interrater reliability of ordinal level data from the history and physical examination. Intraclass correlation coefficients model 2,1 (ICC(2,1)) and the 95% confidence intervals were calculated to determine the interrater reliability for continuous variables. RESULTS: The kappa coefficients ranged from -.06 to .90 for the variables obtained from the history. Reliability values for categorical data collected during the physical examination ranged from no to substantial agreement depending on the particular test and measure. ICC(2,1) for cervical range of motion (ROM) measurements ranged between .66 and .78. CONCLUSIONS: We have reported the interrater reliability of the history and physical examination in a group of patients with a primary report of neck pain. The reliability variables varied considerably for manual assessment techniques and were significantly higher for the examination of muscle length and cervical ROM. Ultimately, it will be up to each clinician to determine if a particular test or measure poses adequate reliability to assist in the clinical decision making process.

Adult↗

Reliability and validity of the pain observation scale for young children and the visual analogue scale in children with burns.

The aim of this study was to assess if the pain observation scale for young children (POCIS) and the visual analogue scale (VAS) are reliable and valid instruments to measure procedural and background pain in burned children aged 0-4 years. Burn care nurses (n=73) rated pain from 24 fragments of videotaped children during wound care procedures and during periods of rest using the POCIS and the VAS. Intraclass correlations were used to assess inter-rater and intra-rater reliability for the POCIS and the VAS. Internal consistency for POCIS was assessed by Cronbach's alpha. The POCIS has shown poor to moderate inter-rater reliability, moderate to good intra-rater reliability and an acceptable internal consistency. The VAS turned out to have poor inter-rater reliability and poor to moderate intra-rater reliability. Due to poor results of inter-rater reliability in both scales, construct validation is left undone until more acceptable results are obtained. Factors explaining the results are the large number of raters, the manner they were trained and a lack of variation between pain classes in video fragments. Although not all results were satisfying, an easy to use scale as POCIS has promising qualities and deserves further reliability research.

Adolescent↗

Interrater reliability of measurements of comorbid illness should be reported.

OBJECTIVE: Comorbidity indices are commonly used to stratify patients to control for treatment selection bias. The objectives here were to review the reporting of interrater reliability when studies use comorbidity indices in clinical research publications and to report the interrater reliability of four common indices in a particular research setting. STUDY DESIGN AND SETTING: Four trained abstractors reviewed the same 40 charts of patients with squamous cell carcinoma of the head and neck from a regional cancer center. Scores for the Charlson Index, the Index of Co-existent Disease, the Cumulative Illness Rating Scale, and the Kaplan-Feinstein Classification were calculated, and the intraclass correlation coefficient was used to assess interrater reliability. RESULTS: The details on the training of abstractors and the results of interrater reliability tests are not commonly reported. In our study setting, the Charlson Index had excellent reliability and the others had acceptable reliability. CONCLUSION: If the quality of a study using an index or scale is to be assessed, the reliability and interrater reliability of the score assignment process should be reported.

Carcinoma, Squamous Cell↗

Optical biometry of the anterior eye segment: interexaminer and intraexaminer reliability of ACMaster.

PURPOSE: To evaluate the interexaminer and intraexaminer reliability of corneal thickness, anterior chamber depth (ACD), and crystalline lens thickness measurements using a commercially available anterior segment optical biometry instrument (ACMaster, Carl Zeiss Meditec) based on partial coherence interferometry (PCI). SETTING: Medical University of Vienna, Vienna, Austria. METHODS: Interexaminer reliability and intraexaminer reliability were evaluated in 10 eyes of 10 young volunteers and 11 eyes of 11 cataract patients. The measurements of the interexaminer reliability were taken by 3 examiners. Corneal thickness, ACD, and lens thickness of the intraexaminer reliability were measured twice in all eyes by 1 examiner. To evaluate the effect of cycloplegia on the variability, the measurements were performed on 5 volunteers under cyclopentolate 1%. Measurements were performed using the prototype of the ACMaster based on PCI. RESULTS: The interexaminer/intraexaminer reliabilities were 99.9% for corneal thickness and ACD. The reliability of lens thickness could not be estimated because of a large number of missing values in the cataract patient group. The median interexaminer variability (SD) was 1.9 microm for corneal thickness, 7.5 microm for ACD, and 10.6 microm for lens thickness. The median intraexaminer variability (SD) was 1.6 microm for corneal thickness, 10.8 microm for ACD, and 8.7 microm for lens thickness. With cycloplegia, both the interexaminer variability and intraexaminer variability were smaller than without cycloplegia. CONCLUSIONS: Partial coherence interferometry measurements of anterior chamber distances (corneal thickness, ACD, lens thickness) using the prototype of ACMaster were highly reliable, allowing examiner-independent measurements. However, lens thickness measurements in cataract eyes were often difficult.

Adult↗

The modulatory effects of nicotine on parietal cortex activity in a cued target detection task depend on cue reliability.

This functional magnetic resonance imaging study investigates the effects of nicotine in a cued target detection task when changing cue reliability. Fifteen non-smoking volunteers were studied under placebo and nicotine (Nicorette polacrilex gum 1 and 2 mg). Validly and invalidly cued trials were arranged in blocks with high, middle and low cue reliability. Two effects of nicotine were investigated: its influence on i) parietal cortex activity underlying the processing of invalid vs. valid trials (i.e. validity effect) and ii) neural activity in the context of low, middle and high informative value of the cue (i.e. cue reliability effect). Nicotine did not affect behavioral performance. However, nicotine reduced the difference in the blood oxygenation level dependent (BOLD) signal between invalid and valid trials in the right intraparietal sulcus. The reduction of parietal activity in invalid trials was smaller in the low cue reliability condition. The same posterior parietal region exhibited a nicotinic modulation of BOLD activity in valid trials which was dependent on cue reliability: Nicotine specifically enhanced the neural activity during valid trials in the context of low cue reliability, i.e. when subjects are already in a state of low certainty. We speculate that the right intraparietal sulcus might be part of two networks working in parallel: one responsible for reorienting attention and the other for the cholinergic modulation of cue reliability. By reducing the use of the cue, nicotine modulates parietal activity related to reorienting attention in conditions with higher cue certainty. On the other hand, nicotine increases parietal activity in states of low certainty. This enhanced activation might influence brain regions, such as the posterior cingulate, directly involved in the processing of cue reliability.

Adult↗

Psychometric properties of ADHD rating scales among children with mental retardation I: reliability.

The reliability of Attention-Deficit/Hyperactivity Disorder (ADHD) rating scales in children with mental retardation was assessed. Parents, teachers, and teaching assistants completed ADHD rating scales on 48 children aged 5-12 diagnosed with mental retardation. Measures included the Child Behavior Checklist (CBCL), Conners Rating Scales, the Attention-Deficit/Hyperactivity Disorder Test (ADHDT), the Swanson, Nolan, and Pelham (SNAP) Checklist, the Werry-Weiss-Peters Activity Rating Scale (WWPARS), the ADD-H Comprehensive Teacher's Rating Scale (ACTeRS), and the Aberrant Behavior Checklist-Community (ABC-C). The internal consistency, test-retest, and interrater reliability of each scale was examined. Results showed best support for teacher completed scales, followed by ratings made by teaching assistants, and parent-report scales. Strong support for the internal consistency of the teacher-report measures was found, and it was quite similar to previously reported internal consistencies with typically developing children. Test-retest reliabilities of the teacher report measures were also quite good but tended to be lower than those reported for typically developing children. For teaching assistant ratings, test-retest reliabilities were adequate to very good. The internal consistency reliabilities for parent completed measures were adequate to excellent, but test-retest reliabilities were low. Interrater reliability was best for teacher-teaching assistants. The ABC-C was the only measure on which the interrater reliability was adequate for clinical purposes.

Attention Deficit Disorder with Hyperactivity↗

The interobserver reliability of pretest probability assessment in patients with suspected pulmonary embolism.

INTRODUCTION: Pretest probability assessment and objective testing are combined to appropriately manage patients with suspected pulmonary embolism (PE). However, the interobserver reliability of pretest probability assessment has not been investigated. We sought to determine (for patients with suspected PE) the interobserver reliability of pretest probability assessment (by overall impression (gestalt) versus an explicit clinical model). MATERIALS AND METHODS: A prospective cohort study was conducted at an urban university hospital. For patients referred for ventilation and perfusion (V/Q) scanning for suspected PE, structured assessments (11 history and 4 physical examination parameters) were performed by a referring physician and a designated thrombosis physician. The referring and thrombosis physicians also assigned a pretest probability for PE (low, moderate, or high) by gestalt. An explicit seven-point clinical model for suspected PE was later applied to each structured assessment to determine the pretest probability. Assessments were performed independently and prior to diagnostic test results. Interobserver reliability (two rater unweighted Kappa (kappa) statistic) was determined for each parameter on the structured assessment and the pretest probability assessments (gestalt vs. explicit clinical model). RESULTS: One hundred and ten patients with suspected PE received duplicate assessments. Historical features demonstrated substantial to almost perfect interobserver reliability (kappa=0.60-0.95). For the physical findings, only heart rate had substantial interobserver reliability (kappa=0.60). Pretest probability assessment was not reliable when using physician's gestalt (kappa=0.33), but produced substantial interobserver reliability using the explicit clinical model (kappa=0.62). CONCLUSIONS: Given the inadequate interobserver reliability of pretest probability assessment by overall impression (or gestalt), physicians should use explicit clinical models in the diagnostic management of patients with suspected pulmonary embolism.

Adolescent↗

Reliability of remembered International Index of Erectile Function domain scores in men with localized prostate cancer.

OBJECTIVES: To test the reliability of recollected International Index of Erectile Function (IIEF) domain scores before and after radical prostatectomy. Recall reliability can be affected by several biases. In men with localized prostate cancer (PCa), conflicting results have been reported. METHODS: Thirty-nine men, aged 44 to 69 years, were invited to participate in a prospectively administered IIEF questionnaire. The survey was administered before and 6 and 12 months after radical prostatectomy. Several months later, a recall IIEF survey targeted the prospectively gathered IIEF data. The independent sample t test, Pearson correlation coefficient, partial correlation, and intraclass correlation coefficient tested the reliability of the recalled IIEF scores versus the prospective ratings. RESULTS: All 39 men completed the prospective and recalled IIEF surveys addressing preoperative erectile function. Surveys targeting function at 6 and 12 months after surgery were completed by 85% and 51% of the participants, respectively. The erectile function domain demonstrated the greatest recall reliability (intraclass correlation coefficient 0.65 to 0.73). Erectile function and sexual desire scale recall reliability was greatest for pretreatment function or function 12 months after surgery. The orgasmic function domain had the lowest recall reliability (intraclass correlation coefficient 0.37 to 0.54). CONCLUSIONS: When restricted to before surgery and 12 months after surgery, most IIEF domains may be reliably used in a retrospective fashion. The erectile function and sexual desire domains appear to be most reliable, possibly because they address more objective areas of men's sexual function.

Adenocarcinoma↗

Reliability of treadmill exercise testing in older patients with chronic hemiparetic stroke.

OBJECTIVE: To assess the test-retest reliability of cardiopulmonary measurements during peak effort and submaximal treadmill walking tests in older patients with gait-impaired chronic hemiparetic stroke. DESIGN: Nonrandomized test-retest. SETTING: Hospital geriatric research stress testing laboratory. PARTICIPANTS: Fifty-three subjects (44 men, 9 women; mean age, 65+/-8y) with chronic hemiparetic gait after remote (>6mo) ischemic stroke. Patients had mild to moderate chronic hemiparetic gait deficits, making handrail support necessary during treadmill walking. INTERVENTIONS: Peak effort and submaximal effort treadmill walking tests were conducted and then repeated on a separate day at least a week later. Main outcome measures Reliability coefficients (r) were calculated for heart rate, systolic blood pressure (SBP), oxygen consumption (Vo(2) [L/min]), Vo(2) (mL.kg(-1).min(-1)), respiratory exchange ratio (RER), rate-pressure product (RPP), and oxygen pulse during peak effort testing. The reliability coefficients for all but SBP and RPP data were calculated from the submaximal tests. RESULTS: Heart rate (r=.87), Vo(2)peak (L/min) (r=.92), Vo(2)peak (mL.kg(-1).min(-1)) (r=.92), and oxygen pulse (r=93) were highly reliable parameters during maximal testing in this population. Submaximal testing produced highly reliable results for V.o(2) (L/min) (r=.89) and oxygen pulse (r=.85). All cardiopulmonary measures except RER had a reliability coefficient greater than.80 during submaximal testing in this population. CONCLUSION: Our study provides the first evidence that peak effort treadmill testing provides highly reliable oxygen consumption measures in chronic hemiparetic stroke patients using minimal handrail support. The submaximal tests were at or near the threshold level of reliability for the 2 most important measures of V.o(2) (L/min) and V.o(2) (mL.kg(-1).min(-1)) (r=.89, r=.84, respectively), with the remaining measures falling above.70.

Adult↗