PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

The reliability of H-reflex recordings in standing subjects.

OBJECTIVE: Several studies have used the H-reflex to investigate the effect of upright stances and locomotion on spinal reflex excitability. The reliability of eliciting this reflex during weight-bearing has however yet to be addressed. This study was undertaken to determine the reliability of individual differences in the H-reflex recorded from healthy subjects during quiet standing. Secondary aims of the study were to evaluate individual reliability during prolonged standing, and to establish the minimum number of trials required to provide reliable measurements of the H-reflex. DESIGN AND PARTICIPANTS: Twenty neurologically healthy volunteers participated in a repeated measures design consisting of 8 blocks of 20 trials evenly distributed over two testing sessions. H-reflex recordings were elicited from the subject's dominant side soleus muscle by percutaneously stimulating the posterior tibial nerve. Stimuli were presented every 10 s, with 2 min seated rest provided between blocks of trials. Peak-to-peak amplitude of H-reflexes and m-responses determined for individual trials were used for subsequent analysis. RESULTS: It was found that the reliability of measuring the H-reflex and m-response during quiet standing was extremely robust (r = .97 for both measures). This pattern of individual differences remained consistent over 80 trials confirming the stability of the measures. High reliabilities (r = .96 and .87 for the H-reflex and m-responses respectively) were also observed when as few as four trials were analysed. When measures obtained during the first session of testing were compared with those obtained for session two, the correlation coefficients were generally of a lower order (r = .54 to .90). DISCUSSION: The results demonstrate that the H-reflex in quiet standing provides high intra-individual reliability, suggestive of a stable reflex resistant to potentially confounding postural influences, or other sources of biological variation. The between-session reliability underscores the difficulty in reproducing conditions between sessions, and emphasises the need for within-session comparisons of H-reflex amplitudes. Given the functional challenge of maintaining an upright posture, the H-reflex appears to be a well maintained and stable phenomenon.

Adolescent↗

Test-retest reliability of the soleus H-reflex in three different positions.

PURPOSE: H-reflex has been clinically useful in the diagnosis of radiculopathies, developmental disorders, and measurement of motoneuron excitability. However, variability of the H-reflex precluded its routine application. The purpose of this study is to evaluate the test-retest and within-subject reliability of the soleus H-reflex tested in three different positions. SUBJECTS: Seven men and eight women healthy volunteered (20-50 y) with no history of significant low back pain or radiculopathy consented to the study. METHODS: The soleus H-reflexes for both lower extremities were elicited and recorded using Cadwell 5200-A EMG unit and surface recording. The tibial nerve was electrically stimulated at the popliteal fossa using 0-5 ms., 0.2 pps pulses at intensity equivalent to H-max. Each subject was tested randomly in three different positions: pronelying, free standing, and standing while lifting 20% of his/her body weight. Signal were amplified (1-5 K) using surface electrodes applied on the soleus muscle at midline and 3 cm below the gastrocnemius musculotendinous junction. The peak-to-peak amplitude and onset latencies of four separate traces were averaged for each trial. Subjects were re-tested within 10 days by the same tester following the same protocol. RESULTS: Test-retest reliability of the H-reflex amplitude ranged from r = .29 in prone position to r = .56 in the loading position. Within day reliability of the H-amplitude was high between the three different positions and ranged from r = .56 to r = .97. The test-retest reliability of the H-latency were extremely high and robust, with the coefficients ranged from r = .92 to r = .94. Also the within day reliability of the H-latency ranged from r = .96 to r = .99. CONCLUSIONS: Results indicated that, when the H-amplitude is the measure of choice, testing the H-reflex in standing and loading positions is more reliable than testing in pronelying. Also testing the subject during various procedures in the same session is more reliable than testing subject in different days/sessions. The H-latency is highly reliable in all three testing positions.

Adult↗

Consultation competence in general practice: testing the reliability of the Leicester assessment package.

BACKGROUND: An acceptable assessment must be both valid and reliable; the face validity of the Leicester assessment package has already been established. AIM: This study set out to test the reliability of the Leicester assessment package, and the factors influencing it, when used by multiple assessors to assess performance in general practice consultations. METHOD: Six randomly selected course organizer assessors simultaneously used the package to conduct independent assessments of the performance of five doctors of widely varying abilities in consultation with six simulated patients. The scores allocated were subjected to generalizability analysis. RESULTS: The mean scores allocated for consultation performance of individual doctors ranged from 51% to 70%, with the lower scores being allocated to the less experienced doctors. Scores of each assessor across the cases were examined for internal consistency and five of the six assessors consistently scored the doctors with an alpha coefficient of the minimum accepted level of 0.80 or greater. The other assessor had a consistency of only 0.22. Measurements of consistency within cases between markers indicated that the first case produced unreliable results (alpha coefficient 0.25) but all other cases were scored consistently. Two independent assessors scoring eight consultations are the requisite numbers to achieve acceptable levels of reliability in a formal assessment process; seven consultations produce the minimum acceptable generalizability coefficient of 0.80 plus the first 'non-counting' consultation. CONCLUSION: Required levels of reliability can be achieved when the package is used by multiple markers assessing the same consultations over a wide range of consultation performance. To achieve reliability only two hours of assessment time are required using the Leicester package compared with the previously regarded minimum of 32 hours. Although assessors can produce reliable scores with minimal training, intra-assessor reliability cannot be taken for granted and all assessors should be trained and calibrated before being sanctioned to conduct assessments, particularly for regulatory purposes. The Leicester assessment package has now been shown to be valid, reliable, feasible and easy to use in practice. It can, therefore, be recommended for use in both formative and summative assessment of consultation competence in general practice.

Communication↗

Chiropractic biophysics lateral cervical film analysis reliability.

OBJECTIVE: To determine the degree to which the geometric line drawings used in Chiropractic Biophysics Technique (CBP) on lateral cervical radiographs are reliable. DESIGN: A blind, delayed repeated measures design was used. Three examiners were presented radiographs in random order. All identifying marks were removed prior to each examiner's individual marking and measurement. Each examiner was blinded as to how the previous examiners marked and measured the radiographs. SETTING: Primary care private chiropractic clinic. PATIENTS PARTICIPANTS: Sixty-five subject films were provided from the patient records of a primary care private chiropractic clinic. The 65 radiographs qualified for inclusion in the study based on two criteria: C1 through C7 had to be clearly visible, and there had to be no identifying artifacts. MAIN OUTCOME MEASURES: Anterior head translation in millimeters, atlas plane to horizontal, Ruth Jackson's cervical stress lines, and five relative rotation angles for C2-C3, C3-C4, C4-C5, C5-C6, C6-C7. Inter- and intrareliability of the three examiners were statistically analyzed. RESULTS: Intraexaminer for a) C1 to horizontal reliability was .98-.99 with confidence intervals of .96-.99, b) absolute rotation angle from C2 to C7 reliability was .82-.95 with confidence intervals of .80-.99, c) anterior head translation [+Sz] reliability was .86-.99, with confidence intervals of .74-.99, d) relative rotation angle reliability ranges were (C2-C3) .99, and (C3-C4) .98-.99, (C4-C5) .88-.99, (C5-C6) .80-.99, and (C6-C7) .94-.98. Interexaminer reliabilities across examiners ranged from a) Winer:.89-.99 and b) Bartko: .72-.96. CONCLUSIONS: The reliabilities for intra- and interexaminer were all greater than .70, indicating that these measurements in CBP technique would be considered accurate enough to provide measurements for future clinical studies. The data indicated that the C6-C7 relative rotation angle was the least reliable measurement. This might be due to the very small angles found at this level.

Analysis of Variance↗

Intra- and interexaminer reliability of the chiropractic biophysics lateral lumbar radiographic mensuration procedure.

OBJECTIVE: To determine the intra- and interexaminer reliability of a specific method of mensuration commonly used to evaluate the positional configuration of the lumbopelvic spine viewed on lateral lumbar radiographs. DESIGN: A blind, repeated-measures design was used. Lateral lumbopelvic radiographs were presented to each of three examiners in random order. Each film was marked and measurements were recorded. The films were cleaned of all markings and randomized again for a second run by each examiner. Each examiner's measurements were unavailable to the other examiners. SETTING: Private, primary-care chiropractic clinic. MAIN OUTCOME MEASURES: Anterior/posterior thoracic translation in millimeters, Ferguson's sacral-plane angle to horizontal, arcuate line angle to horizontal, L1 to L5 absolute rotation angle and four relative rotation angles for L1-L2, L2-L3, L3-L4 and L4-L5. Intra- and interrelibility of the three radiographic examiners were analyzed. RESULTS: Intraexaminer reliability for (a) L1-L5 absolute rotation angle was .98, with confidence intervals included in the range of 0.95-0.99, (b) anterior/posterior thorax translation [+/- Sz] was .97-.99, with confidence intervals included in the range of 0.94-1.00, (c) arcuate angle (AA) .40-.81, with confidence intervals included in the range of 0.07-0.90, (d) Ferguson's angle (FA) was .91-.97, with confidence intervals included in the range of 0.82-0.98, (e) relative rotation angle reliability ranges were L1-L2, .84-.94; L2-L3, .80-.85; L3-L4, .78-.89; L4-L5, .87-.92. Interexaminer reliabilities for the three examiners ranged from .66-.98. CONCLUSION: With the exception of the arcuate angle measurement, the reliabilities for all other measurements were at least .78. Those measurements with reliabilities approaching .80 or better would be considered accurate enough for use in future clinical studies. The arcuate angle measurement may have been least reliable because of the subjective nature of the method of affixing a best-fit line to a radiographic landmark that often takes on the appearance of a mild curvature. Establishing reliability is an important first step toward evaluating these and other similar radiographic measurements that have yet to be examined for their validity.

Analysis of Variance↗

Reliability of nerve conduction studies among active workers.

Nerve conduction studies play an important role in clinical practice and research. Given their widespread use, reliability of tests merits careful attention. We assessed interexaminer and intraexaminer reliability of median and ulnar sensory nerve measures of amplitude, onset latency, and peak latency. In a two-phase cross-sectional study, two examiners tested 158 workers. Reliability was assessed with intraclass correlations (ICC) and kappa statistics. Median nerve measures were more reliable (ICC range, 0.76 to 0.92) than ulnar measures (ICC range, 0.22 to 0.85). Ulnar-onset latencies had the worst reliability. The median-ulnar peak latency difference was a particularly stable measure (ICC range, 0.79 to 0.92). The median-ulnar peak latency difference had high interexaminer reliability (kappa range, 0.71 to 0.79) for normal tests defined by cut points of 0.8 ms and 0.5 ms. Intraexaminer reliability was higher with the 0.8-ms cut point (kappa = 0.90 and kappa = 0.85 for examiners 1 and 2, respectively). Rather than absolute cut points to describe normality, a more rational interpretation of results can be made with ordered categories or continuous measures.

Adult↗

Validity and test-retest reliability of a disability questionnaire for essential tremor.

BACKGROUND: One important outcome in clinical trials is patients' own opinions about whether the medication alleviates their symptoms and improves their ability to function. A valid and reliable method with which to assess this subjective information is important. OBJECTIVE: To determine the validity and test-retest reliability of the Columbia University Disability Questionnaire for Essential Tremor (ET). METHODS: Patients with ET underwent a 2.5-hour evaluation, including a 36-item tremor disability questionnaire, to assess the functional impact of tremor, a 26-item videotaped tremor examination rated by a neurologist, a 15-item performance-based test, and quantitative computerized tremor analysis. We determined the validity and test-retest reliability of the tremor disability questionnaire. Correlations between variables were assessed using Pearson's correlation coefficients and test-retest reliability with the weighted kappa statistic. RESULTS: Ninety-five patients with ET participated. The score on tremor disability questionnaire correlated with the neurologist's clinical ratings (r = 0.57, p <0.001) and the total score on the performance-based test (r = 0.69, p < 0.001). Correlations with quantitative computerized tremor analysis results were less robust, but each remained significant, including mean amplitude of dominant arm tremor while arms were extended (r = 0.56, p <0.001), while drawing a spiral (r = 0.42, p = 0.01), and while pouring (r = 0.34, p = 0.04). The questionnaire was readministered to 32 subjects, and the test-retest reliability was substantial (weighted kappa = 0.67). CONCLUSIONS: This Tremor Disability Questionnaire demonstrated substantial reliability, and it correlated with multiple measures of tremor severity, including a neurologist's clinical ratings, a performance-based test of function, and quantitative computerized tremor analysis results. The questionnaire would be useful in clinical trials in which it could be used as a reliable and valid tool to assess disability in ET.

Activities of Daily Living↗

Spatial information content and reliability of hippocampal CA1 neurons: effects of visual input.

The effects of darkness on quantitative spatial firing characteristics of 235 hippocampal CA1 "complex spike" (CS) cells were studied in young and old Fischer-344 rats during food-motivated performance of a randomized, forced-choice task on an eight-arm radial maze. The room lights were turned on or off on alternate blocks of all eight arms. In the dark, a lower proportion of CS cells had "place fields," and the fields were less specific and less reliable than in the light. A small number of cells had place fields unique to the dark condition. Like CS cells, Theta cells showed a reduction in spatially related firing in the dark. The specificity and reliability of the place fields under both light and dark conditions were similar for both age groups. Increasing the salience of the environment, by increasing the light level and the number of visual cues in the light condition, did not affect the specificity or reliability of the place fields. Even though all rats had substantial prior experience with the environment, and were placed on the maze center under normal illumination before the first dark trial, the correlation between the firing pattern in the light and dark increased after the rat first traversed the maze in the light. Thus, even after considerable experience with the environment over days, experiencing the illuminated environment from different locations on a given day was a significant factor affecting subsequent location and reliability of place fields in darkness. While the task was simple and errors rare, rats that made fewer errors (i.e., re-entries into the previously visited arm) also had more reliable place cells, but no such correlation was found with place cell specificity. Thus, the reliability of spatial firing in the hippocampus may be more important for spatial navigation than the size of the place fields per se. Alternatively, both spatial memory and place field reliability may be modulated by a common variable, such as attention.

Action Potentials↗

Reliability between two observers using a protocol for diagnosing essential tremor.

Protocols with demonstrated reliability have been established for the diagnosis of numerous movement disorders. whereas in the essential tremor (ET) literature, there is no discussion about the reliability of diagnostic protocols. Lack of knowledge of the reliability of diagnostic protocols in ET limits the use of these protocols because reliability is an essential requirement for scientific quality in data management. The objective of this study was to determine the reliability of a protocol for diagnosing ET. The protocol consists of a Tremor Interview, a videotaped Tremor Examination, and a diagnostic algorithm. Eighty-three subjects with ET, identified in a community-based health study in Washington Heights-Inwood, New York, were matched with 83 control subjects from the same community. These subjects and their relatives are being recruited to participate in the Washington Heights-Inwood Genetic Study of ET. Two hundred twenty-six subjects have been evaluated to date (35 ET cases, 40 controls, 151 relatives). All 226 underwent an 84-item Tremor Interview and 26-item videotaped Tremor Examination. Diagnoses (normal, possible ET, probable ET, definite ET) were independently assigned by two blinded neurologists specializing in movement disorders. The kappa statistic, k, was used to determine diagnostic agreement between these two neurologists. The concordance rate between two raters using diagnostic categories definite ET, probable ET. possible ET, and normal was 80%; kw = 0.84 (near perfect to perfect agreement). The concordance rate between two raters using two diagnostic categories (definite ET and normal) was 100%; k = 1.00 (perfect agreement). There was high correlation between the two raters' total tremor scores (r = 0.89, p < 0.00001). This diagnostic protocol is highly reliable. Research in ET would greatly benefit from diagnostic protocols with demonstrated reliability.

Adolescent↗

Intratester and intertester reliability and criterion validity of the parallelogram and universal goniometers for active knee flexion in healthy subjects.

BACKGROUND AND PURPOSE: A new parallelogram goniometer was designed by the Rehabilitation Centre of the Royal Ottawa Health Care Group in 1983. The advantage of using such a goniometer is that the clinician is not required to estimate the joint axis of rotation when taking a measurement. The parallelogram goniometer has obtained a good intratester and intertester reliability when measuring active range of motion of hip abduction on eight individuals with hip pathologies. However, the validity of the parallelogram goniometer has not been examined. The purposes of this study were to examine the intratester and intertester reliability and the criterion validity of the parallelogram and universal goniometers for active knee flexion on healthy individuals. SUBJECTS: Sixty healthy university students (44 females and 16 males; mean age of 20.6 yrs.) participated to this study. METHODS: Measurements with the universal and parallelogram goniometers were taken in two different positions, the smaller and larger angles of active knee flexion. All measurements were taken by two trained testers. A radiograph was taken in both positions to serve as the 'gold standard'. The sequence of the measurements and radiographs were randomly selected. The intra and intertester reliability of both goniometers were established by calculating the intraclass correlation coefficients (ICCs) using the repeated-measures ANOVA. The criterion validity was examined by calculating Pearson product-moment correlation coefficients (tau) between each goniometric and radiologic measurements. A 0.05 level of significance was chosen for each statistical test. RESULTS: Intratester reliability ranged from good to excellent for the small angles (ICC = 0.85 and 0.87) and the large angles (ICC = 0.91 and 0.96) when using the parallelogram goniometer. Intertester reliability was fair for the small angles of flexion (ICC = 0.43 to 0.52) and good to excellent for the large angles of flexion (ICC = 0.82 to 0.88). The parallelogram goniometer was found to have greater validity when measuring the large angles of knee flexion (r = 0.73 and 0.77) compared to the small angles of knee flexion (r = 0.33 and 0.41). Similar results of reliability and validity were obtained with the universal goniometer. CONCLUSION: The results of this study have clinical importance. The use of the parallelogram goniometer was found to be as reliable and valid as the universal goniometer when measuring active knee flexion. However, the parallelogram goniometer offered clinicians the advantages of obtaining precise angular measurements with fewer adjustments, and a faster application technique. Further studies on the parallelogram goniometer are necessary among individuals presenting with altered range of motion at different joints.

Adult↗

Reliability of self-reported breast screening information in a survey of lower income women.

BACKGROUND: Self-reported behavior is widely used to estimate the prevalence of breast cancer screening and to evaluate programs for promoting screening, but detailed studies of reliability have not previously been performed. METHODS: Reliability was assessed by comparing responses to questions about screening behavior from repeat personal interviews of 382 women age 40 and older living in low-income census tracts of two Florida communities. Reliability was assessed using Pearson's correlation (r) and kappa (kappa) coefficients. RESULTS: Estimated reliabilities were kappa = 0.38 for "ever had clinical breast examination," kappa = 0.82 for "ever had mammogram," kappa = 0.65 for "mammogram in past year," r = 0.54 for "date of last mammogram," and r = 0.72 for "number of mammograms." The dates of last mammogram reported at the two interviews agreed within 1 month for 64% of the women, while the dates of last clinical breast examination agreed within 1 month for 50% of the women. Reliability of "ever had mammogram" was significantly related to demographic variables. CONCLUSIONS: Women reliably report ever having mammography, but information about timing and frequency has lower reliability. The results have implications for breast screening research because measurement error affects the precision of estimates and the sample sizes needed to detect program effects.

Adult↗

Reliability of psychophysiological responding as a function of trait anxiety.

This study examined the temporal stability of three psychophysiological responses (frontal electromyographic activity, hand surface temperature, and heart rate) recorded over four sessions (days 1, 2, 8, and 28) on 34 subjects, 17 with high Spielberger Trait Anxiety Inventory scores and 17 with low scores. Each session consisted of a 20-minute adaptation period, a baseline condition, and two stressors (one cognitive, the other physical). Two forms of reliability coefficients were employed, intraclass correlations and Pearson Product Moment; the two types of reliability coefficients arrived at the same conclusions. Results indicated that reliability coefficients for the two anxiety groups did not differ on frontal EMG or heart rate responses; however, hand surface temperature responding was considerably less reliable for high anxious individuals than low anxious individuals. Reliability coefficients on absolute scores were, for the most part, reliable. Treating the responses as relative measures (percent change from baseline or simple change scores from baseline) produced smaller and less reliable coefficients. Magnitudes of the three physiological responses did not significantly differ as a function of high or low trait anxiety. Findings are discussed in terms of their clinical, as well as basic psychophysiological, importance.

Adult↗

Reliability and validity of clinical outcome measurements of osteoarthritis of the hip and knee--a review of the literature.

High reliability and validity of clinical rating schemes is crucial for their use as outcome measurements of treatment of hip and knee osteoarthritis. In this paper, we review the empirical evidence on the reliability and validity of commonly used clinical scores. Clinical scores and related reliability and validity studies were identified by systematic literature search. Scores were classified according to the type and joint. Reliability and validity studies were characterized according to design, population, number and qualification of observers, number of measurements, time interval between repeat measurements and results. Reliability and validity studies were reported for only 6 and 15 of the 45 identified clinical scores, respectively. Although comparisons are difficult due to differences in study design, relatively high reliability was reported for most measurements of pain, stiffness, and physical function, while results are less conclusive for clinical signs. Most validity studies focused on the correlation between various scores. Correlation was generally found to be high for overall numerical ratings, but scores often differed with respect to the interpretation of these ratings. Validity has been more comprehensively studied for Lequesne's scores, WOMAC, and ILAS, and these scores have shown satisfactory responsiveness to different treatment effects. Overall, knowledge on reliability and validity of clinical scores of hip and knee osteoarthritis is limited, underlining the need for further properly designed and conducted studies.

Hip↗

Reliability of the Health Utilities Index--Mark III used in the 1991 cycle 6 Canadian General Social Survey Health Questionnaire.

This study presents information on the test-retest reliability of the Health Utility Index--Mark III (HUI) system used in cycle 6 of the Canadian General Social Survey (GSS). The HUI system used in this reliability study consists of an eight-attribute health status classification system (HSCS) and a function for generating a summary score of health-related quality of life. To estimate test-retest reliability, a stratified random sample of individuals (n = 506) completing GSS telephone interviews during August and September, 1991 were interviewed again 1 month later. Weighting adjustments based on the probability of selection were invoked during the analyses to provide unbiased estimates of test-retest reliability for all GSS respondents in the August-September period. The results indicate that the individual questions, attributes and provisional index scores generally provided reliable information on health status in the GSS. The exceptions to this were limitations in speech and dexterity which were reported very infrequently. Kappa estimates of test-retest reliability for individual questions varied from 0.184 to 0.766. For the eight attributes, kappa estimates varied from 0.137 to 0.728. Using the provisional index scores to quantify health overall, a test-retest reliability of 0.767 was obtained (intra-class correlation coefficient).

Activities of Daily Living↗

Factors affecting the reliability of ratings of students' clinical skills in a medicine clerkship.

OBJECTIVE: To determine the overall reliability and factors that might affect the reliability of ratings of students' clinical skills in a medicine clerkship. DESIGN: A nine-item instrument was used to evaluate students' clinical skills. Raters were also asked to provide a grade of each student's overall clinical performance. Generalizability studies were performed to estimate the reliability of the ratings. The effects of rater experience and clerkship setting were investigated by regression analysis. SETTING: Teaching hospitals and community-based sites in three Northwestern states. PARTICIPANTS: All students (328) who had completed the 12-week clerkship in internal medicine at one medical school during the academic years 1987-1989. Raters included attending physicians, chief residents, and other residents. RESULTS: Seven observations were needed to provide a reliable rating of the overall clinical grade. More observations were needed to obtain reliable ratings for individual items, ranging from seven observations needed for the rating of data gathering skills to 27 observations needed for the rating of interpersonal relationships with patients. Rater experience and clerkship setting (i.e., teaching hospitals vs. community-based clinics) were found, in general, not to affect significantly the ratings received by students. CONCLUSIONS: Reliable ratings of students' overall clinical skills, including overall clinical grades, can be achieved by collecting a minimum of seven observations. More observations are needed to measure reliably the interpersonal aspects of clinical performance. These findings support the use of performance ratings to evaluate clinical skills and knowledge of students in clerkship settings.

Clinical Clerkship↗

Reliability of the timeline follow-back sexual behavior interview.

The reliability of self-reported sexual behavior is a question of utmost importance to human immunodeficiency virus (HIV) prevention research. The Timeline Follow-Back (TLFB) interview, which was developed to assess alcohol consumption on the event level, incorporates recall-enhancing techniques that result in reliable information. In this study, the TLFB interview was adapted to assess HIV-related sexual behaviors and their antecedents, and its reliability was assessed. The interview was administered to 110 participants (46% women, M age = 19.7; range = 18-41), and 58 participants who reported sexual behavior during the previous three months returned one week later for a second interview. Test-retest intraclass correlations (rho) from the TLFB protocol showed that all sexual behaviors were reported reliably (rho range = .86 to .97, median = .96). Bootstrapping, a nonparametric statistical technique, was used for significance testing in the reliability analyses. Reliability was equivalent across each of the three months assessed with the TLFB and was equivalent to conventional assessment methods (i.e. single-item questions). These findings show that the TLFB sexual behavior interview provides reliable reports of sexual behavior over three months and yields event-level data that are extremely valuable for sexual behavior and HIV-prevention research.

Adolescent↗

Inter- and intrajudge reliability of a clinical examination of swallowing in adults.

This study investigates inter- and intrajudge reliability of a clinical examination of swallowing in adults. Several investigations have sought correlations between clinical indicators of dysphagia and the actual presence of dysphagia as determined by videofluoroscopy. Whereas some investigations have reported interjudge reliability for the videofluoroscopic measures employed, none have reported reliability for clinical measures. Without established reliability for rating clinical measures, conclusions drawn regarding the utility of a measure for detecting aspiration can be called into question. Results of the present study indicate that fewer than 50% of the measures clinicians typically employ are rated with sufficient inter- and intrajudge reliability. Measures of vocal quality and oral motor function were rated more reliably than were history measures or measures taken during trial swallows. There is a need to define more clearly the measures employed in clinical examinations and to be consistent in reporting reliability for clinical measures of swallowing function in future research.

Adult↗

Reliability of goniometric measurements and visual estimates of ankle joint active range of motion obtained in a clinical setting.

We examined intratester and intertester reliability for goniometric measurements of ankle dorsiflexion (ADF) and ankle plantar flexion (APF) active range of motion (AROM). Parallel-forms intratester reliability for ankle AROM measurements obtained by the universal goniometer (UG) and by visual estimation (VE) and intertester reliability for VE of ADF and APF were examined. Repeated measurements were obtained on 38 patients with orthopedic problems by 10 physical therapists in a clinical setting. For intratester reliability of measurements obtained with UG, intraclass correlation coefficients (ICC) for all physical therapists were 0.64 to 0.92 (median, 0.825) for ADF and 0.47 to 0.96 (median, 0.865) for APF. Intertester reliability was quantified with use of ICC. ICCs for measurements obtained by UG were 0.28 for ADF and 0.25 for APF; ICC of VE for ADF was 0.34 and was 0.48 for APF. ICC for parallel-forms intratester reliability obtained with UG and VE ranged from 0 to 0.94 (median, 0.58) for ADF and 0 to 0.86 (median, 0.625) for APF. Thus, a physical therapist should use a goniometer when making repeated measurements of ankle joint AROM. Considerable inconsistency exists when two or more physical therapists make repeated goniometric and visual measurements of ankle motion on the same subject. Physical therapists may erroneously conclude that a patient's AROM has changed because of treatment when the change could be attributed to a lack of intertester reliability.

Adolescent↗