PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Reliability of spinal palpation for diagnosis of back and neck pain: a systematic review of the literature.

STUDY DESIGN: A systematic review. OBJECTIVES: To determine the quality of the research and assess the interexaminer and intraexaminer reliability of spinal palpatory diagnostic procedures. SUMMARY OF BACKGROUND DATA: Conflicting data have been reported over the past 35 years regarding the reliability of spinal palpatory tests. METHODS: The authors used 13 electronic databases and manually searched the literature from January 1, 1966 to October 1, 2001. Forty-nine (6%) of 797 primary research articles met the inclusion criteria. Two blinded, independent reviewers scored each article. Consensus or a content expert reconciled discrepancies. RESULTS: The quality scores ranged from 25 to 79/100. Subject description, study design, and presentation of results were the weakest areas. The 12 highest quality articles found pain provocation, motion, and landmark location tests to have acceptable reliability (K = 0.40 or greater), but they were not always reproducible by other examiners under similar conditions. In those that used kappa statistics, a higher percentage of the pain provocation studies (64%) demonstrated acceptable reliability, followed by motion studies (58%), landmark (33%), and soft tissue studies (0%). Regional range of motion is more reliable than segmental range of motion, and intraexaminer reliability is better than interexaminer reliability. Overall, examiners' discipline, experience level, consensus on procedure used, training just before the study, or use of symptomatic subjects do not improve reliability. CONCLUSION: The quality of the research on interreliability and intrareliability of spinal palpatory diagnostic procedures needs to be improved. Pain provocation tests are most reliable. Soft tissue paraspinal palpatory diagnostic tests are not reliable.

Back Pain↗

Reliability analysis for digital adolescent idiopathic scoliosis measurements.

OBJECTIVE: Analysis of adolescent idiopathic scoliosis (AIS) requires a thorough clinical and radiographic evaluation to completely assess the three-dimensional deformity. Recently, these radiographic parameters have been analyzed for reliability and reproducibility following manual measurements; however, most of these parameters have not been analyzed with regard to digital measurements. The purpose of this study is to determine the intra- and interobserver reliability of common scoliosis radiographic parameters using a digital software measurement program. METHODS: Thirty sets of preoperative (posteroanterior [PA], lateral, and side-bending [SB]) and postoperative (PA and lateral) radiographs were analyzed by three independent observers on two separate occasions using a software measurement program (PhDx, Albuquerque, NM). Coronal measures included main thoracic (MT) and thoracolumbar-lumbar (TL/L) Cobb, SB MT Cobb, MT and TL/L apical vertical translation (AVT), C7 to center sacral vertical line (CSVL), T1 tilt, LIV tilt, disk below lowest instrumented vertebra (LIV), coronal balance, and Risser, whereas sagittal measures included T2-T5, T5-T12, T2-T12, T10-L2, T12-S1, and sagittal balance. Analysis of variance for repeated measures or Cohen three-way kappa correlation coefficient analysis was performed as appropriate to calculate the intra- and interobserver reliability for each parameter. RESULTS: The majority of the radiographic parameters assessed demonstrated good or excellent intra- and interobserver reliability. The relationship of the LIV to the CSVL (intraobserver kappaa = 0.48-0.78, fair to excellent; interobserver kappaa = 0.34-0.41, fair to poor), interobserver measurement of AVT (rho = 0.49-0.73, low to good), Risser grade (intraobserver rho = 0.41-0.97, low to excellent; interobserver rho = 0.60-0.70, fair to good), intraobserver measurement of the angulation of the disk inferior to the LIV (rho = 0.53-0.88, fair to good), apical Nash-Moe vertebral rotation (intraobserver rho = 0.50-0.85, fair to good; interobserver rho = 0.53-0.59, fair), and especially regional thoracic kyphosis from T2 to T5 (intraobserver rho = 0.22-0.65, poor to fair; interobserver rho = 0.33-0.47, low) demonstrated lesser reliability. In general, preoperative measures demonstrated greater reliability than postoperative measures, and coronal angular measures were more reliable than sagittal measures. CONCLUSIONS: Most common radiographic parameters for AIS assessment demonstrated good or excellent reliability for digital measurement and can be recommended for routine clinical and academic use. Preoperative assessments and coronal measures may be more reliable than postoperative and sagittal measurements. The reliability of digital measurements will be increasingly important as digital radiographic viewing becomes commonplace.

Adolescent↗

The reliability of patients' judgements of care in general practice: how many questions and patients are needed?

OBJECTIVES: To estimate the number of questions and patients that are needed to achieve reliable measurements of patients' judgements of care in general practice. DESIGN: Sensitivity study, using generalisibility theory and real data from surveys of patients. SUBJECTS: 739 patients with chronic illness from 23 general practitioners in The Netherlands. MAIN MEASURES: The reliability coefficients of scores per patient and scores per general practitioner for patients' judgements of nine dimensions of care in general practice. RESULTS: For most dimensions the reliability per patient was 0.80 or higher if three questions were used, but for the evaluation of the "organisation of appointments" and "premises" five questions had to be used. To reach a reliability coefficient of 0.80 per general practitioner three questions and 90 patients, or five questions and 60 patients, were needed for most dimensions. Even more patients or questions were needed for the dimensions "availability for emergencies", premises, and "continuity". A reliability of 0.70 per general practitioner could be achieved if three questions and 60 patients were used, except for availability for emergencies and premises, for which more patients or questions were required. CONCLUSIONS: Surveys of patients can only provide reliable information if the samples of questions and patients are large enough. It is important to distinguish between the reliability of scores per patient and the reliability per care provider, as well as between different dimensions of care. The reliability per patient is good for most dimensions if three questions are used, but a good reliability per care provider requires more questions or patients.

Aged↗

[Test-retest and inter-rater reliability of the French version of the Ontario Society of Occupational Therapy (OSOT) Perceptual Evaluation].

This article presents the results of a study conducted to verify the test-retest and inter-rater reliability of the French version of the Ontario Society of Occupational Therapy (OSOT) Perceptual Evaluation. Designed to evaluate the perceptual deficits in patients with brain injuries, this tool uses a 5-points scale (0-4) to measure 18 different tasks. The scores obtained for each task are added to establish a total score. In the early 90s, the instruction manual of the OSOT Perceptual Evaluation was translated in French by a group of occupational therapists from l'Institut universitaire en gériatrie de Sherbrooke. To ensure the reliability of this version, a study was conducted to determine the test-retest reliability and the simultaneous inter-rater reliability. Thirty-two francophone subjects with brain injuries were each evaluated twice by the same therapist to determine the test-retest reliability of this tool. At one of the two encounters, a second therapist completed the score sheet to verify the simultaneous inter-rater reliability. The results show that, despite a few weak kappas' scores for certain tasks, the test-retest reliability and the inter-rater reliability of the total score were excellent (test-retest reliability: intra-class correlation coefficient = 0.93, with a confidence interval of 0.87 to 0.97; and inter-observer reliability: = 0.98, with a confidence interval of 0.97 to 0.99). The findings of this study show that the French version of the OSOT Perceptual Evaluation can therefore be used confidently by francophone occupational therapists.

Brain Injuries↗

Intra- and inter-observer reliability of the distal metatarsal articular angle in adult hallux valgus.

There is some uncertainty as to whether the distal metatarsal articular angle (DMAA) is a real entity or just radiographic artifact and whether it can be reliably measured. If it is intrinsic to the bone, it should not change with bone position. If it is clinically useful, it should be reproducible. Pre-operative and post-operative radiographs of 32 patients undergoing a proximal bony procedure of the first ray were evaluated independently by three foot and ankle specialists in order to determine the intra and inter-observer reliability of the distal metatarsal articular angle (DMAA). In addition, the hallux valgus angle (HVA), intermetatarsal angle (IMA) and joint congruency/subluxation were determined. We used ANOVA (Scheffe's F-test) to determine reliability of the angular measurements; a p value of less than 0.05 indicates poor reliability and a p value of greater than 0.05 indicates reliability. Intra-observer reliability was good for all angular measurements (HVA, IMA, DMAA pre-op, and DMAA post-op) with p values ranging from 0.33 to 0.95. Inter-observer reliability of the HVA and IMA was good (p=0.63 and p=0.32). Inter-observer reliability of the pre-op DMAA approached statistically poor reliability (p=0.09) and the post-op DMAA reliability was poor (p=0.002). The DMAA reduced after the proximal procedure as measured by all observers, and averaged a reduction of 3.9 degrees. Weighted kappa analysis also revealed that there was poor agreement in the determination of congruency and subluxation (Kappa statistic ranged from 0.07 to 0.19). This study suggests that there may be limited value in the DMAA as a clinical measure as it varies with examiner and with the hallux valgus angle.

Adult↗

Reliability of videotaped observational gait analysis in patients with orthopedic impairments.

BACKGROUND: In clinical practice, visual gait observation is often used to determine gait disorders and to evaluate treatment. Several reliability studies on observational gait analysis have been described in the literature and generally showed moderate reliability. However, patients with orthopedic disorders have received little attention. The objective of this study is to determine the reliability levels of visual observation of gait in patients with orthopedic disorders. METHODS: The gait of thirty patients referred to a physical therapist for gait treatment was videotaped. Ten raters, 4 experienced, 4 inexperienced and 2 experts, individually evaluated these videotaped gait patterns of the patients twice, by using a structured gait analysis form. Reliability levels were established by calculating the Intraclass Correlation Coefficient (ICC), using a two-way random design and based on absolute agreement. RESULTS: The inter-rater reliability among experienced raters (ICC = 0.42; 95%CI: 0.38-0.46) was comparable to that of the inexperienced raters (ICC = 0.40; 95%CI: 0.36-0.44). The expert raters reached a higher inter-rater reliability level (ICC = 0.54; 95%CI: 0.48-0.60). The average intra-rater reliability of the experienced raters was 0.63 (ICCs ranging from 0.57 to 0.70). The inexperienced raters reached an average intra-rater reliability of 0.57 (ICCs ranging from 0.52 to 0.62). The two expert raters attained ICC values of 0.70 and 0.74 respectively. CONCLUSION: Structured visual gait observation by use of a gait analysis form as described in this study was found to be moderately reliable. Clinical experience appears to increase the reliability of visual gait analysis.

Adolescent↗

Reliability of gait speed measured by a timed walking test in patients one year after stroke.

OBJECTIVE: To assess the reliability of gait speed in late-stage stroke patients. DESIGN: Test-retest reliability of three timed walks to 10 metres repeated during two assessments one week apart. SETTING: The patient's home. SUBJECTS: Twenty-two stroke patients with mobility problems more than one year after stroke. MAIN OUTCOME MEASURE: Gait speed measured in seconds taken to walk 10 metres. STATISTICAL ANALYSIS: Intraclass correlations (ICCs) with 95% confidence interval (CI) and the Bland and Altman method for assessing agreement by calculating the mean difference between measurements (d); the 95% CI for d; the standard deviation of the difference (SDdiff); a reliability coefficient and the 95% limits of agreement. RESULTS: There was a trend for decreased times taken to walk 10 metres both within each assessment and between assessments. ICCs for within-assessment reliability were 0.95-0.99. The d (SDdiff) for the second and third walks for assessment 1 was -1.00 (2.63) seconds and for assessment 2 was -0.70 (1.58) seconds. The reliability coefficient was 5.26 for assessment 1 and 3.17 for assessment 2. ICCs for between-assessment reliability were 0.87-0.88. The d (SDdiff) for the comparison of the third walks at assessment 1 and assessment 2 was -0.90 (5.01) seconds. The reliability coefficient was 10.02 and the 95% limits of agreement were -10.92 to +9.12 seconds. CONCLUSION: Within-assessment gait speed measured at home is highly reliable. The between-assessment reliability of gait speed measurement is less reliable but comparable with other studies.

Adult↗

Reliability of individual diagnostic criterion items for psychoactive substance dependence and the impact on diagnosis.

OBJECTIVE: Reliability of diagnostic criterion items for psychoactive substance dependence and the impact of each on the reliability of the diagnosis were analyzed. METHOD: As part of a reliability study for a new interview developed for the multisite Collaborative Study on the Genetics of Alcoholism (COGA), data were collected from both within-center and across centers. The impact of each diagnostic item on the reliability of the substance dependence diagnosis was studied by forcing each item to be reliable one at a time and recomputing the kappa statistic for the diagnosis. RESULTS: Findings indicated that the majority of individual diagnostic criterion items were reliable; 87% and 81% were in the fair or better range of reliability for the within- and cross-center studies, respectively. Individual kappa estimates were statistically similar for the two studies. Reliability findings for two classes of substance, alcohol and cocaine, were good, while those for stimulants were less satisfactory. CONCLUSIONS: Forcing items one at a time to be reliable did not affect reliability of the overall substance dependence diagnosis, because more than one criterion item changed from Time 1 to Time 2. Because no single item was influential, weighting criteria equally, as is done in the DSM and ICD classification systems, appears to be a reasonable approach.

Adolescent↗

Statistical methods for assessing measurement error (reliability) in variables relevant to sports medicine.

Minimal measurement error (reliability) during the collection of interval- and ratio-type data is critically important to sports medicine research. The main components of measurement error are systematic bias (e.g. general learning or fatigue effects on the tests) and random error due to biological or mechanical variation. Both error components should be meaningfully quantified for the sports physician to relate the described error to judgements regarding 'analytical goals' (the requirements of the measurement tool for effective practical use) rather than the statistical significance of any reliability indicators. Methods based on correlation coefficients and regression provide an indication of 'relative reliability'. Since these methods are highly influenced by the range of measured values, researchers should be cautious in: (i) concluding acceptable relative reliability even if a correlation is above 0.9; (ii) extrapolating the results of a test-retest correlation to a new sample of individuals involved in an experiment; and (iii) comparing test-retest correlations between different reliability studies. Methods used to describe 'absolute reliability' include the standard error of measurements (SEM), coefficient of variation (CV) and limits of agreement (LOA). These statistics are more appropriate for comparing reliability between different measurement tools in different studies. They can be used in multiple retest studies from ANOVA procedures, help predict the magnitude of a 'real' change in individual athletes and be employed to estimate statistical power for a repeated-measures experiment. These methods vary considerably in the way they are calculated and their use also assumes the presence (CV) or absence (SEM) of heteroscedasticity. Most methods of calculating SEM and CV represent approximately 68% of the error that is actually present in the repeated measurements for the 'average' individual in the sample. LOA represent the test-retest differences for 95% of a population. The associated Bland-Altman plot shows the measurement error schematically and helps to identify the presence of heteroscedasticity. If there is evidence of heteroscedasticity or non-normality, one should logarithmically transform the data and quote the bias and random error as ratios. This allows simple comparisons of reliability across different measurement tools. It is recommended that sports clinicians and researchers should cite and interpret a number of statistical methods for assessing reliability. We encourage the inclusion of the LOA method, especially the exploration of heteroscedasticity that is inherent in this analysis. We also stress the importance of relating the results of any reliability statistic to 'analytical goals' in sports medicine.

Bias↗

Reliability analysis in therapeutic research: practice and procedures.

Twenty studies examining the reliability of assessment devices and outcome measures in therapeutic research were reviewed and analyzed. The 20 investigations contained 215 quantitative reliability values published in either the American Journal of Occupational Therapy or Physical Therapy during the past 5 years. The reliability studies were classified as interrater, intrarater, test-retest, or internal consistency. Examination of interrater reliability accounted for 41% of all reported reliability values. Studies published in Physical Therapy were more likely to be concerned with test-retest reliability, whereas studies published in the American Journal of Occupational Therapy more often focused on interrater reliability. Examination of the data revealed that the intraclass correlation coefficient (ICC) was the most frequently reported estimate of reliability, accounting for 57% of all reported reliability coefficients. Further review of the results indicated that Pearson product-moment correlations and percentage of agreement indexes accounted for 22% of all reliability values reported in the studies examined. The Pearson product-moment correlation measures association or covariation among variables, but not agreement, and percentage agreement indexes do not correct for chance agreement. The argument is made that product-moment correlations and percentage agreement indexes are inadequate measures of interrater, intrarater or test-retest agreement. They should be used and interpreted with caution.

Humans↗

Test-retest reliability of balance tests in children with cerebral palsy.

To investigate intrasession and intersession reliability of balance tests in children with or without disabilities, 50 children without disabilities (ND) and 36 children with cerebral palsy (CP) aged from 5 to 12 years were tested. Intrasession reliability of postural stability of the Smart Balance Master System and one-leg standing test were assessed in both groups and intersession reliability of the Smart Balance Master System and balance subtest of the Bruininks-Oseretsky Test of Motor Proficiency (BOTMP) were assessed in ND children. Intersession reliability of the postural stability test in ND children, obtained using the Smart Balance Master System, was of moderate to good reliability in centre target (CT), sway vision (SV), eyes open and sway surface (EOSS), and sway vision and sway surface (SVSS; ICC 0.72 to 0.84). In children with CP, intrasession reliability was high in CT (ICC 1). One-leg standing tests in both groups also had moderate to good intersession reliability (ICC 0.56 to 0.99). Agreement of failure score of lateral rhythmic shifting (LRS) at 1 second and 2 seconds pace was 85% and 93% respectively in ND children. Within the balance subtest of BOTMP, only two items had 100% agreement. Results suggest that postural stability tests in four conditions (CT, SV, EOSS, and SVSS), LRS, one-leg standing, and walking on a line are reliable and can be used to monitor balance control in ND children. Postural stability in CT condition and one-leg standing test are reliable in children with CP. Further study is needed to establish more reliable balance tests for children with CP.

Cerebral Palsy↗

Interobserver reliability of osteopathic palpatory diagnostic tests of the lumbar spine: improvements from consensus training.

CONTEXT: Establishing reliable palpatory tests continues to be a critical, yet elusive, step in osteopathic medical research and evidence-based clinical practice. OBJECTIVE: The authors investigated the interobserver reliability of common osteopathic palpatory tests used to evaluate the lumbar spine. DESIGN AND METHODS: Subjects (N=119) were recruited from the faculty, staff, and students of Kirksville (MO) College of Osteopathic Medicine (KCOM) of A.T. Still University of Health Sciences. Three osteopathic medical examiners residency trained in neuromusculoskeletal medicine initially evaluated lumbar segments on subjects from one subgroup (n=42) in a blinded assessment. The examiners performed palpatory tests of tenderness and tissue texture changes, as well as--in three planes--vertebral positional asymmetry and motion asymmetry. Kappa statistics (kappa) were used to evaluate interobserver reliability. Following a period of consensus training, subjects from another subgroup (n=77) were evaluated in a blinded assessment for those palpatory tests that seemed most likely to produce reliable findings. The interobserver reliability was then re-evaluated. RESULTS: During the initial evaluation of interobserver reliability, kappa ranged from -0.02 to 0.34, within the poor-to-fair reliability range. Following consensus training, reliability improved, rising into the moderate range for tissue texture changes (kappa=0.45) and into the substantial range for tenderness assessments (kappa=0.68). Reliability for positional asymmetry in the transverse plane (kappa=0.34) and rotational motion asymmetry (kappa=0.20) were improved but remained in the fair range. CONCLUSION: The authors concluded that consensus training improved the interobserver reliability of common osteopathic palpatory tests of the lumbar spine. In two of the four tests that were studied--tissue texture and tenderness--acceptable kappa values for clinical tests were achieved after consensus training.

Adult↗

An examination of reliability in developmental research.

The purpose of this investigation was to examine quantitative methods used to determine reliability in developmental research. Procedures used to compute reliability estimates in 30 studies published in three developmental journals were examined. Four types of reliability studies were identified and analyzed. These included interrater reliability, stability (test-retest and intrarater reliability), equivalence reliability, and internal consistency. Interrater reliability investigations were the most frequently reported in the developmental literature reviewed (45%). The Pearson product moment correlation (r) was the most commonly reported reliability statistic. The findings reveal that researchers in developmental pediatrics frequently analyze reliability data using the Pearson product moment correlation and interpret the results as indicating consensus (agreement) among raters or across instruments. The Pearson product moment correlation (r) provides information on covariation among variables but does not indicate agreement. Thus, the findings suggest that developmental researchers may be misinterpreting the statistical results of reliability investigations. The argument is made that the intraclass correlation coefficient (ICC) is a more appropriate method of analysis when the purpose of the research is to examine consensus.

Adolescent↗

[Computer technology for genogenographic study of the gene pool. V. Evaluation of the reliability of maps].

An important aspect of gene geography, the estimation of the reliability of interpolation maps, is considered. The introduced quantitative parameter, map reliability, was estimated as the probability of predicted character values in interpolated map regions. The obtained estimates characterize the statistic significance of the mapped values of the character. The estimation algorithm is based on concepts and mathematical methods of the reliability theory. The proposed approach involved the estimation of the reliability at each point of the mapped area and resulted in a new map (a reliability map) expressing the reliability of gene geographic mapping in probability terms. Approaches to estimate the reliability of mapping, as dependent on various parameters of the initial data, were proposed, a general computer-based technology was elaborated, and a standardized reliability scale was proposed. Reliability maps are considered necessary for the correct interpretation of gene geographic maps. The estimation of reliability as dependent on the number and distribution of initially tested populations was illustrated by the example of frequency maps of individual genes (HP*1, HLA*A1) and synthetic maps (100 alleles of 34 polymorphic loci) of Eastern Europe.

Algorithms↗

The NEO-FFI is a reliable measure of premorbid personality in patients with probable Alzheimer's disease.

AIM: To assess the inter-informant reliability, intra-informant reliability and internal consistency of the NEO-FFI as a measure of premorbid personality in patients with Alzheimer's Disease (AD). SUBJECTS: One hundred and five persons with NINCDS-ADRDA probable AD for the assessment of inter-informant reliability and internal consistency, and 30 for the assessment of intra-informant reliability. METHODS: Premorbid personality was rated retrospectively by close relatives remembering the patient as he/she had been when aged in his/her forties. One hundred and five AD patients were rated by two separate informants. Thirty AD patients were rated by the same informant on separate occasions one year apart. RESULTS: Inter-informant reliability for the five domain scores of the NEO-FFI was shown to range from fair to good when measured using the single measure Intraclass Correlation Co-efficient (ICC) (0.52-0.64), and to range from good to excellent when measured using the average ICC (0.68-0.78). Intra-informant reliability for four out of the five domains was shown to be excellent when measured using the single ICC (0.81-0.92), and good for the remaining domain (0.72). Intra-informant reliability was found to be excellent for all five domains when measured using the average ICC (0.84-0.96). Internal consistency of the five domains was good. CONCLUSIONS: The NEO-FFI can be used reliably to measure premorbid personality in patients with probable AD. It may be useful to maximise reliability by using a mean domain score based on questionnaires completed by two or more informants who knew the patient well earlier in life.

Aged↗

Reliability of pelvic floor muscle strength assessment using different test positions and tools.

AIMS: The aims of this study were to determine the intra-therapist reliability for digital muscle testing and vaginal manometry on maximum voluntary contraction strength and endurance. In addition, we assessed how reliability varied with different tools and different testing positions. METHODS: Subjects included 20 female physiotherapists. The modified Oxford scale was used for the digital muscle testing, and the Peritron perineometer was used for the vaginal resting pressure and vaginal squeeze pressure assessments. Strength and endurance testing were performed. The highest of the maximum voluntary contraction scores was used in strength analysis, and a fatigue index value was calculated from the endurance repetitions. Bent-knee lying, supine, sitting, and standing positions were used. The time interval for between-session reliability was 2-6 weeks. RESULTS: Kappa values for the between-session reliability of digital muscle testing were 0.69, 0.69, 0.86, and 0.79 for the four test positions, respectively. Intra-class correlation coefficient (ICC) values for squeeze pressure readings for the four positions were 0.95, 0.91, 0.96, and 0.92 for maximum voluntary contraction, and 0.05, 0.42, 0.13, and 0.35 for endurance testing. ICC values for resting pressure were 0.74, 0.77, 0.47, and 0.29. CONCLUSIONS: Reliability of digital muscle testing was very good in sitting and good in the other three positions. vaginal resting pressure demonstrated very good reliability in all four positions for maximum voluntary contraction, but was unreliable for endurance testing. Vaginal resting pressure was not reliable in upright positions. Both measurement tools are reliable in certain positions, with manometry demonstrating higher reliability coefficients.

Adult↗

Reliability of physical functioning measures in ambulatory subjects with MS.

BACKGROUND AND PURPOSE: One of the primary reasons for measuring outcomes during rehabilitation is to determine the effect of physiotherapy. Repeated measurement situations are susceptible to several sources of error, including inconsistencies caused by the subject, the procedure, the instrument and the examiner. Therefore, the reliability of the measures needs to be examined. METHOD: The present study used a repeated-measures design. Two studies were undertaken to examine the test-retest and inter-rater reliability for physical functioning measures. The interval between the measurements was one week. The sample consisted of 19 ambulatory subjects with multiple sclerosis (MS) in the test-retest and nine subjects in the inter-rater reliability study. The measures were selected to assess different domains of the World Health Organization International Classification of Functioning, Disability and Health (WHO, 2001). Several parameters of the Box and Block Test (BBT), the Berg Balance Scale (BBS), the Kela Coordination test, the postural stability test, the timed 10-metre gait test, the six-minute walk test, the shoulder tug test, grip strength, maximal isometric force of the knee extensors, muscle endurance tests, the modified Ashworth Scale and passive straight leg raise test were examined in terms of reliability. RESULTS: The intra-class coefficient (ICC) values for test-retest reliability were >0.80 in 17 of 23 parameters, and correspondingly so in 20 out of 26 parameters for inter-rater reliability. Poor reliability (defined as ICC < or = 0. 60) was obtained only for the patient classification index (PCI) of the six-minute walk test in the test-retest reliability study. In general, the coefficient of variation was good. A moderate amount of variability was discovered for the Kela Coordination test, and for postural stability and muscle endurance tests. The data obtained from the modified Ashworth Scale and the shoulder tug test were highly skewed and the percentage of agreement ranged between 63.9% and 93.4%. CONCLUSIONS: The study revealed acceptable test-retest and inter-rater reliability of these measures in ambulatory subjects with MS, with the exception of the Modified Ashworth Scale and the shoulder tug test.

Adult↗

The reliability of three perceptual evaluation scales for dysphonia.

Perceptual rating scales are widely used in voice quality assessment, yet apart from the GRBAS scale, their reliability has been poorly demonstrated. There are no studies that have compared the optimal reliability of experienced judges using different auditory rating scales in a controlled experimental environment. This study aimed to assess the reliability of three common scales (The Buffalo Voice Profile, The Vocal Profile Analysis Scheme (VPA) and GRBAS. Sixty-five varyingly dysphonic and five normal voices were recorded onto CD in random order. Thirty voices were recorded twice. Seven experienced and trained speech and language therapists rated all voices on the three scales. Only the overall grade was found to be reliable for the Buffalo Voice Profile. The reliability of the VPA scheme was found to be poor to moderate. The VPA may have a use as a multi-dimensional and in-depth evaluation of voice types, but its greater scope is at the expense of reliability. The GRBAS was reliable across all parameters except Strain. Our detailed reliability analysis comparing performance of three commonly used rating scales provides further evidence to support the GRBAS as a simple reliable measure for clinical use.

Adolescent↗