PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “reliability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Reliability of the Tone Assessment Scale and the modified Ashworth scale as clinical tools for assessing poststroke spasticity.

OBJECTIVES: To establish reliability of the Tone Assessment Scale and modified Ashworth scale in acute stroke patients. SETTING: A North Liverpool university hospital. PATIENTS: Eighteen men and 14 women admitted with acute stroke and still in hospital at the study start date (median age, 74 yrs; median Barthel score, 8). MAIN OUTCOME MEASURES: The modified Ashworth scale and the Tone Assessment Scale. STUDY DESIGN: The 32 patients were examined with both scales on the same occasion by two raters (interrater comparison) and on two occasions by one rater (intrarater comparison). RESULTS: The reliability of the modified Ashworth scale was very good (kappa = .84 for interrater and .83 for intrarater comparisons). The reliability of the Tone Assessment Scale was not as strong as the modified Ashworth scale, with marked variability in the assessment of posture (kappa = .22 to .50 for interrater and .29 to .55 for intrarater comparisons) and associated reaction (kappa/kappaW = -.05 to .79 for interrater and .19 to .83 for intrarater comparisons). However, those aspects of the Tone Assessment Scale that addressed response to passive movement and that are scored similarly to the modified Ashworth scale showed good to very good interrater reliability (kappaW = .79 to .92) and good to very good intrarater reliability (kappaW = .72 to .86), except for the question related to movement at the ankle where agreement was only moderate (kappaW = .59). CONCLUSIONS: The modified Ashworth scale is reliable. The section of the Tone Assessment Scale relating to response to passive movement is reliable at various joints, except the ankle. It may assist in studies on the prevalence of spasticity after stroke and the relationship between tone and function. Further development of a measure of spasticity at the ankle is required. The Tone Assessment Scale is not reliable for measuring posture and associated reactions.

Aged↗

Assessment of published reliability studies for cervical spine range-of-motion measurement tools.

OBJECTIVE: To assess the reliability of tools to measure cervical spine range of motion in clinical settings and discuss the necessary components for reliability studies. DATA SOURCES: Database searches included Bandolier, Bath Information and Data Services including Index of Scientific and Technical Proceedings, British Nursing Index, Cumulative Index to Nursing and Allied Health Literature, English National Health Care Database, MEDLINE, Occupational Therapy Index, Physiotherapy Index, and Rehabilitation index for English language articles from 1966. In addition, citations were searched. STUDY SELECTION: Studies were selected that assessed the tool for intraobserver or interobserver reliability, evaluated it on movements of flexion/extension, lateral flexion, or rotation, and measured range of motion of the whole cervical spine. DATA EXTRACTION: All papers were read by one nonclinical researcher with a data extraction sheet. A consultant rheumatologist and a physiotherapist were each asked to read a sample of the papers to give a clinical viewpoint. DATA SYNTHESIS: Evidence for the reliability of measurement tools was assessed qualitatively based on the quality of the study designs, appropriateness of analysis, and strength of the reliability based on reported intraclass correlation coefficients (the most appropriate analysis technique for reliability studies of this nature). Measurement tools were found to have not been fully tested for reliability, particularly in terms of adequate sample size and appropriate analysis techniques. There were also wide variations in the research design, including the protocol for movement, the characteristics of observers and study population, whether warm-ups were allowed, whether the movement was active or passive, and time intervals between repeated measurements. CONCLUSION: Although a range-of-motion device has shown promise in reliability and has many advocates, its practicality for clinical use is questionable. Further work must be performed on all measurement tools. Researchers need to produce more rigorous studies and consider the issues discussed here.

Cervical Vertebrae↗

Use of bioelectrical impedance in hydration status assessment: reliability of a new tool in psychophysiology research.

Adequate hydration is crucial in maintaining optimal physical and mental functioning and the need for a fast and reliable hydration status assessment in behavioral medicine research has become increasingly important. The goal of this study was to determine the reliability of bioelectrical impedance assessment (BIA) in assessing total body water (TBW), extracellular water (ECW) and intracellular water (ICW) and to assess whether individuals can be reliably classified as being hypohydrated or hyperhydrated using lower and upper quartiles, respectively. TBW, ECW and ICW were assessed via BIA (Bodystat, Isle of Man, UK) in 52 male and 48 female college students on 2 separate days within 1 week. Results revealed strong test-retest reliability for TBW (r=0.983), ECW (r=0.972) and ICW (r=0.988) (all P's<0.001). Following the initial and follow-up assessments, participants were then classified as being either hypohydrated or hyperhydrated based on the percentage of body weight accounted for by TBW. Test-retest reliability of hydration status within classifications was then assessed by gender. Test-retest reliability was found for TBW, ECW and ICW among hypohydrated (r=0.985, r=0.972 and r=0.99, respectively) and hyperhydrated (r=0.994, r=0.989 and r=0.994, respectively) males (all P's<0.001). Significant test-retest correlations were also found for females classified as being hypohydrated (r=0.97, r=0.956 and r=0.976, respectively) and hyperhydrated (r=0.973, r=0.976 and r=0.976, respectively) (all P's<0.001). These findings suggest that hydration status, as indexed by bioelectrical impedance technique, is reliable across time and is also reliable within individuals who are chronically hyperhydrated or hypohydrated.

Adolescent↗

Two simple methods for improving the reliability of joint center locations.

OBJECTIVE: A clinically oriented technique is proposed for evaluating the reliability of methods for estimating joint center locations from surface markers, as is an optimization method for estimating joint center locations during planar movements. DESIGN: Segment length variability is used as a measure of reliability, and three simple methods for locating joint centers are compared via repeated measures analysis. Rigorous evaluation is achieved by applying adjustment parameters to a data set, other than the one from which parameters were derived. BACKGROUND: Although more sophisticated techniques are available, many clinical and experimental studies use visual observation and palpation to locate joint centers. This study offers a simple means to evaluate the reliability of that method, and it offers two simple post-hoc methods to improve reliability. METHODS: Single-joint movements are used to generate adjustment parameters from three-dimensional (3D) measurements of surface markers; these are applied to multi-joint movement trials. Segment length variability is compared before and after adjustment with each of two post-hoc methods. RESULTS: As shown by lowered segment length standard deviations, the proposed optimization technique improved reliability compared to the observational and the two-dimensional (2D) post-hoc methods. CONCLUSIONS: The segment length technique offers a simple means to evaluate the reliability with which joint centers are located, and the new optimization method improves reliability for planar multi-joint movements. RELEVANCE: These improvements in reliability are easily implemented within settings where sophisticated technical support may be unavailable.

Journal Article↗

Reliability of the Cardiff Test of basic life support and automated external defibrillation version 3.1.

The introduction of the European Resuscitation Guidelines (2000) for cardiopulmonary resuscitation (CPR) and automated external defibrillation (AED) prompted the development of an up-to-date and reliable method of assessing the quality of performance of CPR in combination with the use of an AED. The Cardiff Test of basic life support (BLS) and AED version 3.1 was developed to meet this need and uses standardised checklists to retrospectively evaluate performance from analyses of video recordings and data drawn from a laptop computer attached to a training manikin. This paper reports the inter- and intra-observer reliability of this test. Data used to assess reliability were obtained from an investigation of CPR and AED skill acquisition in a lay responder AED training programme. Six observers were recruited to evaluate performance in 33 data sets, repeating their evaluation after a minimum interval of 3 weeks. More than 70% of the 42 variables considered in this study had a kappa score of 0.70 or above for inter-observer reliability or were drawn from computer data and therefore not subject to evaluator variability. 85% of the 42 variables had kappa scores for intra-observer reliability of 0.70 or above or were drawn from computer data. The standard deviations for inter- and intra-observer measures of time to first shock were 11.6 and 7.7 s, respectively. The inter- and intra-observer reliability for the majority of the variables in the Cardiff Test of BLS and AED version 3.1 is satisfactory. However, reliability is less acceptable with respect to shaking when checking for responsiveness, initial check/clearing of the airway, checks for signs of circulation, time to first shock and performance of interventions in the correct sequence. Further research is required to determine if modifications to the method of assessing these variables can increase reliability.

Automation↗

Reliability of the Melbourne assessment of unilateral upper limb function.

This study examines the reliability of the Melbourne Assessment of Unilateral Upper Limb Function: a quantitative test of quality of movement in children with neurological impairment. The assessment was administered to 20 children aged from 5 to 16 years (mean age 9 years 10 months, SD 2 years 10 months) who had various types and degrees of cerebral palsy (CP). The performances of the 20 children during assessment were videotaped for subsequent scoring by 15 occupational therapists. Scores were analyzed for internal consistency of test items, inter- and intrarater reliability of scorings of the same videotapes, and test-retest reliability using repeat videotaping. Results revealed very high internal consistency of test items (alpha=0.96), moderate to high agreement both within and between raters for all test items (intraclass correlations of at least 0.7) apart from item 16 (hand to mouth and down), and high interrater reliability (0.95) and intrarater reliability (0.97) for total test scores. Test-retest results revealed moderate to high intrarater reliability for item totals (mean of 0.83 and 0.79) for each rater and high reliability for test totals (0.98 and 0.97). These findings indicate that the Melbourne Assessment of Unilateral Upper Limb Function is a reliable tool for measuring the quality of unilateral upper-limb movement in children with CP.

Adolescent↗

Interobserver reliability of the gross motor performance measure: preliminary results.

Although assessment of the quality of movement in children with cerebral palsy (CP) is difficult, the development of the Gross Motor Performance Measure (GMPM) has facilitated this process. In order to determine the interobserver reliability of the GMPM, 36 children with spastic neuromuscular disorders (mean age 7 years, range 4 to 15 years) were evaluated using four of the five dimensions of the GMPM. Percent Agreement, Intraclass Correlations, and Kappas were calculated by both dimension and attribute to determine reliability. In addition, reliability measures were evaluated over time to determine whether reliability improved with continual use of the GMPM. Overall, interobserver reliability was in the 'fair to good' category regardless of the reliability measure used in the analysis. Reliability scores improved over time with a greater number of individual item scores moving from the 'fair to good' category to the 'excellent' category. Results from this study indicate that it is possible to assess reliably the quality of movement in children with CP.

Adolescent↗

Reliability analysis for hazardous waste treatment processes.

The reliability of a treatment process is addressed in terms of achieving a regulatory effluent concentration standard and the design safety factors associated with the treatment process. This methodology was then applied to two aqueous hazardous waste treatment processes: packed tower aeration and activated sludge (aerobic) biological treatment. The designs achieving 95 percent reliability were compared with those designs based on conventional practice to determine their patterns of conservatism. Scoping-level treatment costs were also related to reliability levels for these treatment processes. The results indicate that the reliability levels for the physical/chemical treatment process (packed tower aeration) based on the deterministic safety factors range from 80 percent to over 99 percent, whereas those for the biological treatment process range from near 0 percent to over 99 percent, depending on the compound evaluated. Increases in reliability per unit increase in treatment costs are most pronounced at lower reliability levels (less than about 80 percent) than at the higher reliability levels (greater than 90 percent, indicating a point of diminishing returns. Additional research focused on process parameters that presently contain large uncertainties may reduce those uncertainties, with attending increases in the reliability levels of the treatment processes.

Aerobiosis↗

Test-retest reliability of static EMG scan configural profiling.

Measurement techniques or instruments are typically evaluated along the dimensions of reliability and validity. The focus of this investigation was to assess the test-retest reliability of a static EMG scan profile (sESP) method using 64 chronic pain participants. The test-retest interval was 30-33 days. Reliability coefficients were expressed using the Profile Similarity Coefficient (rp) in place of the more traditional Pearson Product Moment Correlation. sESP reliabilities were calculated for posture laterality for the head and neck, back, and overall profiles (head, neck, and back combined). The reliability coefficients ranged from .57 to .80. The back profile was the least reliable with a range of .55-.59 whereas the overall profiles were the most reliable, .78-.80. The analysis method was judged to be very conservative with its use of rp, a protracted intertest interval period, and weighting the data by their variances. These results can be viewed as setting the lower reliability limit for sESP.

Adult↗

A meta-analysis of job analysis reliability.

Average levels of interrater and intrarater reliability for job analysis data were investigated using meta-analysis. Forty-six studies and 299 estimates of reliability were cumulated. Data were categorized by specificity (generalized work activity or task data), source (incumbents, analysts, or technical experts), and descriptive scale (frequency, importance, difficulty, time-spent, and the Position Analysis Questionnaire). Task data initially produced higher estimates of interrater reliability than generalized work activity data and lower estimates of intrarater reliability. When estimates were corrected for scale length and number of raters by using the Spearman-Brown formula, task data had higher interrater and intrarater reliabilities. Incumbents displayed the lowest reliabilities. Scales of frequency and importance were the most reliable. Implications of these reliability levels for job analysis practice are discussed.

Employment↗

Tri-word presentations with phonemic scoring for practical high-reliability speech recognition assessment.

Speech recognition test reliability is optimized with 450 test items, and the Computer Assisted Speech Recognition Assessment (CASRA) test is a practical approach for achieving this goal by combining 50 presentations of 3 consonant-vowel nucleus-consonant (CNC) words each with phonemic scoring (S. A. Gelfand, 1998). However, optimized reliability might not be essential if reliability is as high as possible in light of practical constraints and what the clinician is trying to do with the results. The CASRA paradigm addresses these compromise goals with a reduced number of 3-word sets: 25 sets yield 25 (groups) x 3 (words) x 3 (phonemes) = 225 test items, 20 sets give 20 x 3 x 3 = 180 items, and 10 sets provide 10 x 3 x 3 = 90 items. This study addressed the empirical reliability of such an approach, and the extent to which results on shortened versions predict full-test scores. Test and retest scores were obtained for 10-, 20-, and 25-set versions of the CASRA for 144 participants with a wide range of hearing ability. For group data, first and second scores were highly correlated and not significantly different from each other for all 3 test sizes. Performance based on 20 and 25 sets accounted for roughly 97% of the variance of full (50-set) test scores, and scores based on 10 sets accounted for about 88% of the full-test variance. Individual test-retest reliability agreed with theoretical expectations based on 95% binomial confidence intervals. Cases outside the 95% confidence limits were 7.6% for 10 sets, and 3.5% for 20 and 25 sets with phoneme scoring, and 4.9% for 10 and 20 sets and 3.5% for 25 sets with word scoring. The shortened CASRA is a practical way to achieve improvements in reliability over traditional word tests. The 20-set version may approximate the strongest compromise when trying to shorten test size without appreciably reducing reliability for clinical purposes. However, the 10-set version is probably a more practical approach for routine use because it accounts for 88% of full-test variance, is more reliable than a traditional 75-word test, and does not appear to be subject to significant short-term learning effects.

Adolescent↗

Composite undergraduate clinical examinations: how should the components be combined to maximize reliability?

BACKGROUND: Clinical examinations increasingly consist of composite tests to assess all aspects of the curriculum recommended by the General Medical Council. SETTING: A final undergraduate medical school examination for 214 students. AIM: To estimate the overall reliability of a composite examination, the correlations between the tests, and the effect of differences in test length, number of items and weighting of the results on the reliability. METHOD: The examination consisted of four written and two clinical tests: multiple-choice questions (MCQ) test, extended matching questions (EMQ), short-answer questions (SAQ), essays, an objective structured clinical examination (OSCE) and history-taking long cases. Multivariate generalizability theory was used to estimate the composite reliability of the examination and the effects of item weighting and test length. RESULTS: The composite reliability of the examination was 0.77, if all tests contributed equally. Correlations between examination components varied, suggesting that different theoretically interpretable parameters of competence were being tested. Weighting tests according to items per test or total test time gave improved reliabilities of 0.93 and 0.81, respectively. Double weighting of the clinical component marginally affected the reliability (0.76). CONCLUSION: This composite final examination achieved an overall reliability sufficient for high-stakes decisions on student clinical competence. However, examination structure must be carefully planned and results combined with caution. Weighting according to number of items or test length significantly affected reliability. The components testing different aspects of knowledge and clinical skills must be carefully balanced to ensure both content validity and parity between items and test length.

Clinical Competence↗

Reliability of the MRCP(UK) Part I Examination, 1984-2001.

OBJECTIVES: To assess the reliability of the MRCP(UK) Part I Examination over the period 1984-2001, and to assess how the reliability is related to the difficulty of the examination (mean mark) and to the spread of the candidates' marks (standard deviation). METHODS: Retrospective analysis of the reliability (KR20) of the MRCP(UK) examination recorded in examination records for the 54 diets between 1984 and 2001. RESULTS: The reliability of the examination showed a mean value of 0.865 (SD 0.018, range 0.83-0.89). There were fluctuations in the reliability over time, and multiple regression showed that reliability was higher when the mean mark was relatively high, and when the standard deviation of the marks was high. CONCLUSIONS: The reliability of the MRCP(UK) Examination was maintained over the period 1984-2001. As theory predicted, the reliability was related to the average mark and to the spread of marks.

Clinical Competence↗

Inter-examiner and intra-examiner reliability of the standing flexion test.

The practice of musculoskeletal medicine requires the use of a wide variety of clinical examination procedures to establish a diagnosis, plan treatment, and monitor patient progress. Many of these examination procedures constitute a significant part of daily practice. Despite their extensive use, the reliability and validity of many of these assessment procedures remains questionable. The aim of this study was to determine the inter- and intra-examiner reliability of palpatory findings for the standing flexion test; one test for sacroiliac joint (SIJ) dysfunction. Nine examiners performed the standing flexion test on nine asymptomatic subjects. Inter-examiner reliability data, with a mean percentage agreement of 42% and a kappa coefficient of 0.052, demonstrated statistically insignificant reliability. Intra-examiner reliability data demonstrated a mean percentage agreement of 68% and a kappa coefficient of 0.46 indicating moderate reliability. These results suggest that the reliability of the standing flexion test as an indicator of SIJ dysfunction still remains questionable. Before this test can be relied upon as an accurate indicator of SIJ dysfunction it must undergo further research. This research must not only further standardize the procedure, but also ascertain reliability and validity.

Adult↗

Are chiropractic tests for the lumbo-pelvic spine reliable and valid? A systematic critical literature review.

OBJECTIVE: To systematically review the peer-reviewed literature about the reliability and validity of chiropractic tests used to determine the need for spinal manipulative therapy of the lumbo-pelvic spine, taking into account the quality of the studies. DATA SOURCES: The CHIROLARS database was searched for the years 1976 to 1995 with the following index terms: "chiropractic tests," "chiropractic adjusting technique," "motion palpation," "movement palpation," "leg length," "applied kinesiology," and "sacrooccipital technique." In addition, a manual search was performed at the libraries of the Nordic Institute of Chiropractic and Clinical Biomechanics, Odense, Denmark, and the Anglo-European College of Chiropractic, Bournemouth, United Kingdom. STUDY SELECTION: Studies pertaining to intraexaminer reliability, interexaminer reliability, and/or validity of chiropractic evaluation of the lumbo-pelvic spine were included. DATA EXTRACTION: Data quality were assessed independently by the two reviewers, with a quality score based on predefined methodologic criteria. Results of the studies were then evaluated in relation to quality. DATA SYNTHESIS: None of the tests studied had been sufficiently evaluated in relation to reliability and validity. Only tests for palpation for pain had consistently acceptable results. Motion palpation of the lumbar spine might be valid but showed poor reliability, whereas motion palpation of the sacroiliac joints seemed to be slightly reliable but was not shown to be valid. Measures of leg-length inequality seemed to correlate with radiographic measurements but consensus on method and interpretation is lacking. For the sacrooccipital technique, some evidence favors the validity of the arm-fossa test but the rest of the test regimen remains poorly documented. Documentation of applied kinesiology was not available. Palpation for muscle tension, palpation for misalignment, and visual inspection were either undocumented, unreliable, or not valid. CONCLUSION: The detection of the manipulative lesion in the lumbo-pelvic spine depends on valid and reliable tests. Because such tests have not been established, the presence of the manipulative lesion remains hypothetical. Great effort is needed to develop, establish, and enforce valid and reliable test procedures.

Chiropractic↗

Palpation of the upper thoracic spine: an observer reliability study.

OBJECTIVE: To assess the intraobserver reliability (in terms of hour-to-hour and day-to-day reliability) and the interobserver reliability with 3 palpation procedures for the detection of spinal biomechanic dysfunction in the upper 8 segments of the thoracic spine. DESIGN: A repeated-measures design was used in all substudies. SETTING: Department of Nuclear Medicine, Odense University Hospital, Denmark. PARTICIPANTS: Two chiropractors examined 29 patients and 27 subjects in the interobserver part and 1 chiropractor examined 14 patients and 15 subjects in the intraobserver studies. INTERVENTION: Three types of palpation were performed: Sitting motion palpation and prone motion palpation for biomechanic dysfunction and paraspinal palpation for tenderness. Each dimension was rated as "absent" or "present" for each segment. All examinations were performed according to a standard written procedure. RESULTS: Using an "expanded" definition of agreement that accepts small inaccuracies (+/-1 segment) in the numbering of spinal segments, we found--based on the pooled data from the thoracic spine--kappa values of 0.59 to 0.77 for the hour-to-hour and the day-to-day intraobserver reliability with all 3 palpation procedures. Kappa coefficients were 0.24 and 0.22 for the interobserver reliability with prone and sitting motion palpation and 0.67 and 0.70, respectively, with paraspinal palpation for tenderness. CONCLUSION: With expanded agreement we found good hour-to-hour and day-to-day intraobserver reliability with all 3 palpation procedures and good interobserver reliability for paraspinal tenderness. The interobserver reliability was unacceptably poor with prone and sitting motion palpation.

Adult↗

Intersession reliability for H-reflex measurements arising from the soleus, peroneal, and tibialis anterior musculature.

The Hoffman reflex (H-reflex) has been widely used throughout neuroscience research, as it allows for the assessment of alpha motoneuron excitability arising from a specific motoneuron pool. Recently, a protocol has been developed allowing for the simultaneous examination of the soleus, peroneal, and tibialis anterior motoneuron pools elicited from a single peripheral stimulus. In order for this protocol to be useful, the reliability of the measures must be established. The purpose of the current study was to determine the intersession reliability of the soleus, peroneal, and tibialis anterior H-reflexes and their corresponding M-waves elicited from a single stimulus to the sciatic nerve. Ten healthy neurologically sound individuals (age: 23 +/- 7 yrs; height: 175 +/- 12 cm; mass: 76 +/- 22 kg) volunteered to participate in this investigation. To obtain the measurements, the sciatic nerve was stimulated just prior to its bifurcation into the tibial and common peroneal nerves in the popliteal fossa. A 1-ms square wave pulse was delivered in 0.2 V increments until the maximum M wave was seen in each muscle. The maximum H-reflex and M-waves were collected from each muscle and their ratios calculated. Intersession reliability over 2 consecutive days was estimated using intraclass correlation coefficients (ICC [2.1]). Intersession reliability for the soleus Max H, Max M, and H:M ratio were 0.9953, 0.9514, and 0.9747, respectively. The peroneal reliability measurements were as follows: 0.9979 (Max H), 0.9924 (Max M), and 0.9664 (H:M ratio). Intersession reliability was 0.8591, 0.9968, and 0.7810 for the tibialis anterior Max H. Max M. and H:M ratio, respectively. These results indicate that the H-reflex measured from the soleus, peroneal, and tibialis anterior musculature elicited with a single peripheral stimulus to the sciatic nerve is reliable between sessions. This protocol allows the clinician/researcher to reliably investigate the alpha motoneuron excitability of multiple motoneuron pools about the ankle at a single point in time.

Adolescent↗

The reliability of distance run tests for children in grades K-4.

The purpose of this study was to determine test-retest reliability for the 1-mile, 3/4-mile, and 1/2-mile distance run/alk tests for children in Grades K-4. Fifty-one intact physical education classes were randomly assigned to one of the three distance run conditions. A total of 1,229 (621 boys, 608 girls) completed the test-retests in the fall (October), with 1,050 of these students (543 boys, 507 girls) repeating the tests in the spring (May). Results indicated that the 1-mile run/walk distance, as recommended for young children in most national test batteries, has acceptable intraclass reliability (.83 less than R less than .90) for both boys and girls in Grades 3 and 4, has minimal (fall) to acceptable (spring) reliability for Grade 2 students (.70 less than R less than .83), but is not reliable for children in Grades K and 1 (.34 less than R less than .56). The 1/2 mile was the only distance meeting minimal reliability standards for boys and girls in Grades K and 1 (.73 less than R less than .82). Results also indicated that reliability estimates remained fairly stable across gender and age groups from the fall to spring testing periods, with the exception of the noticeably improved values for Grade 2 students on the 1-mile run/walk test. Criterion-referenced reliability (P, percent agreement) was also estimated relative to Physical Best and Fitnessgram run/walk standards. Reliability coefficients for all age group standards were acceptable to high (.70 less than P less than .95), except for Fitnessgram standards for 5-year-old girls on the 1-mile test for both fall and spring and for 6-year-old boys and girls on the 1-mile test administered in the spring.

Child↗