PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 793 records · Page 44Linked to original sources

Reliability, validity and factor structure of the upper limb subscale of the Motor Assessment Scale (UL-MAS) in adults following stroke.

PURPOSE: The upper limb items of the Motor Assessment Scale (MAS) have been shown to be a sensitive, valid and reliable measure of upper limb function for adults following stroke, however the validity and reliability of summing these items into an independent subscale has not yet been evaluated. The stability, internal consistency and construct validity of the upper limb MAS subscale (UL-MAS) was assessed in this study. METHOD: Twenty-seven inpatients following stroke (mean age = 67 years, range = 40 - 80) were sampled from an acute, inpatient rehabilitation setting. Patients were evaluated with 'Upper Arm Function', 'Hand Movements', and 'Advanced Hand Activities' items of the MAS by masked physiotherapists who had received standardized training in administration of the MAS. RESULTS: All items were explained by one factor on confirmatory factor analysis and correlated significantly with one another and with the composite (summed total) score. Internal consistency analysis produced a Cronbach's alpha of 0.83 which did not benefit from removal of any items. CONCLUSIONS: The acceptable internal consistency score obtained verifies the validity and reliability of using the UL-MAS as an independent scale. This study has also verified the construct validity of the UL-MAS subscale and provides a valuable extension of previous work, which together demonstrates the value of the UL-MAS as a responsive, valid and reliable measure of upper limb function in adults following stroke. The UL-MAS produced a single, composite score that could be interpreted as a total score for upper limb function in this population.

Adult↗

Validation of a bacteremia prediction model.

OBJECTIVE: To validate a previously published model for predicting bacteremia in hospitalized patients. DESIGN: Application of a published bacteremia prediction model to a prospective validation cohort of patients and comparison of its predictability to that found in the derivation cohort. SETTING: Urban, university-affiliated, 550-bed public hospital. PATIENTS: The validation cohort consisted of 342 patients with 559 blood culture episodes between October 14, 1992, and December 5, 1992. Each blood culture episode was scored based on the presence or absence of seven predictors of bacteremia and the findings compared with published results (derivation cohort). INTERVENTIONS: None. RESULTS: Application of the bacteremia prediction model to the validation cohort identified episodes with a low risk (3%) and a high risk (17%) for true bacteremia, similar to the findings in the derivation cohort (1% and 16%, respectively). Comparison of the predictions of the model in the two cohorts by receiver operator characteristic curve analysis revealed that the overall predictability of the model in the validation cohort was not as good as in the derivation cohort. CONCLUSIONS: Although the bacteremia prediction model did not perform as well overall in the validation cohort, the model still was able to clearly define two extreme groups: those with a low risk and those with a high risk for true bacteremia. This predictive capability may aid physicians in prescribing empiric antimicrobial therapy and also may be useful to hospital epidemiologists in assessing quality of care.

Adolescent↗

The predictive validity for mortality of the index of mobility-related limitation--results from the EPESE study.

BACKGROUND: self-reported disability reflects physical, environmental and attitudinal factors. We have previously reported the empirical identification of three simple tests to provide an index of (ambulatory) mobility-related physiological limitations (MOBLI). Evidence of the MOBLI 's responsiveness over time has been presented. Evidence of the predictive validity of the index is needed. OBJECTIVE: we aimed to measure the predictive validity for future mortality of the MOBLI and of self-reported mobility disability in a longitudinal cohort study. METHODS: data are from the sixth annual interview for two sites in the Established Populations for Epidemiologic Studies of the Elderly study. Included were 3,040 people, with information about self-reported walking difficulties, walking speed, time to complete five chair stands and peak expiratory flow. Age- and sex-adjusted death rates over a 4-year follow-up were computed, and proportional hazards regression models were used in the analysis. RESULTS: the MOBLI score is associated with subsequent mortality over 4 years, with evidence of a 'dose-response' relationship. The predictive value for mortality of the MOBLI score is similar to that of self-reported mobility disability in the studied population. CONCLUSIONS: the 'objective' MOBLI index has predictive validity as a continuous or dichotomised measure of the physiological component of mobility limitation in older populations. Given its empirical basis and face validity, predictive validity and responsiveness to change, MOBLI should be considered for local validation and use in epidemiological comparisons of older populations across countries or over longer periods of time.

Activities of Daily Living↗

Biochemical markers as additional measurements in dietary validity studies: application of the method of triads with examples from the European Prospective Investigation into Cancer and Nutrition.

The validity coefficient of dietary questionnaire measurements can be estimated from a triangular comparison between questionnaire, reference, and biochemical marker measurements with the method of triads. The method assumes that the measurements are linearly related to true intake and have independent random errors. We applied the method of triads to examples from the European Prospective Investigation into Cancer and Nutrition. In some examples, Heywood cases occurred, ie, the estimated validity coefficients were > 1 or the validity coefficients were not estimable. Such results are caused by random sampling fluctuations or violation of the model assumptions. One possible violation is a positive correlation between the random errors of questionnaire and reference measurements. We used a bootstrap method to estimate CIs for the validity coefficients. Validity studies with several hundred subjects, more accurate biochemical indicators of dietary intake, or both, are needed to estimate validity coefficients precisely and avoid complications with the bootstrap method.

Biomarkers↗

External validity in the assessment of intellectual development in adulthood.

The relation of intelligence and competence is discussed and external validity issues are examined for the dimensions of settings, measurement variables, treatment variables and experimental units. It is argued that external validity across situations and life stages cannot be obtained for any single measure of intellectual ability. External validity problems are exacerbated beyond young adulthood since single criterion goals comparable to that of educational aptitude in work with the young are not available, and tasks do not retain ecological validity when the situational context of the individual under study changes due to developmental progression and idiosyncratic modification of individual life situation and roles. External validity in adulthood must therefore be addressed by examining task-by-person-by-situation interfaces separately for different life stages and across cohort groupings. A major test construction and validation program is outlined, and examples are given showing how some of the aspects of such a program can be operationalized.

Cognition↗

An adapted version of the U.S. Department of Agriculture Food Insecurity module is a valid tool for assessing household food insecurity in Campinas, Brazil.

Until recently, Brazil did not have a national instrument with which to assess household food insecurity (FI). The objectives of this study were as follows: 1) to describe the process of adaptation and validation of the 15-item USDA FI module, and 2) to assess its validity in the city of Campinas. The USDA scale was translated into Portuguese and subsequently tested for content and face validity through content expert and focus groups made up of community members. This was followed by a quantitative validation based on a convenience (n = 125) and a representative (n = 847) sample. Key adaptations involved replacing the term "balanced meal" with "healthy and varied diet," to construct items as questions rather than statements, and to ensure that respondents understood that information would not be used to determine program eligibility. Chronbach's alpha was 0.91 and the scale item response curves were parallel across the 4 household income strata. FI severity level was strongly associated in a dose-response manner (P < 0.001) with income strata and the probability of daily intake of fruits, vegetables, meat/fish, and dairy. These findings were replicated in the 2 independent survey samples. Results indicate that the adapted version of the USDA food insecurity module is valid for the population of Campinas. This validation methodology has now been replicated in urban and/or rural areas of 4 additional states with similar results. Thus, Brazil now has a household food insecurity instrument that can be used to set national goals, to follow progress, and to evaluate its national hunger and poverty eradication programs.

Adult↗

Validity of derived measurements of leg-length differences obtained by use of a tape measure.

Determining the difference in the length of an individual's legs is often an important component of a musculoskeletal examination. Although measurements are easily obtained with a tape measure, the validity of these measurements is not known. The purpose of this study was to examine the validity of determinations of leg-length differences (LLDs) obtained by use of a specified tape measure method (TMM). Leg-length differences using the TMM and a radiographic technique were determined for 10 subjects who were candidates for clinical leg-length measurements and for 9 healthy control subjects. Validity of the TMM measurements was determined by assessing the degree of agreement between TMM-obtained LLDs and those obtained by the radiographic method. Validity estimates as determined by intraclass correlation coefficients (ICCs) were .770 for patients, .359 for healthy subjects, and .683 for all subjects. When the means of the two values obtained by use of the TMM were compared with the radiographic measurements, the ICCs were .852 for the patient group, .637 for the healthy subjects, and .793 for all subjects. This study suggests that TMM-derived LLD measurements are valid indicators of leg-length inequality and that the estimates of validity are improved by using the average of two determinations rather than a single determination.

Adult↗

Validity and reliability of joint indices. A longitudinal study in patients with recent onset rheumatoid arthritis.

This prospective longitudinal study evaluates the validity and reliability of joint indices (JIs) used to measure disease activity in patients with RA. From seven traditional JIs (Ritchie Articular Index (RAI), Modified RAI, Thompson score, 28 JI, 36 JI, total tender and total swollen joints) 37 'new' JIs were computed by considering three different characteristics of joint inflammation, tenderness, swelling and the combination of tenderness and swelling, and by grading for tenderness and/or weighting for surface area of the joints. Several aspects of validity were investigated, the construct (correlation with radiographic damage), correlational (correlation with ESR, general health) and criterion validity (correlation with a Health Assessment Questionnaire, discrimination between high and low disease activity). It was found that the validity and reliability of traditional JIs do not differ substantially. Graded JIs are almost always more valid than ungraded JIs. Weighted JIs are almost always less valid and reliable than unweighted JIs. Therefore no JI proved to be superior for measuring the disease activity under consideration. Taking simplicity into account the 28 JI, not graded and not weighted, was preferable.

Aged↗

Validity of three clinical performance assessments of internal medicine clerks.

PURPOSE: To analyze the construct validity of three methods to assess the clinical performances of internal medicine clerks. METHOD: A multitrait-multimethod (MTMM) study was conducted at the Case Western Reserve University School of Medicine to determine the convergent and divergent validity of a clinical evaluation form (CEF) completed by faculty and residents, an objective structured clinical examination (OSCE), and the medicine subject test of the National Board of Medical Examiners. Three traits were involved in the analysis: clinical skills, knowledge, and personal characteristics. A correlation matrix was computed for 410 third-year students who completed the clerkship between August 1988 and July 1991. RESULTS: There was a significant (p < .01) convergence of the four correlations that assessed the same traits by using different methods. However, the four convergent correlations were of moderate magnitude (ranging from .29 to .47). Divergent validity was assessed by comparing the magnitudes of the convergence correlations with the magnitudes of correlations among unrelated assessments (i.e., different traits by different methods). Seven of nine possible coefficients were smaller than the convergent coefficients, suggesting evidence of divergent validity. A significant CEF method effect was identified. CONCLUSION: There was convergent validity and some evidence of divergent validity with a significant method effect. The findings were similar for correlations corrected for attenuation. Four conclusions were reached: (1) the reliability of the OSCE must be improved, (2) the CEF ratings must be redesigned to further discriminate among the specific traits assessed, (3) additional methods to assess personal characteristics must be instituted, and (4) several assessment methods should be used to evaluate individual student performances.

Clinical Clerkship↗

Validating the standardized-patient assessment administered to medical students in the New York City Consortium.

PURPOSE: To test the criterion validity of existing standardized-patient (SP)-examination scores using global ratings by a panel of faculty-physician observers as the gold-standard criterion; to determine whether such ratings can provide a reliable gold-standard criterion to be used for validity-related research; and to encourage the use of these gold-standard ratings for validation research and examination development, including scoring and standard setting, and for enhancing understanding of the clinical competence construct. METHOD: Five faculty physicians independently observed and rated videotaped performances of 44 students from one medical school on the seven SP cases that make up the fourth-year assessment administered at The Morchand Center of Mount Sinai School of Medicine to students in the eight member schools in the new York City Consortium. RESULTS: The validity coefficients showed correlations between scores on the examination and the overall ratings ranging from .60 to .70. The reliability coefficients for ratings of overall examination performance reached the commonly recommended .80 level and were very close at the case level, with interrater reliabilities generally in the .70 to .80 range. CONCLUSION: The results are encouraging, with validity coefficients high enough to warrant optimism about the possibility of increasing them to the recommended .80 level, based on further studies to identify those measurable performance characteristics that most reflect the gold-standard ratings. The high interrater reliabilities indicate that faculty-physician ratings of performance on SP cases and examinations can or may be able to provide a reliable gold standard for validating and refining SP assessment.

Clinical Clerkship↗

Reliability and validity of the Frail Elderly Functional Assessment questionnaire.

Measuring functional activity for elderly at very low functional levels remains a challenge because many functional instruments have not been standardized in a frail elderly population. The Frail Elderly Functional Assessment questionnaire (FEFA) is a 19-item, interviewer-administered questionnaire designed to assess function in frail elderly at a very low activity level. The purpose of this study was to determine the reliability and validity of this instrument in a frail elderly population. Two groups of subjects over 65 yr old were selected to test the reliability and validity of this questionnaire. Test-retest reliability was determined by correlating the responses of 29 homebound (including nursing home-bound) subjects who answered the questionnaire on two occasions 2 wk apart. To assess the validity of the FEFA, the questionnaire was administered to 23 frail, homebound (including nursing home-bound) elderly subjects who had a Mini-Mental State Examination score of > or = 18. Validity was determined by correlating patient responses to direct observations by the investigators of tasks addressed in the questionnaire. Correlation was also determined against the Katz's Activity of Daily Living index, Lawton's Instrumental Activity of Daily Living index, and the Barthel index. The reliability coefficient was 0.82. Correlation between the FEFA questionnaire and direct observation of questionnaire task performance was 0.90. Construct validity against the Katz's Activity of Daily Living, Lawton's Instrumental Activity of Daily Living, and the Barthel index showed correlations of 0.86, 0.67 and 0.91, respectively. Initial data indicate that the FEFA is a valid and reliable instrument that may be useful in assessing function in frail elderly people.

Activities of Daily Living↗

Reliability and validity of utilization review criteria. Appropriateness Evaluation Protocol, Standardized Medreview Instrument, and Intensity-Severity-Discharge criteria.

A study was conducted to assess the reliability and validity of the Appropriateness Evaluation Protocol (AEP), the Standardized Medreview Instrument (SMI) and the Intensity-Severity-Discharge criteria set (ISD), three utilization review instruments used to determine whether inpatient care is required. Reliability and validity were assessed for retrospective application of these instruments to charts of a sample of 119 medical cases from 21 hospitals in the state of Michigan. The reliability of each instrument was determined by having the instrument applied by two different nurse reviewers to each hospital record. Results indicated that the AEP and ISD were moderately reliable, while the SMI had low reliability. The validity of each instrument was tested by comparing the judgments of nurse reviewers using the instruments with the judgment of a panel of physicians. The AEP and ISD were found to be moderately valid and the SMI was found to have low validity. Results suggested that the SMI should not be used. The modest level of validity of the other two instruments suggests that payment should never be denied on the basis of the instrument alone. Payment should be denied only if a physician confirms the judgment based on the instrument that inpatient care was not required.

Health Maintenance Organizations↗

The MOS 36-Item Short Form Health Survey: reliability, validity, and preliminary findings in schizophrenic outpatients.

OBJECTIVES: The authors test the reliability and validity of the Medical Outcomes Study Short Form 36-Item Health Survey (SF-36) as a written, self-administered survey in outpatients with chronic schizophrenia. METHODS: Thirty-six schizophrenic outpatients completed a written and oral form of the SF-36. A psychiatrist rated the patients using the Brief Psychiatric Rating Scale to determine severity of psychopathology. Cognitive functioning and academic achievement were also assessed. Internal consistency, test-retest reliability, concurrent and discriminative validity of the oral and written versions were determined. RESULTS: The SF-36 in both forms was shown to have good internal consistency, stability, and concurrent validity. The mental health SF-36 subscales had poor discriminant validity, compared with the physical functioning scale that demonstrated good discriminant validity. CONCLUSIONS: The validity of using the written form of the SF-36 on a sample of patients with chronic mental illness was demonstrated. The SF-36 appears to be an appropriate outcome measure for changes in physical and role functioning in consumers of outpatient mental health programs.

Adult↗

An empirical assessment of the validity of explicit and implicit process-of-care criteria for quality assessment.

OBJECTIVE: To evaluate the validity of three criteria-based methods of quality assessment: unit weighted explicit process-of-care criteria; differentially weighted explicit process-of-care criteria; and structured implicit process-of-care criteria. METHODS: The three methods were applied to records of index hospitalizations in a study of unplanned readmission involving roughly 2,500 patients with one of three diagnoses treated at 12 Veterans Affairs hospitals. Convergent validity among the three methods was estimated using Spearman rank correlation. Predictive validity was evaluated by comparing process-of-care scores between patients who were or were not subsequently readmitted within 14 days. RESULTS: The three methods displayed high convergent validity and substantial predictive validity. Index-stay mean scores, using explicit criteria, were generally lower in patients subsequently readmitted, and differences between readmitted and nonreadmitted patients achieved statistical significance as follows: mean readiness-for-discharge scores were significantly lower in patients with heart failure or with diabetes who were readmitted; and mean admission work-up scores were significantly lower in patients with lung disease who were readmitted. Scores derived from the structured implicit review were lower in patients eventually readmitted but significantly so only in diabetics. CONCLUSIONS: These three criteria-based methods of assessing process of care appear to be measuring the same construct, presumably "quality of care." Both the explicit and implicit methods had substantial validity, but the explicit method is preferable. In this study, as in others, it had greater inter-rater reliability.

Case-Control Studies↗

Development and validation of a grading system for the quality of cost-effectiveness studies.

PURPOSE: To provide a practical quantitative tool for appraising the quality of cost-effectiveness (CE) studies. METHODS: A committee comprising [corrected] of health economists selected a set of criteria for the instrument from an item pool. Data collected with a conjoint analysis survey on 120 international health economists were used to estimate weights for each criterion with a random effects regression model. To validate the grading system, a survey was sent to 60 individuals with health economics expertise. Participants first rated the quality of three CE studies on a visual analogue scale, and then evaluated each study using the grading system. Spearman rho and Wilcoxon tests were used to detect convergent validity and analysis of covariance (ANCOVA) for discriminant validity. Agreement between the global rating by experts and the grading system was also examined. RESULTS: Sixteen criteria were selected. Their coefficient estimates ranged from 1.2 to 8.9, with a sum of 93.5 on a 100-point scale. The only insignificant criterion was "use of subgroup analyses." Both convergent validity and discriminant validity of the grading system were shown by the results of the Spearman rho (correlation coefficient = 0.78, P < 0.0001), Wilcoxon test (P = 0.53), and ANCOVA (F(3,146) = 5.97, p = 0.001). The grading system had good agreement with global rating by experts. CONCLUSIONS: The instrument appears to be simple, internally consistent, and valid for measuring the perceived quality of CE studies. Applicability for use in clinical and resource allocation decision-making deserves further study.

Cost-Benefit Analysis↗

Validity and factorial invariance of the Social Physique Anxiety Scale.

PURPOSE: The present study 1) tested whether the two-factor model to the 12-item Social Physique Anxiety Scale (SPAS) was substantively meaningful or a methodological artifact representing positively and negatively worded items, 2) assessed the factorial validity of the nine-item unidimensional model to the SPAS, 3) examined whether modifying the number of SPAS items would improve the factorial validity. 4) evaluated the factorial invariance of the SPAS across gender, and 5) explored the construct validity of SPAS scores. METHODS: Female (N = 146) and male (N = 166) college students (22.2 +/- 4.0 yr) in lecture (N = 103) and physical activity (N = 209) courses completed the SPAS, Physical Self-Efficacy Scale (PSES), Surveillance subscale of the Objectified Body Consciousness Scale (S-OBCS), and short form of the Marlowe-Crowne Social Desirability Scale (SDS-C). RESULTS: Confirmatory factor analyses (CFA) revealed that the two-factor model to the 12-item SPAS was a methodological artifact representing positively and negatively worded items. CFA indicated that the nine-item unidimensional model represented an acceptable fit to the SPAS, but it also could be improved. Modifications based on standardized residuals and item content led to the removal of two items and a seven-item unidimensional solution to the SPAS. The nine- and seven-item models demonstrated factorial invariance across gender. Correlation analyses between nine- and seven-item SPAS scores to PSES, S-OBCS, and SDS-C provided support for the construct validity. CONCLUSIONS: The nine- and seven-item unidimensional models to the SPAS demonstrated evidence of factorial validity, factorial invariance, and construct validity; the two-factor model to the SPAS represented a methodological artifact.

Adolescent↗

Reliability and validity of the Borg and OMNI rating of perceived exertion scales in adolescent girls.

PURPOSE: To examine the reliability and validity of the Borg and OMNI rating of perceived exertion (RPE) scales in adolescent girls during treadmill exercise. METHODS: Adolescent girls (N = 57, age = 15.3+/-1.5 yr) were randomly assigned to use an RPE scale (Borg or OMNI) during one of three treadmill submaximal exercise conditions (walking, walking uphill, or jogging). After RPE assessment, exercise intensity was increased until participants achieved volitional exhaustion (O2max). Expired respiratory gases and heart rate (HR) were measured continuously during exercise. Reliability of the RPE scales was assessed using ANOVA (intraclass) and Spearman-Brown prophecy formula (single trial) measures. Validity estimates were calculated using Pearson Product Moment correlations, with % HRmax and % O2max as criterion measures. RESULTS: Intraclass and single-trial reliability estimates were higher for the OMNI (r(xx) = 0.95 and r(kk) = 0.91, respectively) compared with the Borg (r(xx) = 0.78 and r(kk) = 0.64, respectively) RPE scale. Validity estimates were also higher for the OMNI scale compared with the Borg scale. Validity coefficients (r(xy)) for %HRmax and %O2max comparisons were 0.86 and 0.89, respectively, for the OMNI, compared with 0.66 and 0.70, respectively, for the Borg. CONCLUSION: The OMNI cycle pictorial scale was found to be reliable and valid for use with adolescent girls. It also appears to be more reliable and valid than the Borg scale for use in this population during treadmill exercise.

Adolescent↗

Assessment of reliability and validity of the behavioral observation record for developmental care.

Due to time constraints, clinicians are rarely able to carry out neurobehavioral assessments that use the Naturalistic Observation of Newborn Behavior Instrument. The content validity, interrater reliability, and criterion-related validity for a less time-consuming instrument, the Modified Infant Behavioral Observation Record (MIBOR) for developmental care was evaluated in this study. Eight developmental care specialists evaluated the MIBOR for content validity. Fifteen infants (birth weight < 1,500 g) were observed to determine interrater reliability, and three were observed to evaluate criterion-related validity. The content validity of the MIBOR, as determined by average congruence, was 96.9%. Interrater reliability for each developmental care subsystem ranged from 67% to 98%. Three of the four subsystems on the MIBOR achieved criterion-related validity, achieving an agreement of r = .60.

Attention↗