PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Internal validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

[Quasi experimental evaluation of public health interventions (author's transl)].

The classic experiment, the randomised controlled trial, is the best known and most revered of evaluation research methods. Randomization in community-based intervention trials, however, is not always possible because of ethical problems arising from with holding the experimental treatment from the control groups or the difficulties in conducting experiments in field settings which do not approach controlled laboratory conditions. In such circumstances, quasi-experimental or observational designs must be used. Two major principles are involved in using quasi-experimental methods: (1) the logic for establishing causality between treatment and effect is the same as that for randomised experiments, but the problems of assessing causality or internal validity are greater, and (2) assessment of the external validity or generalizability of quasi-experimental findings crucial to the interpretation of results. Selected quasi-experimental designs using time series and comparison groups are described with examples from public health intervention trials where threats to internal validity have been assessed by using different analytic techniques or gathering additional evidence. Quasi-experimental evaluations are most useful when opportunities exist for testing rival hypotheses concerning the internal and external validity, or the findings can be used to complement true experiments.

Epidemiologic Methods↗

Feasibility and validity of International Classification of Diseases based case mix indices.

BACKGROUND: Severity of illness is an omnipresent confounder in health services research. Resource consumption can be applied as a proxy of severity. The most commonly cited hospital resource consumption measure is the case mix index (CMI) and the best-known illustration of the CMI is the Diagnosis Related Group (DRG) CMI used by Medicare in the U.S. For countries that do not have DRG type CMIs, the adjustment for severity has been troublesome for either reimbursement or research purposes. The research objective of this study is to ascertain the construct validity of CMIs derived from International Classification of Diseases (ICD) in comparison with DRG CMI. METHODS: The study population included 551 acute care hospitals in Taiwan and 2,462,006 inpatient reimbursement claims. The 18th version of GROUPER, the Medicare DRG classification software, was applied to Taiwan's 1998 National Health Insurance (NHI) inpatient claim data to derive the Medicare DRG CMI. The same weighting principles were then applied to determine the ICD principal diagnoses and procedures based costliness and length of stay (LOS) CMIs. Further analyses were conducted based on stratifications according to teaching status, accreditation levels, and ownership categories. RESULTS: The best ICD-based substitute for the DRG costliness CMI (DRGCMI) is the ICD principal diagnosis costliness CMI (ICDCMI-DC) in general and in most categories with Spearman's correlation coefficients ranging from 0.938-0.462. The highest correlation appeared in the non-profit sector. ICD procedure costliness CMI (ICDCMI-PC) outperformed ICDCMI-DC only at the medical center level, which consists of tertiary care hospitals and is more procedure intensive. CONCLUSION: The results of our study indicate that an ICD-based CMI can quite fairly approximate the DRGCMI, especially ICDCMI-DC. Therefore, substituting ICDs for DRGs in computing the CMI ought to be feasible and valid in countries that have not implemented DRGs.

Diagnosis-Related Groups↗

[Validation of French translation of the "Tinnitus Reaction Questionnaire", Wilson et al. 1991].

The present study proposes a validation of a french translation of the TRQ initially published by Wilson et al., for evaluating the psychological distress of tinnitus sufferers. The 26 items translated into french were used on a sample of 173 tinnitus sufferers, who also filled out the Mini-Mult, a short version of the MMPI proposed by Kincannon. Internal validity was demonstrated by strong correlations (i) between each item (except items 5 and 20) and total TRQ score (0.33 < or = r < or = 0.87, p < or = 0.0001), (ii) between each internal TRQ factor (0.58 < r < 0.81, p < 0.0001) and the others. Cronbach's alpha test also showed the questionnaire to have a good internal validity (alpha = 0.94). The external factors used for testing concurrent validity were the scores on depression, psychaesthenia and anxiety Mini-Mult scales. The strong correlations (one factor ANOVA and simple linear regression tests) between scores on depression and psychaesthenia scales and (1) each TRQ item, (2) each TRQ factor, (3) total TRQ scores, confirmed concurrent validity. Scores obtained on anxiety index showed high correlations only with TRQ score, factor 3 score and some TRQ items (most of them included in factor 3). The internal and concurrent validities of the French version of the TRQ justify the use of this questionnaire, with the reserve that items 5 and 20 appeared irrelevant for the measuring of tinnitus distress in French-speaking countries. Such a questionnaire should improve our knowledge of tinnitus' life-impact and enable detection of patients whose psychological distress necessitates rapid intervention.

Adaptation, Psychological↗

A Rationale for Using Synthetic Designs in Medical Education Research.

The extent to which the results of a study can be attributed to the intervention under investigation (i.e., internal validity) is an important consideration in interpreting study findings. There are many threats to the internal validity of designs frequently used in medical education research. Synthetic designs, which involve the integration of two or more weak designs, or the addition of design elements, may afford investigators greater control over confounding variables in medical education research. A rationale for using synthetic designs is presented and two examples of their use in medical education settings are examined. The concluding proposition is that synthetic designs allow investigators flexibility in planning research that is feasible in medical education settings. In addition, they may permit stronger causal inferences between interventions and results than traditional research designs.

Journal Article↗

Statistical aspects of clinical trials of antibiotics in acute infections.

Controlled clinical trials are important tools for evaluating antibiotics in acute infections. External and internal validity, definition of efficacy criteria, and size of the patient sample constitute special statistical problems in such studies. Critical issues regarding external validity pertain to the selection of patients and to the concept of consecutive patients. The internal validity of a study is influenced by the withdrawal of patients from the evaluation after randomization and the comparability of treatment groups with regard to prognostic factors. The definition of efficacy criteria on the basis of bacteriologic outcomes across control visits is not straightforward. Particularly, the evaluation of efficacy at the last follow-up visit must take into account the accumulated information rather than the cross-sectional information. The most common situation in comparative trials of antibiotics is that rather small differences in efficacy can be anticipated. Sometimes, the question at issue is the demonstration of antibiotic equivalence. For valid conclusions to be made in such situations, large samples must be used. A basic problem affecting many studies of antibiotics is that this criterion is not fulfilled.

Acute Disease↗

Validity of international, time trend, and migrant studies of dietary factors and disease risk.

A linear form relative risk model is used to identify circumstances in which various types of aggregate data lead to valid inferences on relative risk parameters. Upon making a random effects assumption, international or time trend data can lead to appropriate relative risk parameter estimation using iteratively reweighted least-squares procedures. Adequate confounding factor control, however, will typically require data on the distribution of confounding factors in each country or time period. For a simple interpretation of relative risk parameters one may also require data on the joint distribution of primary and confounding factors in each country or time period. Hence disease rate data need to be supplemented by dietary and risk factor survey data in order to avoid confounding bias. Measurement error in individual dietary assessment may, however, limit the ability to quantify the dependence of relative risk on dietary factors, unless the relative risk function is approximately linear in the dietary factors of interest. Most studies of migrant populations involve a comparison of migrant mortality rates with those of their countries of emigration and immigration, with little or no data collection on the dietary habits and risk factors of the migrants themselves. The potential of more comprehensive aggregate data and analytic migrant studies in the diet and disease area is briefly indicated. These issues and methods are illustrated using various types of data pertinent to the association between dietary fat and breast cancer.

Breast Neoplasms↗

Problems translating a questionnaire in a cross-cultural setting.

Questionnaires are widely used in veterinary epidemiological studies. With careful design, assessment and administration, questionnaires can collect accurate data. In cross-cultural settings, questionnaire design is complicated by the added step of translation. As part of a cross-sectional study of smallholder pig raisers in the Philippines, we assessed the importance of translation as a source of error in our questionnaire. The questionnaire contained 120 questions that collected data for entry into 361 separate data-entry fields in a computerised database. The translation was verified using the methods of back-translation and pilot-testing to identify and quantify errors of literal translation, omission and mistranslation. Errors were identified in 44 (37%) questions that obtained data for 147 (41%) data-entry fields. Although errors might have remained in the final translated version of the questionnaire, we have no doubt that the verification procedures substantially improved the internal validity of the questionnaire. It is important that veterinary epidemiologists use such procedures to check the internal validity of translated questionnaires.

Animal Husbandry↗

Testing the effects of nutrient deficiencies on behavioral performance.

The association between specific nutrient deficiencies and poor performance on behavioral tests has been documented for several nutrients. The determination of causality, however, remains elusive. This paper presents the essential criteria for a valid test of causality. Findings from experimental studies in which a nutritional treatment was randomly allocated can be summarized in a statistical statement about the probability that the nutrient treatment caused the behavioral response. Criteria for assessing the internal validity of these studies are examined in terms of whether alleviation of a nutrient deficiency did or did not produce a detectable behavioral response. The plausibility of such a causal inference is dependent on its congruency with known or theorized biological and behavioral mechanisms. External validity describes the extent to which inferences from internally valid studies may be applicable to other populations or circumstances. In addition to these scientific considerations, some of the ethical issues of nutrient-treatment trials are also discussed. All of these considerations provide a better basis for judging whether public health action would be worthwhile than do observed associations that could actually be due to other causes.

Behavior↗

Family planning field research projects: balancing internal against external validity.

This report discusses the experience of a two-year family planning and maternal/child health project in Nepal. Although the project was planned as an experimental field research endeavor, a series of unanticipated events repeatedly compromised the internal validity of the project and forced design changes. While unexpected events are common in the history of most field projects, they present the research evaluator with the fundamental dilemma of trying to maintain a high degree of internal validity without sacrificing external validity. Rigid research designs with tight control over the introduction and measurement of experimental variables may serve to increase internal validity but they may also create an atypical and artificial situation that fails to mirror real field conditions and thus threatens external validity.

Community Health Services↗

Disability in depression and back pain: evaluation of the World Health Organization Disability Assessment Schedule (WHO DAS II) in a primary care setting.

The World Health Organization Disability Assessment Schedule (WHO DAS II) is a new measure of disability based on the ICIDH-2 model of functioning and disability. This study evaluates the measurement properties of the WHO DAS II in two disorders commonly encountered in the primary care setting. Seventy-three patients with depression and 76 patients with back pain were interviewed at baseline and after 3 months of usual primary care. Internal validity, convergent validity, and responsiveness to change of the WHO DAS II were evaluated. The WHO DAS II had excellent internal validity and convergent validity in the primary care setting. The responsiveness to change of the WHO DAS II was comparable to that of the SF-36. The WHO DAS II appears to be a useful health status instrument for measuring the disability associated with both physical and mental disorders in the primary care setting. This instrument facilitates the use of the ICIDH-2 as a framework for evaluating activity limitations and participation.

Adult↗

Validity of International Classification of Diseases, Ninth Revision, Clinical Modification Codes for Acute Renal Failure.

Administrative and claims databases may be useful for the study of acute renal failure (ARF) and ARF that requires dialysis (ARF-D), but the validity of the corresponding diagnosis and procedure codes is unknown. The performance characteristics of International Classification of Diseases, Ninth Revision, Clinical Modification (ICD-9-CM) codes for ARF were assessed against serum creatinine-based definitions of ARF in 97,705 adult discharges from three Boston hospitals in 2004. For ARF-D, ICD-9-CM codes were compared with review of medical records in 150 patients with ARF-D and 150 control patients. As compared with a diagnostic standard of a 100% change in serum creatinine, ICD-9-CM codes for ARF had a sensitivity of 35.4%, specificity of 97.7%, positive predictive value of 47.9%, and negative predictive value of 96.1%. As compared with review of medical records, ICD-9-CM codes for ARF-D had positive predictive value of 94.0% and negative predictive value of 90.0%. It is concluded that administrative databases may be a powerful tool for the study of ARF, although the low sensitivity of ARF codes is an important caveat. The excellent performance characteristics of ICD-9-CM codes for ARF-D suggest that administrative data sets may be particularly well suited for research endeavors that involve patients with ARF-D.

Acute Kidney Injury↗

Validation of a nomogram predicting the probability of lymph node invasion among patients undergoing radical prostatectomy and an extended pelvic lymphadenectomy.

INTRODUCTION: Our goal was to develop and internally validate a nomogram for prediction of lymph node invasion (LNI) in patients with clinically localized prostate cancer undergoing extended pelvic lymphadenectomy (ePLND). METHODS: 602 consecutive patients (mean age 65.8 years) underwent an ePLND, where 10 or more nodes were removed. PSA was 1.1-49.9 (median 7.2). Clinical stages were: T1c in 55.6%, T2 in 41.4% and T3 in 3%. Biopsy Gleason sums were: 6 or less in 66%, 7 in 25.4%, 8-10 in 8.6%. Multivariate logistic regression models tested the association between all of the above predictors and LNI. Regression-based coefficients were used to develop a nomogram predicting LNI and 200 bootstrap resamples were used for internal validation. RESULTS: Mean number of lymph nodes removed was 17.1 (range 10-40). LNI was detected in 66 patients (11.0%). Univariate predictive accuracy for total PSA, clinical stage and biopsy Gleason sum was 63%, 58% and 73%, respectively. A nomogram based on clinical stage, PSA and Biopsy Gleason sum demonstrated bootstrap-corrected predictive accuracy of 76%. CONCLUSIONS: A nomogram based on pre-treatment PSA, clinical stage and biopsy Gleason sum can highly accurately predict LNI at ePLND.

Aged↗

Development of a questionnaire to measure quality of life in families with a child with food allergy.

BACKGROUND: Food allergy is potentially severe, affects approximately 5% of children, and requires numerous measures for food avoidance to maintain health. The effect of this disease on health-related quality of life (HRQL) has been documented by using generic instruments, but no disease-specific instrument is available. OBJECTIVE: To create a validated, food allergy-specific HRQL instrument to measure parental burden associated with having a child with food allergy: the Food Allergy Quality of Life-Parental Burden questionnaire. METHODS: After identification of 74 items affecting families with children with food allergy, 88 families were approached for effect scoring. Final items were generated by score results, elimination of redundancies, and content review. Resulting high-effect areas were queried for validation with a 7-point Likert scale. A final instrument including 17 items and 2 expectation of outcome questions was distributed to 352 families for validation. RESULTS: Areas of effect included family/social activities (restaurant meals, social activities, child care, vacation), school, time for meal preparation, health concerns, and emotional issues. Validation steps showed strong internal validity (Cronbach alpha, 0.95) and good correlation with expectation of outcome questions ( r = 0.412; P < .01) and scores on a generic HRQL instrument, the Children's Health Questionnaire-PF50 ( r = -0.36 to -0.4; P < .01). The instrument showed the ability to discriminate by disease burden: parents whose children had multiple (>2) food allergies were more affected than parents whose children had fewer allergies (scores, 3.1 vs 2.6; P < .001). CONCLUSIONS: The Food Allergy Quality of Life-Parental Burden demonstrates strong internal and cross-sectional validity. Its discriminative ability suggests that it will be a useful tool to measure outcomes in treatment studies of food allergy for children.

Adolescent↗

Development and external validation of an extended 10-core biopsy nomogram.

OBJECTIVES: To test the accuracy of a previously externally validated sextant biopsy nomogram in referred men exposed to > or =10 or more biopsy cores. Moreover, we explored the hypothesis that a more accurate predictive tool could be developed. METHODS: Previous nomogram predictors (age, digital rectal examination, prostate-specific antigen, and percent free PSA) were used to assess the accuracy of our previous nomogram in a cohort consisting of 2900 men referred for prostatic evaluation. Moreover, these variables were complemented with sampling density (SD) (i.e., ratio of gland volume and the number of planned biopsy cores) within multivariable logistic regression models (LRM) predicting presence of prostate cancer (pCA) on the initial 10 or more core biopsy. The LRMs were used to develop and internally validate (200 bootstrap resamples) a new nomogram in 1162 men from Hamburg, Germany. The LRMs' external validity was tested in three separate cohorts (Hamburg, n=582; Milan, n=961; Seattle, n=195). RESULTS: The contemporary external validation of the previously validated sextant nomogram demonstrated 70% accuracy. Internal validation of the new nomogram demonstrated 77% accuracy, and external cohorts demonstrated 73-76% accuracy. CONCLUSIONS: In the era of extended biopsy schemes, previously developed predictive models are less accurate in predicting the probability of pCA on initial biopsy. We developed a new tool that allows obtaining more accurate predictions. Moreover, before biopsy, it also allows defining the ideal ratio between gland volume and the number of planned biopsy cores that would yield the ideal biopsy rate.

Adult↗

Screening for drug abuse among adolescents in clinical and correctional settings using the Problem-Oriented Screening Instrument for Teenagers.

Recent research has indicated high rates of substance abuse among adolescents with emotional and behavioral disorders. Moreover, adolescents in clinical and correctional settings found to have comorbid disorders involving substance abuse experience higher morbidity and mortality rates when compared to adolescents having one or no condition. The present study examines the ability of the Problem-Oriented Screening Instrument for Teenagers (POSIT) to identify DSM-III-R-defined psychoactive substance use disorders among 342 adolescents aged 12-19 years. Participants were sampled from school, clinical, and correctional settings. Optimal-scale cut scores for drug abuse diagnosis classification were derived by a minimum loss function method that minimized false classifications. When using the optimal cut score of two for the total sample, the standard POSIT substance use/abuse scale obtained a drug abuse diagnosis classification accuracy of 84% with sensitivity and specificity ratios of 95% and 79%, respectively. The internal validity of the standard 17-item substance use/abuse scale was subsequently examined by principle component analysis, item analysis, and coefficient alpha. The internal validity analyses were conducted to determine if a shortened scale could be developed and yet retain acceptable classification accuracy. When using the optimal cut score of two for the total sample, the revised 11-item scale obtained a drug abuse diagnosis classification accuracy of 85% with sensitivity and specificity ratios of 91% and 82%, respectively. The results suggest that the POSIT can serve as a useful first-gate instrument to identify adolescents in need of further drug abuse assessment.

Affective Symptoms↗

Internal and external validity in two studies that compared treatment methods.

Research comparing the effectiveness of two treatments offers both strengths and weaknesses for occupational therapy. Although it is worthwhile to determine which of two treatments works best for a particular problem, methodological problems may arise that preclude a valid conclusion. To draw valid conclusions from research, criteria for internal validity and external validity must be satisfied. The two preceding articles in this issue are examples of studies that used between-groups experimental methodology to compare the effectiveness of two different treatments. This paper evaluates the above-mentioned studies on the basis of principles of internal and external validity. One of these studies (Jongbloed, Stacey, & Brighton, 1989) was truly experimental, whereas the other study (Groves & Rider, 1989) was quasi-experimental. Results from both studies were similar because, in each study, both of the treatment groups improved, but there were no significant differences between treatments. Absence of a true control group in both studies presented limitations on the conclusion that both treatments worked equally well.

Analysis of Variance↗

A Dynamic Nomogram to Predict Metabolic Dysfunction-Associated Fatty Liver Disease in Patients with Metabolic Syndrome.

BACKGROUND: Metabolic syndrome (MetS) involves multiple metabolic disorders. This study aimed to identify high-risk populations for metabolic dysfunction-associated fatty liver disease (MAFLD) in patients with MetS and to establish a dynamic predictive nomogram. METHODS: A total of 627 patients with MetS from six regions in Zhejiang Province were enrolled and categorized into MAFLD and non-MAFLD groups, then randomly assigned to training and validation sets at a ratio of 7:3. Independent predictors of MAFLD were identified using least absolute shrinkage and selection operator regression and multivariable logistic regression analyses. These predictors were then used to construct a dynamic nomogram. RESULTS: A total of 627 patients with MetS were included in the final analysis, of whom 77.0% (483/627) were diagnosed with MAFLD. Multivariable logistic regression analysis identified body mass index (BMI), waist circumference (WC), total cholesterol (TC), alanine aminotransferase (ALT), MetS-defined dysglycemia, and education level as independent risk factors for MAFLD. MetS-defined dysglycemia showed the highest odds ratio (OR) for MAFLD development [OR = 1.87, 95% confidence interval (CI): 1.07-3.29]. Although the number of MetS components and the metabolic syndrome score were significantly associated with MAFLD in univariate analysis, they were not independently associated with MAFLD in the multivariate model. A dynamic nomogram for predicting MAFLD risk in patients with MetS was developed and internally validated. The area under the receiver operating characteristic curve was 0.834 (95% CI: 0.787-0.880) in the training set and 0.839 (95% CI: 0.771-0.899) in the validation set, indicating strong predictive performance. Bootstrap internal validation demonstrated good agreement between predicted and observed outcomes in calibration curves. Decision curve analysis further indicated favorable clinical applicability of the nomogram. CONCLUSION: BMI, WC, TC, ALT, MetS-defined dysglycemia, and education level are independent risk factors for MAFLD. A dynamic nomogram for predicting MAFLD risk in patients with MetS was successfully developed and validated.

Humans↗

Prediction of target range of intact parathyroid hormone in hemodialysis patients with artificial neural network.

The application of artificial neural network (ANN) to predict outcome and explore potential relationships among clinical data is increasing being used in many clinical scenarios. The aim of this study was to validate whether an ANN is a useful tool for predicting the target range of plasma intact parathyroid hormone (iPTH) concentration in hemodialysis patients. An ANN was constructed with input variables collected retrospectively from an internal validation group (n = 129) of hemodialysis patients. Plasma iPTH was the dichotomous outcome variable, either target group (150 ng/L 300 ng/L). After internal validation, the ANN was prospectively tested in an external validation group (n = 32) of hemodialysis patients. The final ANN was a multilayer perceptron network with six predictors including age, diabetes, hypertension, and blood biochemistries (hemoglobin, albumin, calcium). The externally validated ANN provided excellent discrimination as appraised by area under the receiver operating characteristic curve (0.83 +/- 0.11, p = 0.003). The Hosmer-Lemeshow statistic was 5.02 (p= 0.08 > 0.05) which represented a good-fit calibration. These results suggest that an ANN, which is based on limited clinical data, is able to accurately forecast the target range of plasma iPTH concentration in hemodialysis patients.

Female↗