PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “External validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Assessing socioeconomic status in adolescents: the validity of a home affluence scale.

STUDY OBJECTIVE: To examine the completion rate, internal reliability, and external validity of a home affluence scale based on adolescents' reports of material circumstances in the home as a measure of family socioeconomic status. DESIGN: Cross sectional survey. SETTING: Data were collected from a school based study in seven schools in the north of England Cheshire over a five month period from September 1999 to January 2000. PARTICIPANTS: 1824 students (1248 girls, 567 boys) aged 13-15 years who were attending normal classes in Years 9 and 10 in 7 schools on the days of data collection. MAIN RESULTS: Comparatively poor completion rates were found for questions on parental education and occupation while material deprivation items had much higher completion rates. There was evidence that students with poorer material circumstances were less able to report parental education and occupation whereas material based questions showed less bias. A home affluence scale composed of material items was found to have adequate internal reliability and good external validity. CONCLUSIONS: A home affluence scale based on material markers provides a useful alternative in assessing family affluence in adolescents. Additionally, it prevents exclusion of those less materially well off adolescents who fail to complete conventional socioeconomic status items.

Adolescent↗

Evaluation and application of models for the prediction of ready biodegradability in the MITI-I test.

Three existing models and one newly developed model for the prediction of ready biodegradability of organic compounds are evaluated by comparing the descriptors they use, and the consistency of the models when applied to the set of High Production Volume Chemicals (HPVC) in the European Union. Linear regression models developed for the OECD showed the best performance in the external validation (84.7% correct), although comparison with the other three models is flawed because of the class specificity of these models. With these models 567 of the 894 compounds could be predicted in the validation. The multivariate statistical model showed the best performance in the external validation (82.7% correct) combined with the broadest applicability of the model. The evaluation of the predictions of the models for the HPVC shows that all models are highly consistent in their prediction of not-ready biodegradability, but much less consistency is seen in the prediction of ready biodegradability. This complies with the observation that all 4 models show better performance in their predictions of not-ready biodegradability.

Analysis of Variance↗

Validation of a tool to safely triage selected patients with chest pain to unmonitored beds.

OBJECTIVE: To externally validate a chest pain protocol that triages low risk patients with chest pain to an unmonitored bed. METHODS: Retrospective study of all patients admitted from the emergency department of a tertiary referral public teaching hospital with an admission diagnosis of 'unstable angina' or suspected ischemic chest pain. Data was collected on adverse outcomes and analysed on the basis of intention-to-treat according to the chest pain protocol. RESULTS: There were no life-threatening arrhythmias, cardiac arrests or deaths within the first 72 h of admission in the group assigned to an unmonitored bed by the chest pain protocol ([0/244]; 0.0%: 95% confidence interval 0.0-1.5%). Four patients had an uncomplicated myocardial infarction, two patients had recurrent ischemic chest pain and one patient developed acute pulmonary oedema ([7/244]; 2.9%: 95% confidence interval 1.2-5.8%). CONCLUSION: This retrospective study externally validated the chest pain protocol. Care in a monitored bed would not have altered outcomes for patients triaged to an unmonitored bed by the chest pain protocol. Compared to current guidelines, application of the chest pain protocol could increase the availability of monitored beds.

Aged↗

Randomized and non-randomized patients in clinical trials: experiences with comprehensive cohort studies.

In clinical research, randomized trials are widely accepted as the definitive method of evaluating the efficacy of therapies. Random assignment of patients to treatment ensures internal validity of the comparison of new treatments with controls. An assessment of external validity can best be achieved by comparing the randomized study sample to the population of patients who met the eligibility criteria but did not consent to randomization. The Comprehensive Cohort Study (CCS) is designed to recruit all patients fulfilling the clinical eligibility criteria regardless of their consent to randomization. The CCS concept was adopted in the major clinical trials of the German Breast Cancer Study Group (GBSG) conducted between 1983 and 1989. In this period 124 centres recruited 2084 patients in three clinical trials. 734 (35 per cent) of these patients accepted being randomized, while 1350 (65 per cent) chose one of the treatments under study; the randomization rates differed remarkably between trials. In this paper we examine the representativeness of the randomized patients in the three trials. Based on a median follow-up of about 5 years we present results on the external validity of the treatment effects estimated in the randomized patients by means of Cox's proportional hazards model and compare them between trials. We discuss advantages and disadvantages of the CCS design and conclude that its use is only justified under extraordinary circumstances.

Breast Neoplasms↗

Pros and cons of permutation tests in clinical trials.

Hypothesis testing, in which the null hypothesis specifies no difference between treatment groups, is an important tool in the assessment of new medical interventions. For randomized clinical trials, permutation tests that reflect the actual randomization are design-based analyses for such hypotheses. This means that only such design-based permutation tests can ensure internal validity, without which external validity is irrelevant. However, because of the conservatism of permutation tests, the virtues of permutation tests continue to be debated in the literature, and conclusions are generally of the type that permutation tests should always be used or permutation tests should never be used. A better conclusion might be that there are situations in which permutation tests should be used, and other situations in which permutation tests should not be used. This approach opens the door to broader agreement, but begs the obvious question of when to use permutation tests. We consider this issue from a variety of perspectives, and conclude that permutation tests are ideal to study efficacy in a randomized clinical trial which compares, in a heterogeneous patient population, two or more treatments, each of which may be most effective in some patients, when the primary analysis does not adjust for covariates. We propose the p-value interval as a novel measure of the conservatism of a permutation test that can be defined independently of the significance level. This p-value interval can be used to ensure that the permutation test have both good global power and an acceptable degree of conservatism.

Humans↗

Factor structure and clinical validity of competing models of positive symptoms in schizophrenia.

BACKGROUND: The factor structure of four competing models of positive symptoms and their clinical validity was studied in a sample of 253 schizophrenia inpatients. METHODS: The following models were tested using confirmatory factor analysis: a one-dimension severity model, a two-dimension model comprising a psychosis factor and a disorganization factor, a four-dimension model based on the Scale for the Assessment of Positive Symptoms (SAPS) structure in subscales, and a five-dimension model derived from the previous one by further differentiating Schneiderian delusions from non-Schneiderian ones. RESULTS: More complex multifactorial models fit the data better than simpler models. The five-dimension model was the best adjusted (goodness of fit index = .844, nonnormed fit index = .812, normed fit index = .728). Whereas the one-dimension model did not display significant association with the clinical variables, multidimensional models were related to age at onset and illness severity. The two-dimension model captured well the clinical correlates of the more complex models. CONCLUSION: None of the tested models showed good fit to the data. The one-dimension model displayed both poor factor validity and poor external validity; therefore, research relying on the SAPS total score may reach misleading conclusions.

Adult↗

The future of pharmacoeconomics: bridging science and practice.

In the context of new challenges, issues facing the science, practice, and future of pharmacoeconomics will be discussed. Certain methodologic weaknesses have been observed in published pharmacoeconomic studies, and compromises need to be made between developing an "ideal" method and allowing a study to remain practicable. The objective is to reach a balance between clinical trial-based studies and projective models; trials have high internal validity but low external validity, while models can help explore relevance to real-life settings. Cross-national differences also have an important impact on pharmacoeconomic data; however, using some basic standardized guidelines results from pharmacoeconomic studies may be generalized to other settings. The use of pharmacoeconomic results by decision makers in the United Kingdom has been restrained by unclear priorities within their authority and by the limited availability of credible studies. The future of pharmacoeconomics lies in developing both trial-based and modeling studies, improving their credibility, and meeting the needs of decision makers.

Clinical Trials as Topic↗

Developing a measure of attitudes: the holistic complementary and alternative medicine questionnaire.

We have developed an 11-item scale, the Holistic Complementary and Alternative Medicine Questionnaire (HCAMQ). Six of the HCAMQ items relate to beliefs about the scientific validity of complementary and alternative medicine (CAM), and five to beliefs about holistic health (HH). The HCAMQ was completed by 50 patients attending a CAM clinic and 50 attending rheumatology outpatients; the former completed it twice. Factor analysis (oblique rotation) showed that the CAM and HH items measured distinct but related constructs. The HCAMQ has good test retest reliability (r=0.86, 0.82 and 0.77 for the total, CAM subscale and HH subscale, respectively). The individuals attending CAM clinics were significantly more positive on the CAM but not the HH subscale of the HCAMQ and also used less antibiotics than those attending rheumatology outpatients. Positivity towards CAM on the total HCAMQ and subscales was significantly associated with lower age, increased vitamin use, reduced painkiller use, and, other than on the HH subscale, less antibiotic use. The reason why the HH subscale failed to distinguish between the two patient groups or predict less antibiotic use is unknown. The HCAMQ appears to have good internal validity, but its external validity remains to be established.

Age Factors↗

ADME evaluation in drug discovery. 3. Modeling blood-brain barrier partitioning using simple molecular descriptors.

In this paper, QSPR models were developed for in vivo blood-brain partitioning data (logBB) of a large data set consisting of 115 diverse organic compounds. The best model is based on three descriptors: n-octanol/water partition coefficient calculated using the SLOGP approach, logP; high-charged polar surface areas based on the Gasteiger partial charges, HCPSA, and the excessive molecular weight larger than 360, MW(360). The model bears good statistical significance, n = 78, r = 0.88, q = 0.86, s = 0.36, F = 81.5. The actual prediction potential of the model was validated through two external validation sets of 37 diverse compounds. The predicted results demonstrate that the model bears better prediction potential than many other models and can be used for logBB estimations for drug and drug-like molecules. Comparison of several logP calculation approaches suggests that logP calculated by SLOGP can be used as a significant descriptor for the prediction of molecular transport properties because SLOGP gives the most similar results with CLOGP. The QSPR model indicates that larger polar surface areas have a more negative contribution to logBB, but the absolute partial charges on the atoms surrounded by the polar surfaces should be larger than 0.10|e|. Meanwhile, tight junction membranes limit the size of hydrophilic molecules that can cross the membrane with a molecular weight of approximately 360, because when a molecule's weight is larger than 360 it shows a negative contribution to logBB. The computations of molecular surface, partial charges, logP, and logBB have been accomplished using a program called Drug-BB. Moreover, to improve the efficiency of the computations of logP, we made an extensive reparametrization of SLOGP, and the newly developed SLOGP model is only based on simple atomic addition. Further, we developed a set of parameters to calculate the topological polar surface area (TPSA), thus the high-charged topological polar surface area (HCTPSA) could be estimated from the 2D connection information of a molecule. Adopting the new strategies, the estimations of logP, HCTPSA, and logBB are only based on the topological structure of a molecule and therefore, can be used for fast screening of virtual libraries having millions of molecules.

Absorption↗

Nomogram for overall survival of patients with progressive metastatic prostate cancer after castration.

PURPOSE: To develop a pretreatment prognostic model for survival of patients with progressive metastatic prostate cancer after castration using parameters that are measured during routine clinical management. PATIENTS AND METHODS: Pretreatment clinical and biochemical determinants from 409 patients enrolled onto 19 consecutive therapeutic protocols from June 1989 through January 2000 were evaluated. The factors selected were age, Karnofsky performance status (KPS), hemoglobin (HGB), prostate-specific antigen (PSA), lactate dehydrogenase (LDH), alkaline phosphatase (ALK), and albumin. These factors were combined in an accelerated failure time regression model to produce a nomogram to predict median, 1-year, and 2-year survival. The nomogram was validated internally and externally using data from a multicenter randomized trial of suramin plus hydrocortisone versus hydrocortisone alone. RESULTS: The median survival of the entire group was 15.8 months (range, 0.9 to 77.8 months); 87% have died. In multivariable analysis, KPS, HGB, ALK, albumin, and LDH were significantly associated with survival (P <.05), whereas age and PSA were not. All seven factors were included in the nomogram. When applied to the external validation data set, the nomogram achieved a concordance index of 0.67. Calibration plots suggested that the nomogram was well calibrated for all predictions. CONCLUSION: A nomogram derived from pretreatment parameters that are measured on a routine basis was constructed. It can be used to predict the median, 1-year, and 2-year survival of patients with progressive castrate metastatic disease with reasonable accuracy. The information is useful to assess prognosis, guide treatment selection, and design clinical trials.

Adult↗

Empirically supported psychosocial interventions for children: an overview.

Discusses issues related to the identification of psychosocial interventions for children that have demonstrated efficacy. Recent debate concerning differences between clinical trials research and clinical practice is summarized, including the tradeoff between interpretability (internal validity) and generalizability (external validity) of outcome studies. This article serves as an introduction to the special issue containing articles that have as their focus the identification of empirically supported psychosocial interventions for children as part of a task force. The article provides an overview of the history, agenda, and methodology used by the task force to define and identify specific empirically supported interventions for children with specific disorders. Whereas a number of well-established or probably efficacious interventions are identified within the series, more work directed at closing the gap between research and practice is needed.

Adolescent↗

Analysis of randomized and nonrandomized patients in clinical trials using the comprehensive cohort follow-up study design.

In clinical research, randomized trials are widely accepted as the definitive method of evaluating the efficacy of therapies. The random assignment of patients to their treatment ensures the internal validity of the comparison of new treatments with controls. An assessment of the external validity of trial results can best be achieved by comparing the study population to the population of patients who met the eligibility criteria but did not consent to randomization. A part of the data of the Coronary Artery Surgery Study (CASS), in which coronary artery bypass surgery is compared to conventional medical therapy in patients with coronary artery disease, is used to illustrate a strategy of multivariate analysis of randomized and nonrandomized patients which allows an investigation of both internal and external validity. The method used Cox's proportional hazards regression model with inclusion of covariates for randomization status and corresponding interactions in addition to the usual covariates for treatment and the important prognostic factors.

Cohort Studies↗

[Validity of the clinical prediction rule for the diagnosis of renal arterial stenosis in hypertensive patients resistant to treatment].

PURPOSE: To perform an external validation of the clinical prediction rule established by Krijnen et al. (Ann Intern Med 1998; 129: 705-11) designed to identify renal artery stenoses (RAS) in hypertensive patients. METHODS: We included 102 patients with a refractory hypertension treated with at least two antihypertensive drugs. All subjects had the research of RAS by renal angiography, or angio-computed tomography, or doppler ultrasound. Probability to detect RAS was calculated with Krijnen's algorithm (Pre-test probability) from the following parameters: age, smoking status, diffuse atherosclerosis, recent hypertension (< 2 y), obesity (BMI > 25), abdominal bruit, hypercholesterolemia (> 6.5 mmol/L), creatinine. ROC curves were plotted for each pre-test probability value. A "post-test probability" was obtained from the likelihood ratio calculated at each pre-test probability level. RESULTS: RAS prevalence in this population was 49%. Area under the ROC curve was 0.79 and Youden index was maximal for a pre-test probability of 15%. Maximal likelihood ratio was obtained for a pre-test probability of 46%. Table shows post-test probability as a function of pre-test probability obtained with Krijnen's algorithm. [table: see text] CONCLUSION: Krijnen's algorithm is valid in a population of resistant hypertensives treated with a bi-therapy. This external validation obtained on a population with a high prevalence of RAS should also be tested on a population with a lower prevalence of SAR.

Age Factors↗

[Theory and practice in medical specialization. I. An instrument for measuring learning strategies].

We present the development and validation of a measurement instrument intended to estimate the degree of vinculation between the theoretical and practical learning activities of medical residents in their usual working conditions in hospitals. The main reason for residents to read medical literature is to find support to their decisions when treating patients. Based on this perspective we designed a self-applied questionnaire that explores diverse circumstances in which the vinculation between theory and practice may be expressed. This instrument was validated through rounds of experts in terms of its construction and content. Its external validity was explored with two groups of internal medicine residents with a different degree of vinculation between theory and practice: one high and the other with a low vinculation. In addition, we designed a guide for the direct observation of the theoretical and practical clinical learning activities in order to estimate its concordance with the results of the questionnaire. The questionnaire was able to discriminate the group differences and showed a satisfactory concordance with the information provided by direct observation. A copy of the questionnaire is available by request to the authors. We conclude that in its present stage of development, the instrument has shown internal and external validity and may be used to explore the process training of medical residents.

Education, Medical↗

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning↗

Representativeness and response rates from the Domestic/International Gastroenterology Surveillance Study (DIGEST).

BACKGROUND: The Domestic/international Gastroenterology Surveillance Study (DIGEST) examined the prevalence of upper gastrointestinal symptoms among the general population in 10 countries, and the impact of these symptoms on healthcare usage and quality of life. This report discusses the validation of the DIGEST sample and reviews the response rates from the survey. METHODS: External validation of the DIGEST sample was conducted by comparing the age, age by gender and annual household incomes of the sample with census-derived data. A comparison was also made between Psychological General Well-Being Index (PGWBI) scores from study subjects in the Scandinavian countries and the USA and the total sample population norms. RESULTS: Under- and oversampling, defined as > or =5% difference from the population norms, was evident in eight out of 10 countries, but no systematic bias was evident. The final distribution of the sample by gender was 51% female and 49% male. Although differences in PGWBI scores were noted between DIGEST subjects and population norms, these differences were <0.30 standard deviations--markedly below the difference considered as relevant for the PGWBI. Response for the survey in individual countries ranged from 17% in the USA to 61% in Norway, with a survey-wide rate of 27%. The overall response rate, including primary non-respondents, was 13.4%. The majority of nonresponse (51.4%) was attributed to failure to establish contact with the subjects, with 41.7% of subjects declining to be interviewed and the remaining 6.9% of subjects not meeting the age and sex criteria used for the survey. CONCLUSIONS: The DIGEST sample exhibited good external validity, providing a foundation for comparison between data derived from individual countries in the survey.

Adult↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

How generalizable are the effects of smoking prevention programs? Refusal skills training and parent messages in a teacher-administered program.

This study investigated both substantive and methodological issues associated with school-based smoking prevention programs. Substantive issues included the efficacy of a refusal skills training curriculum and of parent messages mailed to students' homes. Methodological issues included the effects of assigning classrooms versus entire schools to experimental conditions and determination of the effects of attrition on internal and external validity. Results revealed differential impact for different subgroups of adolescents. The refusal skills program produced lower rates of smoking than the control condition for students who were smokers at the pretreatment assessment but may have produced detrimental effects among males who were nonsmokers at pretest. The provision of parent messages did not affect outcome. Method of assignment (schools versus classrooms) failed to produce significant effects, and attrition did not affect internal validity. However, the above differential findings, as well as the impact of attrition on external validity, raise questions concerning the generalizability of smoking prevention programs.

Adolescent↗