PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “External validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Empirically supported psychosocial interventions for children: an overview.

Discusses issues related to the identification of psychosocial interventions for children that have demonstrated efficacy. Recent debate concerning differences between clinical trials research and clinical practice is summarized, including the tradeoff between interpretability (internal validity) and generalizability (external validity) of outcome studies. This article serves as an introduction to the special issue containing articles that have as their focus the identification of empirically supported psychosocial interventions for children as part of a task force. The article provides an overview of the history, agenda, and methodology used by the task force to define and identify specific empirically supported interventions for children with specific disorders. Whereas a number of well-established or probably efficacious interventions are identified within the series, more work directed at closing the gap between research and practice is needed.

Adolescent↗

Analysis of randomized and nonrandomized patients in clinical trials using the comprehensive cohort follow-up study design.

In clinical research, randomized trials are widely accepted as the definitive method of evaluating the efficacy of therapies. The random assignment of patients to their treatment ensures the internal validity of the comparison of new treatments with controls. An assessment of the external validity of trial results can best be achieved by comparing the study population to the population of patients who met the eligibility criteria but did not consent to randomization. A part of the data of the Coronary Artery Surgery Study (CASS), in which coronary artery bypass surgery is compared to conventional medical therapy in patients with coronary artery disease, is used to illustrate a strategy of multivariate analysis of randomized and nonrandomized patients which allows an investigation of both internal and external validity. The method used Cox's proportional hazards regression model with inclusion of covariates for randomization status and corresponding interactions in addition to the usual covariates for treatment and the important prognostic factors.

Cohort Studies↗

[Validity of the clinical prediction rule for the diagnosis of renal arterial stenosis in hypertensive patients resistant to treatment].

PURPOSE: To perform an external validation of the clinical prediction rule established by Krijnen et al. (Ann Intern Med 1998; 129: 705-11) designed to identify renal artery stenoses (RAS) in hypertensive patients. METHODS: We included 102 patients with a refractory hypertension treated with at least two antihypertensive drugs. All subjects had the research of RAS by renal angiography, or angio-computed tomography, or doppler ultrasound. Probability to detect RAS was calculated with Krijnen's algorithm (Pre-test probability) from the following parameters: age, smoking status, diffuse atherosclerosis, recent hypertension (< 2 y), obesity (BMI > 25), abdominal bruit, hypercholesterolemia (> 6.5 mmol/L), creatinine. ROC curves were plotted for each pre-test probability value. A "post-test probability" was obtained from the likelihood ratio calculated at each pre-test probability level. RESULTS: RAS prevalence in this population was 49%. Area under the ROC curve was 0.79 and Youden index was maximal for a pre-test probability of 15%. Maximal likelihood ratio was obtained for a pre-test probability of 46%. Table shows post-test probability as a function of pre-test probability obtained with Krijnen's algorithm. [table: see text] CONCLUSION: Krijnen's algorithm is valid in a population of resistant hypertensives treated with a bi-therapy. This external validation obtained on a population with a high prevalence of RAS should also be tested on a population with a lower prevalence of SAR.

Age Factors↗

[Theory and practice in medical specialization. I. An instrument for measuring learning strategies].

We present the development and validation of a measurement instrument intended to estimate the degree of vinculation between the theoretical and practical learning activities of medical residents in their usual working conditions in hospitals. The main reason for residents to read medical literature is to find support to their decisions when treating patients. Based on this perspective we designed a self-applied questionnaire that explores diverse circumstances in which the vinculation between theory and practice may be expressed. This instrument was validated through rounds of experts in terms of its construction and content. Its external validity was explored with two groups of internal medicine residents with a different degree of vinculation between theory and practice: one high and the other with a low vinculation. In addition, we designed a guide for the direct observation of the theoretical and practical clinical learning activities in order to estimate its concordance with the results of the questionnaire. The questionnaire was able to discriminate the group differences and showed a satisfactory concordance with the information provided by direct observation. A copy of the questionnaire is available by request to the authors. We conclude that in its present stage of development, the instrument has shown internal and external validity and may be used to explore the process training of medical residents.

Education, Medical↗

Data-centric, robust, and explainable multimodal deep learning for clinical decision support: A systematic review.

PURPOSE: Multimodal deep learning is increasingly proposed for clinical decision support (CDS) under a "data-centric" framing that prioritizes label quality, missing-modality robustness, distribution shift, calibration, and explainability. Prior reviews have examined multimodal medical AI, CDS, and data-centric methods separately, but none address their intersection. We mapped the modalities, fusion strategies, and data-centric and explainability techniques used in this recent literature, quantified how often each is implemented rather than merely mentioned, assessed deployment-relevant evidence (external validation, clinical-outcome measurement, equity), and formally appraised study-level risk of bias. METHODS: Following the PRISMA 2020 statement (PROSPERO CRD420261427815; registered retrospectively), we screened 150 records and included primary, clinical, multimodal studies that applied machine or deep learning to a decision-support task and reported at least one quantitative result. Two reviewers screened and extracted data with consensus adjudication. Each study was coded against pre-specified operational definitions, separating implemented or empirically evaluated techniques from those only mentioned. Study-level risk of bias was assessed with PROBAST + AI. Synthesis was narrative. RESULTS: Thirty-one studies met inclusion; 30 (97%) were published between 2024 and 2026, with a median of three modalities (range 2-6), most commonly structured EHR (71%) and imaging (39%). Data-centric techniques were frequently reported (74-84% across label-noise, distribution-shift, calibration, missing-modality and class-imbalance handling; equity 61%). However, external validation was reported in only 4/31 studies (13%), a clinical or provider outcome in 3/31 (10%), and no study reported routine deployment. Overall risk of bias was high in 27/31 studies (87%), driven by the analysis domain. CONCLUSION: Within this recent, self-selected slice of the field, technical robustness and explainability techniques are widely reported but rarely validated out-of-distribution or against clinical outcomes, and the underlying evidence is at high risk of bias. Progress requires external multi-site validation, clinical-outcome measurement, formal bias appraisal, and adherence to AI reporting standards (e.g., TRIPOD + AI) before deployment can be justified.

Deep Learning↗

Representativeness and response rates from the Domestic/International Gastroenterology Surveillance Study (DIGEST).

BACKGROUND: The Domestic/international Gastroenterology Surveillance Study (DIGEST) examined the prevalence of upper gastrointestinal symptoms among the general population in 10 countries, and the impact of these symptoms on healthcare usage and quality of life. This report discusses the validation of the DIGEST sample and reviews the response rates from the survey. METHODS: External validation of the DIGEST sample was conducted by comparing the age, age by gender and annual household incomes of the sample with census-derived data. A comparison was also made between Psychological General Well-Being Index (PGWBI) scores from study subjects in the Scandinavian countries and the USA and the total sample population norms. RESULTS: Under- and oversampling, defined as > or =5% difference from the population norms, was evident in eight out of 10 countries, but no systematic bias was evident. The final distribution of the sample by gender was 51% female and 49% male. Although differences in PGWBI scores were noted between DIGEST subjects and population norms, these differences were <0.30 standard deviations--markedly below the difference considered as relevant for the PGWBI. Response for the survey in individual countries ranged from 17% in the USA to 61% in Norway, with a survey-wide rate of 27%. The overall response rate, including primary non-respondents, was 13.4%. The majority of nonresponse (51.4%) was attributed to failure to establish contact with the subjects, with 41.7% of subjects declining to be interviewed and the remaining 6.9% of subjects not meeting the age and sex criteria used for the survey. CONCLUSIONS: The DIGEST sample exhibited good external validity, providing a foundation for comparison between data derived from individual countries in the survey.

Adult↗

Beyond predictive performance: A systematic review and critical methodological appraisal of AI/ML and conventional modelling strategies in breast, colorectal, and pancreatic Cancer.

BACKGROUND: Predictive modelling for cancer risk, treatment-related complications, and survival is central to precision oncology. Conventional logistic regression (LR) and Cox proportional hazards (CoxPH) regression remain widely used but are limited when modelling nonlinear interactions, high-dimensional imaging features, and multimodal clinical-metabolic predictors. Artificial intelligence (AI) and machine learning (ML) methods offer expanded capability through automated feature extraction, ensemble learning, and flexible survival modelling, but the evidence on when AI/ML adds value over conventional models across cancer sites and predictive tasks remains fragmented. OBJECTIVE: To systematically evaluate the methodological performance, validation strategies, and translational limitations of AI/ML models compared with conventional statistical models in published predictive-modelling studies for breast, colorectal, or pancreatic cancer. METHODS: PubMed, Scopus, and Web of Science were searched for studies published between January 2019 and March 2025. Two reviewers independently conducted title-and-abstract screening, full-text eligibility assessment, and PROBAST risk-of-bias assessment. Sixty-five studies (n&#xa0;=&#xa0;907,567 participants) were narratively synthesised by cancer site, predictive task, model family, comparator, validation strategy, predictor modality, and calibration or explainability reporting. RESULTS: The 65 studies comprised breast cancer (n&#xa0;=&#xa0;35), colorectal cancer (n&#xa0;=&#xa0;21), and pancreatic cancer (n&#xa0;=&#xa0;9). AI/ML superiority over LR and CoxPH was task- and data-dependent. CNN- and U-Net-based models predominated in imaging and body-composition tasks, tree-based ensembles consistently outperformed LR for tabular perioperative complication prediction, and CoxPH remained competitive, and in the largest pancreatic risk study, superior to XGBoost (C-index 0.802 vs 0.723) in well-structured datasets. PROBAST analysis-domain risk was moderate in 54 of 65 studies (83%), driven by limited external validation, sparse calibration reporting (11/65), and few decision-curve analyses (7/65). CONCLUSION: AI/ML adds the most methodological value in imaging-derived feature extraction and nonlinear perioperative prediction, while conventional regression remains preferable in large, structured datasets with linear predictors. Clinical translation requires standardised body-composition definitions, external validation, calibration assessment, decision-curve analysis, and explainability, in line with TRIPOD+AI and CLAIM standards.

Humans↗

How generalizable are the effects of smoking prevention programs? Refusal skills training and parent messages in a teacher-administered program.

This study investigated both substantive and methodological issues associated with school-based smoking prevention programs. Substantive issues included the efficacy of a refusal skills training curriculum and of parent messages mailed to students' homes. Methodological issues included the effects of assigning classrooms versus entire schools to experimental conditions and determination of the effects of attrition on internal and external validity. Results revealed differential impact for different subgroups of adolescents. The refusal skills program produced lower rates of smoking than the control condition for students who were smokers at the pretreatment assessment but may have produced detrimental effects among males who were nonsmokers at pretest. The provision of parent messages did not affect outcome. Method of assignment (schools versus classrooms) failed to produce significant effects, and attrition did not affect internal validity. However, the above differential findings, as well as the impact of attrition on external validity, raise questions concerning the generalizability of smoking prevention programs.

Adolescent↗

Design issues for conducting cost-effectiveness analyses alongside clinical trials.

In response to rising demands for timely economic data on new medical technologies, cost-effectiveness studies are increasingly being conducted alongside clinical trials. Because of the historical differences in perspective and methods between cost-effectiveness studies and clinical trials, the design phase of these hybrid trials requires special consideration. Cost-effectiveness studies require more comprehensive evaluations of outcomes than the endpoints typically measured in clinical trials. Often, these comprehensive outcome measures (such as quality of life) prove useful for interpreting the other endpoints measured in the trial, as well as for estimating the cost-effectiveness of the intervention. In this manuscript, we discuss several aspects related to the design of joint clinical/economic trials, including study perspective, hypothesis testing, sample size estimation, and methods for collecting cost and outcome data. We also discuss issues that may limit the external validity of the cost-effectiveness results of these trials. Many potential threats to external validity can be successfully addressed if they are identified and accounted for in the design phase of the study.

Clinical Trials as Topic↗

[Underlying cause of death from external causes: validation of official data in Recife, Pernambuco, Brazil].

OBJECTIVE: To validate the underlying cause of death recorded on the death certificates for individuals under 20 years of age who died from external causes in 1995 in Recife, Pernambuco, Brazil. METHODS: We divided the study into two stages, coding and validation. In both stages we compared the official data concerning causes of death to the data we obtained during our study. We grouped the death certificates into 5 broad categories according to the cause of death; we later subdivided them into 14 categories. We also individually compared the death certificates applying the four-digit system of the International Classification of Diseases, Ninth Revision (ICD-9). We assessed the agreement between the official data and our data in terms of sensitivity and the kappa coefficient. We took as the standard the categorization of the cause of death that we had made during our investigation. RESULTS: In the coding stage, considering all the external causes of death, the overall agreement between the official data and our study data was 94% for the 5 categories, 92% for the 14 categories, and 81% for the four-digit ICD-9 system. In the validation stage the overall agreement was 94% for the 5 categories, 91% for the 14 categories, and 73% for the four-digit ICD-9 system. CONCLUSIONS: Our results suggest that for the death certificates to be reliable, the Institute of Legal Medicine must fill them out following recommended standards. In addition, hospitals and police departments must use greater care in completing the transfer slips that accompany the bodies that are sent to the Institute. More accurate data need to be generated and disseminated for a society to better understand its patterns of violence.

Adolescent↗

[Interobserver and intraobserver variation: a problem of validity in epidemiologic studies of arterial pressure].

In carrying out blood pressure epidemiologic studies there may be different factors that can affect internal and external validity and thus eliminate the inferential process. As part of the Hypertension and Risk Factors Associated Study conducted in March 1987 in Cuajimalpa de Morelos, Mexico City, 23 nursing students were standardized on the blood pressure auscultatory method using a sound picture and measuring intraobserver and interobserver agreement through intraclass correlation coefficient. Even though initial standardization sessions showed difficulties in the use of instruments and in the reading of blood pressure levels, final K (kappa) values measuring interobserver agreement increased from 0.25 to 0.86. Omega values measuring intraobserver agreement fluctuated between 0.86 and 0.98. This epidemiologic technique is proposed in order to improve internal and external validity of blood pressure studies.

Blood Pressure↗

The WHO (Ten) Well-Being Index: validation in diabetes.

BACKGROUND: In a European trial in 8 countries, the subjective well-being of patients on alternative forms of treatment for insulin-dependent diabetes was compared using the 28-item WHO Well-Being Questionnaire, covering four dimensions of depression, anxiety, energy and positive well-being. The objective of the analysis reported here has been to identify the items of the WHO questionnaire which belong to an overall index of negative and positive well-being. METHODS: Adult patients at 10 study centres in 8 countries who had been on insulin for at least 2 years were invited to participate in a randomised, cross-over trial to compare insulin pump treatment with injection therapy. At each phase, patients completed questions on well-being and general health. Internal validity of the well-being index was evaluated by Cronbach's alpha and Loevinger's and Mokken's homogeneity coefficients, as well as factor analysis. External validity was evaluated by comparisons with results of the general assessment questions and by the ability to discriminate between the alternative forms of treatment. RESULTS: 358 patients had sufficient data for analysis. Ten items were found to constitute a valid index of well-being with respect to internal and external validity. Coefficients of homogeneity were acceptable and there was evidence for both concurrent and discriminant validity. CONCLUSIONS: The WHO (Ten) well-being index includes negative and positive aspects of well-being in a single uni-dimensional scale. Its advantage lies in its ability to show overall change along the continuum of well-being, thus facilitating comparisons between patient groups and treatments. It is not specific to diabetes, and therefore may be useful as a disease-independent index of well-being in a broad range of health care studies.

Adaptation, Psychological↗

Contemporary identification of patients at high risk of early prostate cancer recurrence after radical retropubic prostatectomy.

OBJECTIVES: To develop a model that will identify a contemporary cohort of patients at high risk of early prostate cancer recurrence (greater than 50% at 36 months) after radical retropubic prostatectomy for clinically localized disease. Data from this model will provide important information for patient selection and the design of prospective randomized trials of adjuvant therapies. METHODS: Proportional hazards regression analysis was applied to two patient cohorts to develop and cross-validate a multifactorial predictive model to identify men with the highest risk of early prostate cancer recurrence. The model and validation cohorts contained 904 and 901 men, respectively, who underwent radical retropubic prostatectomy at Johns Hopkins Hospital. This model was then externally validated using a cohort of patients from the Mayo Clinic. RESULTS: A model for weighted risk of recurrence was developed: R(W)'=lymph node involvement (0/1)x1.43+surgical margin status (0/1)x1.15+modified Gleason score (0 to 4)x0.71+seminal vesicle involvement (0/1)x0.51. Men with an R(W)' greater than 2.84 (9%) demonstrated a 50% biochemical recurrence rate (prostrate-specific antigen level greater than 0.2 ng/mL) at 3 years and thus were placed in the high-risk group. Kaplan-Meier analyses of biochemical recurrence-free survival demonstrated rapid deviation of the curves based on the R(W)'. This model was cross-validated in the second group of patients and performed with similar results. Furthermore, similar trends were apparent when the model was externally validated on patients treated at the Mayo Clinic. CONCLUSIONS: We have developed a multivariate Cox proportional hazards model that successfully stratifies patients on the basis of their risk of early prostate cancer recurrence.

Adult↗

Validation of the Hamilton Depression Rating Scale and Montgommery and Asberg Rating Scales in terms of AGECAT depression cases.

OBJECTIVE: To validate the Hamilton Depression (17) and Montgommery and Asberg Depression Scales as research instruments in older depressed community residents. DESIGN: External validation against GMS/AGECAT case level in the recruitment of older community residents for an antidepressant trial. ANALYSES: Receiver operator curves were generated for each rating scale, using GMS/AGECAT case level in external criterion. The sensitivity, specificity, positive and negative predictive values of both rating instruments were examined in the whole sample and age and gender subgroups. MADRS and HAM-D cut-off scores differentiating GMS/AGECAT cases from subcases were identified. RESULTS: HAM-D cut-off score of 16 and MADRS score of 21 were identified as differentiating case from sub-case. Diagnostic accuracy of both instruments was good, reflecting good sensitivity and specificity across both genders and sub-age groups. CONCLUSIONS: Both scales performed well in this population. These scores provide researchers with externally validated and clinically relevant cut-off scores in designing trials in the management of older depressed community residents.

Aged↗

[Clinical usefulness of oligoclonal bands].

The presence of oligoclonal bands (OCB) of immunoglobulin G (IgG) is in our days the most useful finding in the study of the CSF for the diagnosis of multiple sclerosis (MS). The most sensitive method for the detection of OCB is the isoelectric focusing followed by immunoblotting. The prevalence of OCB changes in different populations with a rank of results from 60 to 95 97%. We have determined the prevalence of OCB in our population and the sensitivity and the specificity of the technique used in our laboratory. We have included 391 patients in whom we analysed the presence of OCB, subdivided in; Group 0: Diagnosed of MS, group 1: First episode of demyelinating process, group 2: Neurological disorders considered noninflammatory or nonautoimmune (NINA),group 3: Neurological disorders considered inflammatory, infectious or autoimmune (IIA). The presence of OCB was searched in CSF and serum simultaneously using isoelectric focusing and immunoblotting. In order to standardize the technique we achieved and internal and external validation. Internal validation: sensitivity and specificity (using as a control group first the group NINA and after the group IA). External validation: we choose 10 pairs of CSF/serum from patients with different diagnostics and sent to a reference laboratory ( Karolinska Institute Medical School) that was blind of our results and of the diagnostics. The prevalence of OCB in each group has been: group 0 (MS): 87.7%, group 1: 54.8%, group 2 (NINA): 17.5%, group 3(IIA): 52.7%. Sensitivity: 97.7%, specificity using group NINA as control 82.5% and using group IIA 45.7%. Concordance with the reference laboratory in 9/10 determinations. We conclude that in our population the prevalence of OCB, in patients with MS, is lower than in Northern Europe. The OCB appear in may inflammatory, autoimmune diseases, their specificity for the diagnostic of MS is low.

Autoimmune Diseases↗

Artificial Intelligence for Diagnosis, Risk Stratification, and Prognosis of Neuroblastoma - A Systematic Review and Meta-Analysis.

PURPOSE: To synthesizes evidence on artificial intelligence (AI) performance in neuroblastoma (NB) diagnosis, risk stratification, prognosis, and genomic characterization. MATERIALS AND METHODS: A systematic review and meta-analysis was conducted following PRISMA 2020 guidelines (PROSPERO: CRD42024539475) across five databases. Meta-analyses used random-effects models with logit-transformed Area Under the Curve (AUCs) and cluster-robust standard errors. AI models were classified as Machine Learning Models (MLM) or Hybrid Nomograms (HN) based on their construction methodology. RESULTS: Of 3,742 articles identified, 53 were included. MLMs demonstrated higher point estimates than radiologists in differential diagnosis (AUC: 0.87 vs. 0.83), though this difference was not statistically significant and carried substantial uncertainty. HNs achieved stronger performance in risk stratification (AUC: 0.87). AI-derived nomograms (AUC: 0.9) and gene signatures (AUC: 0.8) outperformed conventional prognostic markers descriptively. Chemotherapy response prediction remained below clinical utility thresholds across all model types. Only 33.9% of models reported calibration and 24.5% underwent external validation. CONCLUSIONS: AI demonstrates proof-of-concept across multiple NB clinical domains. However, clinical adoption remains premature given persistent gaps in external validation, calibration, dataset size, and pediatric-specific model development. Future studies should test these models prospectively in multicenter pediatric cohorts, ideally through COG or SIOPEN, using shared definitions for diagnosis, risk group, treatment response, and survival outcomes.

Humans↗

Predictive validity of the Strain Index in turkey processing.

The Strain Index is a job analysis method for determining if workers are exposed to increased risk of developing distal upper extremity disorders. Its predictive and external validity was initially demonstrated in a pork processing plant. The purpose of this study was to evaluate the predictive validity of the Strain Index in one turkey processing plant. While blinded to health outcomes, investigators analyzed the right and left sides of workers in 28 jobs using the Strain Index and classified them as "hazardous" or "safe" based on the Strain Index score. Subsequently, OSHA 200 logs were used to ascertain the occurrence of distal upper extremity disorders retrospectively. If at least one such disorder had occurred on the right or left side during the previous 3 years, that side was classified as "positive." If no such disorder was reported during the previous 3 years, that side was classified as "negative." When comparing sides, symmetry between morbidity and hazard classification was required. When comparing jobs, such symmetry was not required. Evidence of association between the hazard classifications and the morbidity classifications for the 56 sides and the 28 jobs was evaluated using 2 x 2 contingency tables. For the sides, the association between hazard classification and morbidity classification was statistically significant, with an odds ratio of 22.0. The sensitivity, specificity, positive predictive value, and negative predictive value were 0.86, 0.79, 0.92, and 0.65, respectively. Similar results were noted for the jobs--the odds ratio was 50.0, and the sensitivity, specificity, positive predictive value, and negative predictive value were 0.91, 0.83, 0.95, and 0.71. These results provide additional evidence of the external validity and predictive validity of the Strain Index.

Animals↗

Development and validation of the Headache Needs Assessment (HANA) survey.

OBJECTIVE: To develop and validate a brief survey of migraine-related quality-of-life issues. The Headache Needs Assessment (HANA) questionnaire was designed to assess two dimensions of the chronic impact of migraine (frequency and bothersomeness). METHODS: Seven issues related to living with migraine were posed as ratings of frequency and bothersomeness. Validation studies were performed in a Web-based survey, a clinical trial responsiveness population, and a retest reliability population. Headache characteristics (eg, frequency, severity, and treatment), demographic information, and the Headache Disability Inventory were used for external validation. RESULTS: The HANA was completed in full by 994 adults in the Web survey, with a mean total score of 77.98 +/- 40.49 (range, 7 to 175). There were no floor or ceiling effects. The HANA met the standards for validity with internal consistency reliability (Cronbach alpha =.92, eigenvalue for the single factor = 4.8, and test-retest reliability = 0.77). External validity showed a high correlation between HANA and Headache Disability Inventory total scores (0.73, P<.0001), and high correlations with disease and treatment characteristics. CONCLUSIONS: These data demonstrate the psychometric properties of the HANA. The brief questionnaire may be a useful screening tool to evaluate the impact of migraine on individuals. The two-dimensional approach to patient-reported quality of life allows individuals to weight the impact of both frequency and bothersomeness of chronic migraines on multiple aspects of daily life.

Activities of Daily Living↗