PubMed HealthSearch

SEARCH · PubMed Health

Results for “External validation”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Integrating genetic predictors into subsequent breast cancer risk prediction in survivors of childhood cancer.

PURPOSE: Female survivors of childhood cancer are at high risk for developing breast cancer. The contributions of most general population primary breast cancer genetic predictors to this risk have not been explored. METHODS: Analyses included females who survived &#x2265;5 years after their childhood cancer diagnosis with available array (N&#x2009;=&#x2009;2096, subsequent breast cancer [SBC]=218) or whole-genome sequencing (WGS; N&#x2009;=&#x2009;3292, SBC=101) data from the Childhood Cancer Survivor Study and St. Jude Lifetime Cohort. We computed 99 externally-validated primary breast cancer polygenic risk scores (PRS). Using deep-coverage WGS, ClinVar-annotated pathogenic/likely pathogenic (P/LP) variants in breast cancer susceptibility genes were identified. Cox proportional hazards models assessed associations with SBC risk, adjusting for treatments and genetic ancestry. RESULTS: Among 5388 female survivors (genetic ancestry, European: N&#x2009;=&#x2009;4,752; African: N&#x2009;=&#x2009;444; East Asian: N&#x2009;=&#x2009;192), 319 developed SBC. Most (90.9%) PRSs were nominally associated with SBC risk (P&#x2009;<&#x2009;0.05), but effect sizes varied substantially. PRSs with superior discriminatory ability had greater genome-wide coverage (e.g., 6.4 million-variant PRS, HR per SD&#x2009;=&#x2009;1.71, 95% CI&#x2009;=&#x2009;1.43 to 2.05; P&#x2009;=&#x2009;4.2x10-9) and 7.7-fold higher odds (P&#x2009;=&#x2009;7.0x10-4) of including variants in multiple DNA damage repair pathways compared with PRSs with weaker risk associations. Among survivors with WGS, 1.6% carried P/LP variants in clinical testing panel genes, which was associated with a 7.4-fold greater risk (95% CI&#x2009;=&#x2009;3.16 to 17.19). Including genetic factors improved SBC risk prediction by age 40 (P&#x2009;<&#x2009;0.001) compared to treatment exposures alone. CONCLUSIONS: Externally-validated primary breast cancer genetic susceptibility predictors are relevant for SBC risk prediction and should be prioritized for risk stratification in survivors.

Journal Article

Recruitment issues, health habits, and the decision to participate in a health promotion program.

To understand the external validity of experimental studies, it is important to estimate the extent to which the participants are representative of the general population. This paper describes recruitment methods and considers the representativeness of participants in the San Diego Family Health Project. The study was designed to experimentally evaluate the effectiveness of a family-based behavior change intervention in Anglo and Mexican-American families. Initial contact with the families was made through a household health survey that was sent home with all fifth- and sixth-grade children in 12 participating elementary schools. The survey asked about a variety of demographic characteristics, dietary habits, and physical activity habits. Parents were also asked if they were interested in participating in the project. Respondents were classified by level of participation into one of three groups: not interested, expressed initial interest but did not attend the recruitment meeting, and volunteered to participate. Level of participation was the independent variable in the analyses. In separate analyses for Anglo and Mexican-American responders, our data suggested many similarities and a few differences among participant groups. The differences that were observed suggest that participants may already have healthier diets than nonparticipants, although only one of four dietary variables differed by participation status in each ethnic group. The external validity of these data and general recruitment issues are discussed.

Adolescent

An external construct validity study of Rorschach personality variables.

This study examined (a) hypothesized relationships between Rorschach variables and self-report test measures relating to nominally similar aspects of personality functioning and (b) interrelationships among Rorschach variables. Sixty-two undergraduates were administered the Rorschach, Barron Ego Strength Scale, Kaplan Self-Derogation Scale, Eagly Self-Esteem Scale, Multiple Affective Adjective Checklist (MAACL), Marlowe-Crowne Social Desirability Scale, and the Rotter Locus of Control Scale. Only a few of the predictions received confirmation: inanimate movement (m) correlated, as expected, with MAACL anxiety and hostility, the egocentricity index (3r + 2)/R (R = total responses) correlated significantly with self-esteem, and human movement with minus form level (M-) correlated (inversely) with ego strength. Among the unpredicted findings were some that appear inconsistent with standard Rorschach interpretation. Rorschach variables human movement (M), and experience actual (EA), generally interpreted as reflecting coping resources, related significantly with self-report measures of poor coping and of dysphoric affect. In general, the Rorschach appears better at identifying weaknesses in the ego rather than strengths.

Adolescent

Multimodal features and prognostic risk assessment in locally advanced gastric cancer patients following neoadjuvant therapy based on machine learning algorithms: a multicenter study.

BACKGROUND: Neoadjuvant therapy (NAT) is recommended for locally advanced gastric cancer (LAGC), but some patients respond poorly. We aimed to construct a multimodal model integrating CT images, transcriptomic sequencing, and clinicopathological data to assess prognosis in LAGC patients receiving NAT. MATERIALS AND METHODS: This multicenter study included 505 LAGC patients who underwent NAT. Radiomic features were extracted from preoperative CT images of 505 patients. RNA-seq was performed on 277 post-NAT specimens, with additional data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases (n&#x2009;=&#x2009;804). Patients were divided into training (168 cases), internal validation (72 cases), and external validation cohorts. Machine learning algorithms identified key radiomic, molecular, and clinical features associated with NAT response, which were then integrated into a multimodal model to predict overall survival (OS) and disease-free survival (DFS). RESULTS: Six radiomic and three molecular features significantly associated with NAT response were selected. Radiomic risk (hazard ratio [HR]: 4.0, P&#x2009;<&#x2009;0.001) and molecular risk (HR: 7.1, P&#x2009;<&#x2009;0.001) were independent prognostic factors. By integrating radiomic risk, molecular risk, and clinical characteristics, a multimodal model (MuMo) was constructed.The C-index results (OS, C-index&#x2009;=&#x2009;0.855; DFS, C-index&#x2009;=&#x2009;0.786) demonstrated that MuMo outperformed the single-modality models and ypTNM staging.Mechanistic analysis suggested that the efficacy of neoadjuvant therapy was significantly enriched in immune-inflammatory pathways. CONCLUSIONS: MuMo can effectively predict postoperative survival risk in LAGC patients receiving NAT, serving as a powerful tool for optimizing prognostic assessment.

Humans

Noninvasive detection and differentiation of gastric malignancy using cell-free DNA biomarkers.

INTRODUCTION: Gastric cancer remains a major global health burden, with high mortality driven by late-stage diagnoses that limit treatment options and reduce survival. Current diagnostic methods such as endoscopy and biopsy are invasive, resource-intensive, and impractical for large-scale early detection. OBJECTIVES: This study aimed to develop and validate an ensemble machine learning model integrating four cell-free DNA (cfDNA) fragmentomic feature classes derived from 5&#xa0;&#xd7;&#xa0;whole genome sequencing (WGS) data to non-invasively differentiate malignant gastric cancer from benign gastric lesions in high-risk or symptomatic patients. METHODS: A total of 681 plasma samples were prospectively collected, comprising 329 from patients with gastric cancer or high-grade intraepithelial neoplasia (HGIN) and 352 from individuals with benign gastric conditions. The dataset was divided into a training cohort (n&#xa0;=&#xa0;333) and a temporally independent validation cohort (n&#xa0;=&#xa0;348). An external validation cohort of 305 participants was also included. RESULTS: The ensemble model achieved an AUROC of 0.920 in cross-validation testing on the training cohort, 0.912 in the independent validation cohort, and 0.896 (95% CI 0.860-0.932) in the external cohort. At a pre-specified prediction threshold of 0.402, the model demonstrated 93.3% sensitivity and 71.9% specificity in the validation cohort, yielding a PPV of 71.3% and an NPV of 93.5%. In the external cohort, sensitivity and specificity were 91.7% and 69.1%, respectively (PPV 75.7%, NPV 88.8%). Model scores correlated with clinical stage, tumor grade, and histopathological subtype. Approximately 71% of non-cancer patients could have been spared unnecessary endoscopy. CONCLUSIONS: The cfDNA fragmentomics-based ensemble model enables accurate, non-invasive differentiation between gastric cancer and benign gastric lesions in high-risk or symptomatic patients. This approach demonstrates strong potential as a pre-endoscopy triage tool, supporting earlier detection and more efficient use of diagnostic resources.

Humans

A principal-components analysis of the Narcissistic Personality Inventory and further evidence of its construct validity.

We examined the internal and external validity of the Narcissistic Personality Inventory (NPI). Study 1 explored the internal structure of the NPI responses of 1,018 subjects. Using principal-components analysis, we analyzed the tetrachoric correlations among the NPI item responses and found evidence for a general construct of narcissism as well as seven first-order components, identified as Authority, Exhibitionism, Superiority, Vanity, Exploitativeness, Entitlement, and Self-Sufficiency. Study 2 explored the NPI's construct validity with respect to a variety of indexes derived from observational and self-report data in a sample of 57 subjects. Study 3 investigated the NPI's construct validity with respect to 128 subject's self and ideal self-descriptions, and their congruency, on the Leary Interpersonal Check List. The results from Studies 2 and 3 tend to support the construct validity of the full-scale NPI and its component scales.

Adolescent

Risk Prognostication After Hypomethylating Agents Combined With Venetoclax in AML: The PRISM Risk Model.

PURPOSE: As risk stratification for patients with AML treated with lower-intensity venetoclax-based therapy remains suboptimal, we developed and validated a prognostic model integrating clinical, cytogenetic, and molecular features. METHODS: We assembled a multinational data set comprising 2,092 adults with newly diagnosed AML treated with hypomethylating agents plus venetoclax (HMA + VEN). One thousand nine hundred eighteen patients with complete data were randomly divided into training (70%) and internal validation (30%) cohorts. Two independent external validation cohorts were assembled (n = 500 and n = 222). Modeling overall survival (OS), Elastic Net regression was applied in 1,000 bootstrap samples from the training cohort to select variables for a Ridge regression, which generated a continuous Prognostic Risk Integration for Survival Modeling (PRISM) score and risk categories based on tertiles (PRISM-3: low, moderate, high). These PRISM indices were then computed for the validation cohorts and compared with the 4-gene classifier (based on mutations in FLT3-ITD, N/KRAS, and TP53). RESULTS: PRISM integrated 17 clinical and genomic variables and demonstrated a linear association with OS. PRISM-3 stratified survival consistently across all cohorts (median OS: 25.1-28.8 months for low risk, 12.5-14.7 months for moderate risk, and 5.8-6.7 months for high risk; P < .001). Compared with the 4-gene classifier, PRISM-3 reassigned approximately 40% of patients (and >50% of those with favorable risk) and demonstrated significantly better discrimination in validation cohorts (C-index 0.63-0.65 v 0.59-0.61; P < .05). CONCLUSION: PRISM is a validated prognostic model for patients with AML receiving HMA + VEN that improves survival risk stratification beyond current standard tools and supports individualized, risk-adapted clinical decision making. The model, the PRISM-AML Risk Calculator, is publicly available.

Humans

Evaluation and measurement: some dilemmas for health education.

Seven dilemmas of evaluation and measurement posed by the nature of health education are presented, together with suggestions for their resolution. These include the dilemmas of : 1) rigor of experimental design vs significance or program adaptability; 2) internal validity or "true" effectiveness vs external validity or feasibility; 3) experimental vs placebo effectsl 4) effectiveness vs economy of scale; 5) risk vs payoff; 6) measurement of long-term vs short-term out-comon. Emphasis is placed on the need to develop a more cumulative data base through standardization of measures, replication of experiments in different settings, and better documentation, reporting, and diffusion of experiences in practice.

Cost-Benefit Analysis

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n&#x202f;=&#x202f;38, 74%). Hierarchical clustering (n&#x202f;=&#x202f;20) and K-means clustering (n&#x202f;=&#x202f;14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Integrated Genomic and Tumor Microenvironment Subtyping Improved Risk Stratification in Primary Central Nervous System Lymphoma.

Current prognostic models fail to capture the biological complexity of primary central nervous system lymphoma (PCNSL). We integrated whole-genome sequencing and multiplex immunofluorescence in 68 treatment-na&#xef;ve patients to define four genomic subtypes (C1, C2, C3, and C4) with divergent survival (C4 worst: median overall survival [OS], 26&#x2009;months). In parallel, a novel tumor microenvironment (TME) classification based on CD8+T/M2 macrophage ratio stratified patients into High (>&#x2009;1.5), Intermediate (0.8-1.5), and Low (<&#x2009;0.8) groups. Unexpectedly, the Intermediate TME group showed the poorest outcomes (5-year OS: 10%). Integration revealed a lethal subgroup (C4&#x2009;+&#x2009;Intermediate TME; 9.8% of cohort) with a median OS of 3.0&#x2009;months (hazard ratio&#x2009;=&#x2009;7.24, p&#x2009;=&#x2009;0.006). Prognostic nomograms incorporating these subtypes showed promising discriminative performance in internal validation (C-index >&#x2009;0.78), but external validation is needed. Together, these findings identify a high-risk biological subset and provide a hypothesis-generating framework for future biomarker-driven risk stratification and therapeutic discovery in PCNSL.

Humans

Age-based construct validation using structural equation modeling.

In this paper we describe some mathematical and statistical models based on structural equation modeling (SEM) using computer programs like LISREL. We focus on SEM methodology for the simultaneous examination of the internal validity of psychological constructs and the external validity represented by age relations. To illustrate these ideas we use a latent variable path model to examine the organization of intellectual abilities measured by the WAIS-R in the standardization sample. We also examine different ways in which age can be used to structure this organization. This is primarily a methodological paper, but we try to integrate conceptual principles of modeling with some substantive issues of research on the psychology of aging.

Aging

Confidence intervals versus p-values for interpretation of clinical trial results: introduction.

The following three papers summarize the presentations at a Society for Clinical Trials annual meeting session on the relative merits of estimation versus testing for analysis of randomized clinical trials. By design, randomized clinical trials have internal validity. Whether they also possess quantitative external validity--generalizability of effect size to some population represented by the trial subjects--is one of the main points of disagreement among the three authors. It may be unrealistic to expect a resolution that applies across the wide variety of therapeutic areas and clinical trial goals. Extrapolation from clinical trial to clinical practice is often endorsed in connection with large trials having loose entry criteria and focusing on an objective, clearly meaningful clinical endpoint. By contrast, the relevance of estimates of effect size is less clear in the case of many clinical trials conducted in the course of drug development.

Clinical Trials as Topic

SERPINE1-centric inflammatory signature associates with treatment resistance and survival in laryngeal squamous cell carcinoma.

BACKGROUND: Laryngeal squamous cell carcinoma (LSCC) prognosis remains poor despite treatment advances. More accurate prognostic assessment models can help guide individualized treatment and improve prognosis. Chronic inflammation contributes to tumorigenesis, yet inflammatory response-related genes (IRGs) in LSCC prognosis are underexplored. This study aimed to construct an IRG prognostic signature for LSCC and further dissect core IRG-mediated mechanisms of immune escape and chemoresistance. METHODS: Transcriptional profiles and clinical data from LSCC patients were retrieved from The Cancer Genome Atlas (TCGA). IRGs were sourced from Gene Set Enrichment Analysis (GSEA) hallmark gene set. We identified differentially expressed IRGs linked to survival outcomes in LSCC. Key IRGs were subsequently selected using least absolute shrinkage and selection operator (LASSO) Cox regression analysis to establish an inflammatory risk score model. This model underwent internal validation within the TCGA cohort and external validation using independent Gene Expression Omnibus (GEO) datasets. We further assessed the model's association with the tumor immune microenvironment and the impact of IRGs on chemotherapy response. Finally, the functional roles of interested signature IRG were experimentally validated in LSCC cell lines. RESULTS: Four significant IRGs (AQP9, ITGA5, LCK, SERPINE1) were identified to build the risk score model. The model stratified LSCC patients into distinct prognostic groups: TCGA cohort: 5-year area under the curve (AUC) =0.836, P<0.001; GSE25727 cohort: 5-year AUC =0.706, P=0.02; GSE27020 cohort: 5-year AUC =0.798, P<0.01. Multivariate analysis confirmed the risk score as an independent prognostic factor (P<0.05). High-risk patients showed reduced immune cell infiltration (CD8+ T cells, dendritic cells) and suppressed immune pathways. Multi-algorithm immune analysis further revealed defective antigen presentation and reduced anti-tumor immune infiltration in high-risk LSCC, promoting tumor immune escape. GSEA/Gene Ontology (GO) enrichment combined with drug sensitivity prediction further revealed that high-risk tumors activate invasive signaling and acquire broad chemoresistance alongside impaired anti-tumor immunity. SERPINE1 might be associated with chemotherapy resistance and exhibited the highest alteration frequency (predominantly amplification) and overexpression in LSCC tissues. Its knockdown significantly suppressed proliferation, migration, invasion and chemoresistance in LSCC cells. Immunohistochemistry (IHC) confirmed tumor SERPINE1 overexpression (P=0.002 vs. normal tissues), correlating with poor survival (P<0.001). CONCLUSIONS: The 4-IRG risk signature is a reliable prognostic indicator reflecting immune dysfunction in LSCC. SERPINE1 is validated as a therapeutic target and biomarker, enriching our understanding of gene regulation dynamics in LSCC.

Laryngeal cancer

Clinical trials in the elderly. Pivotal points.

A clinical trial, that is, the scientific assessment of drug action in humans, must be undertaken only if there is reference to an expectation of benefit from the trial. Before the trial starts, a series of questions should be answered, including the need for the trial in elderly patients, particularly in view of the possible vulnerability of the elderly study subjects. Patient selection, randomization, follow-up, analysis, and interpretation must be scientifically valid. Of major importance is the external validity of the trial, that is, the generalizability. Any trial, particularly those involving the elderly, should be designed to develop or contribute to generalizable knowledge; that is, it should be possible to extend the conclusions from the trial beyond the study population to the population at large. When elderly are involved, as they should be when a drug is proposed for use mainly in the elderly population, it is especially important that the study be scientifically valid, medically important, and ethically sound. Studies involving the elderly should have sufficient numbers of females and minorities, the former because most elderly patients are females, the latter because minorities now reach the age of 65 years and beyond more so than in the past. Both risk and benefit should be addressed in terms of potential magnitude and duration. If the study drug has a narrow therapeutic window, it should be intensively studied. When the study is completed, there ought to be clear guidelines for the clinician to design initial, individualized, optimal dosage regimens or for subsequent adjustment of the regimen. The final report, which should be easily evaluable by clinicians, should fully discuss reasons for dropouts, inappropriate patient inclusion, number and types of adverse reactions, and defects in design and conduct of the study.

Aged

Temporal lobe signs: electroencephalographic validity and enhanced scores in special populations.

Internal and external validity tests were completed for an inventory that has been used to infer signs of temporal lobe lability. Strong, positive correlations were reported for a normal (reference) population between the numbers of responses that referred to paranormal experiences (including feelings of a "presence") and separately to religious beliefs and the numbers of spikes per minute within electroencephalographic recordings from the temporal lobe. Numbers of spikes were also correlated with the subjects' scores on the hysteria, schizophrenia, and psychasthenia scales from the MMPI. These clusters of items were not correlated with electrical activity from the occipital lobe (the comparison region). Numbers of responses to control clusters of mundane experiences were not correlated with the temporal lobe measures. A group of student poets scored higher on different subclusters of temporal lobe signs and on the schizophrenia and mania scales of the MMPI than the reference group. For both groups, there were positive correlations between the amount of alpha activity in the temporal lobe only and answers to items such as "hearing inner voices" and "feeling as if things were not real." These results demonstrate that quantitative measures of electrical changes in the temporal lobe are correlated with (or with the report of) specific experiences that are prevalent during surgical or epileptic stimulation of this brain region.

Adult

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (&#x2265;54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans

Ecological validity and cultural sensitivity for outcome research: issues for the cultural adaptation and development of psychosocial treatments with Hispanics.

This article has two objectives. The first is to provide a culturally sensitive perspective to treatment outcome research as a resource to augment the ecological validity of treatment research. The relationships between external validity, ecological validity, and culturally sensitive research are reviewed. The second objective is to present a preliminary framework for culturally sensitive interventions that strengthen ecological validity for treatment outcome research. The framework, consisting of eight dimensions of treatment interventions (language, persons, metaphors, content, concepts, goals, methods, and context) can serve as a guide for developing culturally sensitive treatments and adapting existing psychosocial treatments to specific ethnic minority groups. Examples of culturally sensitive elements for each dimension of the intervention are offered. Although the focus of the article is on Hispanic populations, the framework may be valuable to other ethnic and minority groups.

Cross-Cultural Comparison