PubMed HealthSearch

SEARCH · PubMed Health

Results for “External validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

[Registries of morbimortality in cardiology: methods].

The effectiveness of diagnostic, preventive and therapeutic procedures, whose efficacy has been assessed in clinical trials, should be tested in a real treatment scenario. The procedures used in acute myocardial infarction (AMI) management can be evaluated by means of cohort studies that include all consecutive patients admitted to one or several hospitals. Such studies are called hospital registries. They are simpler to organize and cheaper than clinical trials. On the other hand, the AMI population-based registries allow the establishment of the incidence and mortality rates, as well as case-fatality as they include those patients who die before reaching hospital facilities. In both types of registries a set of variables on co-morbidity, age, sex, severity, and the utilization of procedures along with the course of the disease are systematically recorded in each patient using standard definitions to warrant the internal validity. In hospital registries, the external validity of the results will depend on whether the sample of hospitals represents the population where it was obtained. A good registry should include patients with a wide age range, allow the analysis of specific subgroups of patients such as non-Q wave or first AMI to allow for comparison with other registries. In addition, it should also permit a mid-term follow-up, respect ethical issues, receive appropriate funding and keep a multidisciplinary team involved in its design and development.

Cardiology

Risk Prognostication After Hypomethylating Agents Combined With Venetoclax in AML: The PRISM Risk Model.

PURPOSE: As risk stratification for patients with AML treated with lower-intensity venetoclax-based therapy remains suboptimal, we developed and validated a prognostic model integrating clinical, cytogenetic, and molecular features. METHODS: We assembled a multinational data set comprising 2,092 adults with newly diagnosed AML treated with hypomethylating agents plus venetoclax (HMA + VEN). One thousand nine hundred eighteen patients with complete data were randomly divided into training (70%) and internal validation (30%) cohorts. Two independent external validation cohorts were assembled (n = 500 and n = 222). Modeling overall survival (OS), Elastic Net regression was applied in 1,000 bootstrap samples from the training cohort to select variables for a Ridge regression, which generated a continuous Prognostic Risk Integration for Survival Modeling (PRISM) score and risk categories based on tertiles (PRISM-3: low, moderate, high). These PRISM indices were then computed for the validation cohorts and compared with the 4-gene classifier (based on mutations in FLT3-ITD, N/KRAS, and TP53). RESULTS: PRISM integrated 17 clinical and genomic variables and demonstrated a linear association with OS. PRISM-3 stratified survival consistently across all cohorts (median OS: 25.1-28.8 months for low risk, 12.5-14.7 months for moderate risk, and 5.8-6.7 months for high risk; P < .001). Compared with the 4-gene classifier, PRISM-3 reassigned approximately 40% of patients (and >50% of those with favorable risk) and demonstrated significantly better discrimination in validation cohorts (C-index 0.63-0.65 v 0.59-0.61; P < .05). CONCLUSION: PRISM is a validated prognostic model for patients with AML receiving HMA + VEN that improves survival risk stratification beyond current standard tools and supports individualized, risk-adapted clinical decision making. The model, the PRISM-AML Risk Calculator, is publicly available.

Humans

Evaluation and measurement: some dilemmas for health education.

Seven dilemmas of evaluation and measurement posed by the nature of health education are presented, together with suggestions for their resolution. These include the dilemmas of : 1) rigor of experimental design vs significance or program adaptability; 2) internal validity or "true" effectiveness vs external validity or feasibility; 3) experimental vs placebo effectsl 4) effectiveness vs economy of scale; 5) risk vs payoff; 6) measurement of long-term vs short-term out-comon. Emphasis is placed on the need to develop a more cumulative data base through standardization of measures, replication of experiments in different settings, and better documentation, reporting, and diffusion of experiences in practice.

Cost-Benefit Analysis

[Is the randomized controlled trial overvalued as a basis for clinical decision-making? A review with comments].

The randomized controlled trial (RCT) may have considerable limitations in clinical research. Lacking the possibility of blinding impairs the internal validity of the trials. The external validity is often impaired, as results of RCTs obtained in an ideal situation, may be difficult to generalize to a clinical routine situation. Pragmatic randomized trials move from ideal situations towards routine situations, and by modifying the design it is possible to reduce selection bias due to patient and physician preferences. Quasi-experimental studies have varying degrees of problems with internal validity but are necessary contributions to our knowledge of the effect of treatment in clinical routine situations. Limitations of the usefulness of RCTs as well as pragmatic and quasi-experimental studies in clinical research make it necessary to recognise that different methods complement one another. Research in development of RCTs and new methods in clinical research should be encouraged.

Decision Making

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n&#x202f;=&#x202f;38, 74%). Hierarchical clustering (n&#x202f;=&#x202f;20) and K-means clustering (n&#x202f;=&#x202f;14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Integrated Genomic and Tumor Microenvironment Subtyping Improved Risk Stratification in Primary Central Nervous System Lymphoma.

Current prognostic models fail to capture the biological complexity of primary central nervous system lymphoma (PCNSL). We integrated whole-genome sequencing and multiplex immunofluorescence in 68 treatment-na&#xef;ve patients to define four genomic subtypes (C1, C2, C3, and C4) with divergent survival (C4 worst: median overall survival [OS], 26&#x2009;months). In parallel, a novel tumor microenvironment (TME) classification based on CD8+T/M2 macrophage ratio stratified patients into High (>&#x2009;1.5), Intermediate (0.8-1.5), and Low (<&#x2009;0.8) groups. Unexpectedly, the Intermediate TME group showed the poorest outcomes (5-year OS: 10%). Integration revealed a lethal subgroup (C4&#x2009;+&#x2009;Intermediate TME; 9.8% of cohort) with a median OS of 3.0&#x2009;months (hazard ratio&#x2009;=&#x2009;7.24, p&#x2009;=&#x2009;0.006). Prognostic nomograms incorporating these subtypes showed promising discriminative performance in internal validation (C-index >&#x2009;0.78), but external validation is needed. Together, these findings identify a high-risk biological subset and provide a hypothesis-generating framework for future biomarker-driven risk stratification and therapeutic discovery in PCNSL.

Humans

Age-based construct validation using structural equation modeling.

In this paper we describe some mathematical and statistical models based on structural equation modeling (SEM) using computer programs like LISREL. We focus on SEM methodology for the simultaneous examination of the internal validity of psychological constructs and the external validity represented by age relations. To illustrate these ideas we use a latent variable path model to examine the organization of intellectual abilities measured by the WAIS-R in the standardization sample. We also examine different ways in which age can be used to structure this organization. This is primarily a methodological paper, but we try to integrate conceptual principles of modeling with some substantive issues of research on the psychology of aging.

Aging

Confidence intervals versus p-values for interpretation of clinical trial results: introduction.

The following three papers summarize the presentations at a Society for Clinical Trials annual meeting session on the relative merits of estimation versus testing for analysis of randomized clinical trials. By design, randomized clinical trials have internal validity. Whether they also possess quantitative external validity--generalizability of effect size to some population represented by the trial subjects--is one of the main points of disagreement among the three authors. It may be unrealistic to expect a resolution that applies across the wide variety of therapeutic areas and clinical trial goals. Extrapolation from clinical trial to clinical practice is often endorsed in connection with large trials having loose entry criteria and focusing on an objective, clearly meaningful clinical endpoint. By contrast, the relevance of estimates of effect size is less clear in the case of many clinical trials conducted in the course of drug development.

Clinical Trials as Topic

A stroke-adapted 30-item version of the Sickness Impact Profile to assess quality of life (SA-SIP30).

BACKGROUND AND PURPOSE: In view of the growing therapeutic options in stroke, measurement of quality of life has become increasingly relevant as an outcome parameters. The Sickness Impact Profile (SIP) is one of the most widely used measures to assess quality of life. To overcome the major disadvantage of the SIP, its length, we constructed a short stroke adapted 30-item SIP version (SA-SIP30). METHODS: Data on the original SIP version were collected for 319 communicative patients at 6 months after stroke. The 12 subscales and the 136 items of the original SIP were reduced to 8 subscales with 30 items in a three step procedure, on the basis of relevancy and homogeneity. Reliability of the SA-SIP30 was evaluated by means of an analysis of homogeneity (Cronbach's alpha coefficient). Different types of validity were assessed: construct, clinical, and external validities. RESULTS: Homogeneity of the SA-SIP30 was demonstrated by a high Cronbach's alpha (0.85). Principal component analyses revealed the same two dimensions as in the original SIP (a physical and a psychosocial dimension). The SA-SIP30 could explain 91% of the variation in scores of the original SIP in the same cohort of patients, and 89% in a different cohort. Furthermore, the SA-SIP30 was related to other functional health measures similar to how the original SIP was. We could demonstrate that the SA-SIP30 was able to distinguish patients with lacunar infarctions from patients with cortical or subcortical lesions. CONCLUSIONS: We conclude that the SA-SIP30 is a feasible and clinimetrically sound measure to assess quality of life after stroke.

Aged

SERPINE1-centric inflammatory signature associates with treatment resistance and survival in laryngeal squamous cell carcinoma.

BACKGROUND: Laryngeal squamous cell carcinoma (LSCC) prognosis remains poor despite treatment advances. More accurate prognostic assessment models can help guide individualized treatment and improve prognosis. Chronic inflammation contributes to tumorigenesis, yet inflammatory response-related genes (IRGs) in LSCC prognosis are underexplored. This study aimed to construct an IRG prognostic signature for LSCC and further dissect core IRG-mediated mechanisms of immune escape and chemoresistance. METHODS: Transcriptional profiles and clinical data from LSCC patients were retrieved from The Cancer Genome Atlas (TCGA). IRGs were sourced from Gene Set Enrichment Analysis (GSEA) hallmark gene set. We identified differentially expressed IRGs linked to survival outcomes in LSCC. Key IRGs were subsequently selected using least absolute shrinkage and selection operator (LASSO) Cox regression analysis to establish an inflammatory risk score model. This model underwent internal validation within the TCGA cohort and external validation using independent Gene Expression Omnibus (GEO) datasets. We further assessed the model's association with the tumor immune microenvironment and the impact of IRGs on chemotherapy response. Finally, the functional roles of interested signature IRG were experimentally validated in LSCC cell lines. RESULTS: Four significant IRGs (AQP9, ITGA5, LCK, SERPINE1) were identified to build the risk score model. The model stratified LSCC patients into distinct prognostic groups: TCGA cohort: 5-year area under the curve (AUC) =0.836, P<0.001; GSE25727 cohort: 5-year AUC =0.706, P=0.02; GSE27020 cohort: 5-year AUC =0.798, P<0.01. Multivariate analysis confirmed the risk score as an independent prognostic factor (P<0.05). High-risk patients showed reduced immune cell infiltration (CD8+ T cells, dendritic cells) and suppressed immune pathways. Multi-algorithm immune analysis further revealed defective antigen presentation and reduced anti-tumor immune infiltration in high-risk LSCC, promoting tumor immune escape. GSEA/Gene Ontology (GO) enrichment combined with drug sensitivity prediction further revealed that high-risk tumors activate invasive signaling and acquire broad chemoresistance alongside impaired anti-tumor immunity. SERPINE1 might be associated with chemotherapy resistance and exhibited the highest alteration frequency (predominantly amplification) and overexpression in LSCC tissues. Its knockdown significantly suppressed proliferation, migration, invasion and chemoresistance in LSCC cells. Immunohistochemistry (IHC) confirmed tumor SERPINE1 overexpression (P=0.002 vs. normal tissues), correlating with poor survival (P<0.001). CONCLUSIONS: The 4-IRG risk signature is a reliable prognostic indicator reflecting immune dysfunction in LSCC. SERPINE1 is validated as a therapeutic target and biomarker, enriching our understanding of gene regulation dynamics in LSCC.

Laryngeal cancer

Clinical trials of primary care treatments for major depression: issues in design, recruitment and treatment.

The objective of this article is to consider whether randomized clinical trials (RCTs) are able to determine the validity of transferring treatments for major depression from the psychiatric to the primary care sector. This clinical issue is of growing concern in the United States since both governmental and professional bodies are establishing guidelines for the treatment of medical patients with the affective disorder. The article's method involves analysis of how the competing aims of rigorous scientific methodology (internal validity) and generalization of study findings (external validity) are best balanced within the RCT. Experiences in recruiting medical patients with major depression and providing pharmacologic, psychotherapeutic, and usual care interventions compatible with the sociotechnical characteristics of ambulatory medical centers are described to illustrate the complexities of investigating transferability of treatments for major depression with RCT methodology.

Antidepressive Agents

Clinical trials in the elderly. Pivotal points.

A clinical trial, that is, the scientific assessment of drug action in humans, must be undertaken only if there is reference to an expectation of benefit from the trial. Before the trial starts, a series of questions should be answered, including the need for the trial in elderly patients, particularly in view of the possible vulnerability of the elderly study subjects. Patient selection, randomization, follow-up, analysis, and interpretation must be scientifically valid. Of major importance is the external validity of the trial, that is, the generalizability. Any trial, particularly those involving the elderly, should be designed to develop or contribute to generalizable knowledge; that is, it should be possible to extend the conclusions from the trial beyond the study population to the population at large. When elderly are involved, as they should be when a drug is proposed for use mainly in the elderly population, it is especially important that the study be scientifically valid, medically important, and ethically sound. Studies involving the elderly should have sufficient numbers of females and minorities, the former because most elderly patients are females, the latter because minorities now reach the age of 65 years and beyond more so than in the past. Both risk and benefit should be addressed in terms of potential magnitude and duration. If the study drug has a narrow therapeutic window, it should be intensively studied. When the study is completed, there ought to be clear guidelines for the clinician to design initial, individualized, optimal dosage regimens or for subsequent adjustment of the regimen. The final report, which should be easily evaluable by clinicians, should fully discuss reasons for dropouts, inappropriate patient inclusion, number and types of adverse reactions, and defects in design and conduct of the study.

Aged

Temporal lobe signs: electroencephalographic validity and enhanced scores in special populations.

Internal and external validity tests were completed for an inventory that has been used to infer signs of temporal lobe lability. Strong, positive correlations were reported for a normal (reference) population between the numbers of responses that referred to paranormal experiences (including feelings of a "presence") and separately to religious beliefs and the numbers of spikes per minute within electroencephalographic recordings from the temporal lobe. Numbers of spikes were also correlated with the subjects' scores on the hysteria, schizophrenia, and psychasthenia scales from the MMPI. These clusters of items were not correlated with electrical activity from the occipital lobe (the comparison region). Numbers of responses to control clusters of mundane experiences were not correlated with the temporal lobe measures. A group of student poets scored higher on different subclusters of temporal lobe signs and on the schizophrenia and mania scales of the MMPI than the reference group. For both groups, there were positive correlations between the amount of alpha activity in the temporal lobe only and answers to items such as "hearing inner voices" and "feeling as if things were not real." These results demonstrate that quantitative measures of electrical changes in the temporal lobe are correlated with (or with the report of) specific experiences that are prevalent during surgical or epileptic stimulation of this brain region.

Adult

Artificial Intelligence for Diagnosing Meibomian Gland Dysfunction: A Systematic Review and Meta-Analysis of Diagnostic Test Accuracy Studies.

PURPOSE: To identify, appraise, and synthesize the performance of artificial intelligence-based meibography reading as compared with human graders in diagnosing meibomian gland dysfunction. METHODS: We followed Cochrane methodology and reporting guidelines for diagnostic test accuracy reviews. To assess potential risk of bias and applicability, we used a modified Quality Assessment of Diagnostic Accuracy Studies-2 checklist. We applied bivariate logistic models to estimate summary sensitivity and specificity when appropriate and used the GRADE framework to rate the certainty of the evidence. RESULTS: We identified 14 eligible studies involving 5511 predominantly middle-aged participants (average age: 27-55 years) who were primarily female (&#x2265;54.5%). A total of 18,926 meibography images were obtained through noncontact infrared (11 studies) or in vivo confocal microscopy (three studies). Two studies reported external validation of deep learning models, 12 reported internally validated models, and one reported both. All but one study had high risk of bias in at least one domain; 12 studies raised high or intermediate concern about applicability. Based on three external evaluations, the summary sensitivity and specificity for diagnosing meibomian gland dysfunction from normal glands were 97.5% (95% confidence interval: 77.5%-99.8%) and 85.5% (95% confidence interval: 47.3%-97.5%). Sources of heterogeneity in internally validated models included study population, case mix, and others. The overall evidence was very low to low certainty because of imprecision, high risk of bias, and concerns about applicability. CONCLUSIONS: Artificial intelligence-based meibography grading appears less accurate than human graders. Future studies should adopt rigorous designs, including a more diverse participant pool (or image set), and external validation.

Humans

Multi-national, multi-lingual, multi-professional CATs: (Curriculum Analysis Tools).

A consortium of dental schools and allied dental programs was established in 1991 with the expressed purpose of creating a curriculum database program that was end-user modifiable [1]. In April of 1994, a beta version (Beta 2.5 written in FoxPro(TM) 2.5) of the software CATs, an acronym for Curriculum Analysis Tools, was released for use by over 30 of the consortium's 60 member institutions, while the remainder either waited for the Macintosh (TM) or Windows (TM) versions of the program or were simply not ready to begin an institutional curriculum analysis project. Shortly after this release, the design specifications were rewritten based on a thorough critique of the Beta 2.5 design and coding structures and user feedback. The result was Beta 3.0 which has been designed to accommodate any health professions curriculum, in any country that uses English or French as one of its languages. Given the program's extensive use of screen generation tools, it was quite easy to offer screen displays in a second language. As more languages become available as part of the Unified Medical Language System, used to document curriculum content, the program's design will allow their incorporation. When the software arrives at a new institution, the choice of language and health profession will have been preselected, leaving the Curriculum Database Manager to identify the country where the member institution is located. With these 'macro' end-user decisions completed, the database manager can turn to a more specific set of end-user questions including: 1) will the curriculum view selected for analysis be created by the course directors (provider entry of structured course outlines) or by the students (consumer entry of class session summaries)?; 2) which elements within the provided course outline or class session modules will be used?; 3) which, if any, internal curriculum validation measures will be included?; and 4) which, if any, external validation measures will be included. External measures can include accreditation standards, entry-level practitioner competencies, an index of learning behaviors, an index of discipline integration, or others defined by the institution. When data entry, which is secure to the course level, is complete users may choose to browse a variety of graphic representations of their curriculum, or either preview or print a variety of reports that offer more detail about the content and adequacy of their curriculum. The progress of all data entry can be monitored by the database manager over the course of an academic year, and all reports contain extensive missing data reports to ensure that the user knows whether they are studying complete or partial data. Institutions using the beta version of the program have reported considerable satisfaction with its functionality and have also offered a variety of design and interface enhancements. The anticipated release date for Curriculum Analysis Tools (CATs) is the first quarter of 1995.

Curriculum

Ecological validity and cultural sensitivity for outcome research: issues for the cultural adaptation and development of psychosocial treatments with Hispanics.

This article has two objectives. The first is to provide a culturally sensitive perspective to treatment outcome research as a resource to augment the ecological validity of treatment research. The relationships between external validity, ecological validity, and culturally sensitive research are reviewed. The second objective is to present a preliminary framework for culturally sensitive interventions that strengthen ecological validity for treatment outcome research. The framework, consisting of eight dimensions of treatment interventions (language, persons, metaphors, content, concepts, goals, methods, and context) can serve as a guide for developing culturally sensitive treatments and adapting existing psychosocial treatments to specific ethnic minority groups. Examples of culturally sensitive elements for each dimension of the intervention are offered. Although the focus of the article is on Hispanic populations, the framework may be valuable to other ethnic and minority groups.

Cross-Cultural Comparison

DSM-IV: empirical guidelines from psychometrics.

This commentary addresses the use of psychometric theory and methodology in the development of the 4th edition of the Diagnostic and Statistical Manual of Mental Disorders (DSM-IV). Reliability issues include interdiagnostician reliability, temporally consistent diagnoses, and the relations of diagnostic criteria within categories. Validity issues include content validity of the diagnostic criteria, criterion-related validity (the relation between different criterion sets or their algorithms and alternative diagnostic criteria), and construct validity (the relation between diagnostic categories and external validators). Specific questions and methodology to investigate its utility vary with the different uses proposed for the diagnostic system. Specific psychometric methodologies that may be useful in developing the DSM-IV are noted, as are the limitations of psychometrics and their applicability to DSM-IV.

Humans