PubMed HealthSearch

SEARCH · PubMed Health

Results for “External validity”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

A comparative evaluation of multiple enlarged perivascular space segmentation tools.

BACKGROUND: Enlarged perivascular spaces (ePVS) are a marker of cerebral small vessel disease, potentially reflecting reduced waste clearance. Because manual quantification is unfeasible in large datasets, we developed and evaluated an automated tool. METHODS: Detection Of Regions of Enlarged perivascular Spaces (DORES), a 3D nnU-Net-based deep learning algorithm was developed for ePVS segmentation using T1-weighted and fluid-attenuated inversion recovery magnetic resonance imaging (MRI). DORES was developed in two stages: an initial model trained on 35 manually segmented scans and a final model on 1460 pseudo-labeled sessions from the Vanderbilt Memory and Aging Project (VMAP). A subset of VMAP participants with 3 T brain MRI underwent whole-brain manual ePVS tracing (n = 35, 73 ± 9 years, 51% male) and visual rating (n = 388, 71 ± 8 years, 54% male) by a neuroradiologist. DORES was evaluated and compared against three other segmentation tools using Dice and F1 scores, absolute volume and element differences, correlation, and agreement. External validation used an Alzheimer's Disease Neuroimaging Initiative 3 subset with manual tracings (ADNI3, n = 18, 73 ± 9 years, 67% female). RESULTS: DORES achieved Dice scores of 0.61 ± 0.16 (white matter) and 0.72 ± 0.08 (basal ganglia) in VMAP, with strong correlations and agreement for ePVS count and volume. Performances modestly declined in ADNI3 across algorithms. Scanner-stratified analyses showed stronger correlations for Philips versus Siemens images in the basal ganglia, indicating scanner-dependent differences in measurement consistency. CONCLUSIONS: DORES provides a multimodal nnU-Net-based pipeline for ePVS segmentation in older adults. The model demonstrates robust within-cohort performance and reasonable external validity, though scanner-related effects limit application across sites.

Humans

Systemic Proteome Profiling to Differentiate Primary Glomerular Diseases.

KEY POINTS: Plasma proteome profiling identified distinct signatures across biopsy-proven primary glomerular disease subtypes. An elastic net model using 93 proteins classified primary glomerular disease subtypes and controls, with external validation. Integrating proteomics with machine learning yields biologically interpretable insights in primary glomerular diseases. BACKGROUND: Primary GN is a heterogeneous group of kidney disorders where understanding of their pathophysiology remains incomplete. Despite the diagnostic potential of high-throughput proteomics, constrained proteomic depth and a reliance on binary comparisons have left the feasibility of using systemic signatures to differentiate multiple GN subtypes largely unexplored. METHODS: To identify protein signatures that noninvasively differentiate major primary glomerular disease subtypes and provide mechanistic insights, we performed large-scale systemic proteome profiling of 5416 plasma proteins via Olink Explore HT in a discovery cohort ( n =147) and an external validation cohort ( n =85) of Korean participants (mean age, 41±13 years; 46% female). The study population included patients with four GN subtypes-focal segmental glomerulosclerosis, IgA nephropathy, minimal change disease, and membranous nephropathy-alongside healthy controls. We developed a machine learning (ML) model using logistic regression with elastic net regularization to classify disease groups based on proteomic profiles and evaluated its performance in the independent validation cohort. RESULTS: Plasma proteome profiles were distinct among disease subtypes, emerging as a significant source of data variation independent of conventional markers such as eGFR or proteinuria levels. The ML model performed robustly in both the discovery and validation cohorts, achieving an area under the receiver operating characteristic curve >0.8 for differentiating minimal change disease, membranous nephropathy, and IgA nephropathy. The model, even without clinical information, correctly identified 93% of minimal change disease cases (14 of 15) and 63% of IgA nephropathy cases (20 of 32), but its performance was limited for focal segmental glomerulosclerosis, with only 21% of cases (three of 14) correctly classified. Functional analysis of key proteins highlighted distinct biologic pathways, such as hemostasis in minimal change disease. CONCLUSIONS: We identified distinct systemic proteome signatures for primary glomerular diseases, where disease subtype served as a major determinant of proteomic variance alongside conventional clinical markers. ML models demonstrated robust discriminatory performance for minimal change disease, membranous nephropathy, and IgA nephropathy, underscoring the potential for proteome-based classification.

Humans

Resolution-dependent self-supervised transfer in chest radiograph classification.

BACKGROUND: Self-supervised learning (SSL) has improved visual representation learning, but its value in chest radiography remains uncertain. DINOv3 extends earlier SSL models through Gram-anchored self-distillation and explicit high-resolution adaptation. Whether these changes improve transfer learning for chest radiograph classification has not been established. METHODS: We benchmarked DINOv3 against DINOv2 and supervised ImageNet initialization across seven chest radiograph datasets comprising 816,183 radiographs from pediatric and adult cohorts. ViT-B/16 and ConvNeXt-B were evaluated under full fine-tuning at 224 × 224 and 512 × 512 pixels, with targeted 1024 × 1024 experiments on three cohorts. Additional analyses examined parameter-efficient adaptation, synthetic label corruption, external validation, frozen 7B features, and computational efficiency. The primary outcome was the mean area under the receiver operating characteristic curve across labels. RESULTS: In adult cohorts, DINOv3 did not consistently outperform DINOv2 at 224 × 224 pixels, but became the strongest initialization at 512 × 512 pixels, especially with ConvNeXt-B. Gains were greatest for small focal and boundary-dependent abnormalities, whereas large-structure findings changed little. The pediatric cohort showed no significant benefit from DINOv3, higher resolution, or backbone choice. Scaling to 1024 × 1024 rarely improved performance and markedly increased computational cost. ConvNeXt-B remained superior to ViT-B/16 under both full and parameter-efficient adaptation. External validation preserved the 512 × 512 DINOv3 advantage, whereas synthetic label corruption showed that this benefit should not be interpreted simply as superior noise robustness. Frozen DINOv3-7B features underperformed relative to fully adapted 86 to 89M-parameter backbones. CONCLUSIONS: For adult chest radiograph classification, DINOv3 provides its most reliable benefit at 512 × 512 pixels, particularly with ConvNeXt-B. Fully adapted mid-sized models at 512 × 512 pixels provided the best performance-cost trade-off in our benchmark.

Journal Article

Deep learning-based cross-attention fusion of multimodal MRI for survival prediction and risk stratification in IDH-wildtype glioblastoma: a multicenter study.

BACKGROUND: Glioblastoma (GBM) exhibits profound molecular and spatial heterogeneity, complicating prognostic evaluations. While multiparametric MRI provides crucial multidimensional biological information, conventional end-to-end deep learning integration strategies, such as early or late fusion, often fail to capture complex nonlinear cross-modal interactions. We aimed to systematically evaluate a cross-attention fusion (CAF) architecture for GBM survival prediction and quantify its incremental prognostic value relative to existing clinical tools. METHODS: In this multicenter retrospective study, 386 adults with IDH-wildtype, WHO grade 4 GBM were assembled from an institutional cohort (n = 226), the Chinese Glioma Genome Atlas (CGGA, n = 62), and The Cancer Genome Atlas (TCGA, n = 98). Using a unified 3D ResNet-18 backbone, we compared single-modality models, early fusion, late fusion, and CAF on preoperative T1-weighted, contrast-enhanced T1-weighted (T1CE), and T2-weighted MRI, and integrated the resulting deep learning risk score with routine clinical variables through multivariable Cox regression. Performance was assessed using Harrell's C-index, time-dependent AUC, and decision curve analysis. RESULTS: CAF showed numerically higher, more consistent C-index trends than early fusion, late fusion, and single-modality models (pooled C-index 0.629, 95% CI 0.594-0.664), although pairwise differences in time-dependent AUC were not statistically significant. Integrating clinical variables raised the pooled C-index to 0.691 (95% CI 0.660-0.721) in the treatment-era model, with comparable performance across the three cohorts (Local 0.688; CGGA 0.716; TCGA 0.689); a pre-treatment configuration excluding adjuvant therapy yielded a pooled C-index of 0.642. Under leave-one-cohort-out external validation, the combined model retained significant risk stratification in all held-out cohorts (C-index 0.63-0.71; all log-rank P&#xa0;<&#xa0;0.01), albeit with attenuated discrimination. The deep learning risk score remained independent after multivariable adjustment (HR 1.41 per SD, 95% CI 1.26-1.57; P&#xa0;<&#xa0;0.001). Kaplan-Meier analysis confirmed significant high- versus low-risk separation in all cohorts, and decision curve analysis showed greater net benefit than clinical-only and deep-learning-only models. CONCLUSION: The CAF-derived risk score offers prognostic information complementary to routine clinical variables, representing a promising noninvasive tool for individualized risk stratification when molecular profiling is incomplete or unavailable; these findings warrant prospective external validation before clinical use.

cross-attention fusion

Integrated Genomic and Proteomic Analysis Reveals T-B Lymphocyte Signatures in the MYCN Driven "Immune Desert" of Specific Neuroblastoma Subtypes.

AIMS: This study aims to systematically dissect how MYCN amplification shapes the immunosuppressive tumor microenvironment (TME) in high-risk neuroblastoma, elucidating key mechanisms underlying immune evasion. METHODS: We performed an integrated multi-omics analysis of bulk RNA-seq (n&#x2009;=&#x2009;721), single-cell RNA-seq (n&#x2009;=&#x2009;9), proteomic data (n&#x2009;=&#x2009;49) and spatial transcriptomics (Visium, with external validation in melanoma). Analyses included unsupervised clustering, cell-cell communication inference, transcriptional regulatory network reconstruction, and spatial proximity assessment to map the immune landscape. RESULTS: A distinct molecular subtype (Class C), defined by MYCN amplification and poor prognosis, exhibited a comprehensive "immune desert" phenotype characterized by low immune scores and minimal leukocyte infiltration. Single-cell analysis confirmed significant depletion of T and B lymphocytes within the Class C TME. Dysregulated transcriptional networks were identified, including upregulation of REL and EOMES in T cells-with EOMES potentially driving exhaustion via regulation of Transient Receptor Potential (TRP) genes, and REL inhibition enhancing cytotoxic function in&#xa0;vitro. A unique immunosuppressive B-cell subset (B7) engaged in enhanced crosstalk with exhausted T cells and harbored a MYC-centered network linked to cell cycle dysregulation and poor survival. Spatial transcriptomics revealed significant proximity between B7-active regions and Treg/exhaustion-enriched areas, externally validated in melanoma. Proteomic data validated elevated REL expression in MYCN-amplified tumors. CONCLUSION: This work delineates the immunosuppressive architecture of MYCN-driven neuroblastoma, revealing novel regulatory nodes within specific lymphocyte compartments. Integrating single-cell, spatial, and proteomic evidence, we propose REL inhibition as a therapeutic candidate, the EOMES/TRP axis as a bioinformatically supported hypothesis, and the B7/MYC hub as a hypothesis supported by transcriptomic and spatial evidence.

Humans

Integrating genetic predictors into subsequent breast cancer risk prediction in survivors of childhood cancer.

PURPOSE: Female survivors of childhood cancer are at high risk for developing breast cancer. The contributions of most general population primary breast cancer genetic predictors to this risk have not been explored. METHODS: Analyses included females who survived &#x2265;5 years after their childhood cancer diagnosis with available array (N&#x2009;=&#x2009;2096, subsequent breast cancer [SBC]=218) or whole-genome sequencing (WGS; N&#x2009;=&#x2009;3292, SBC=101) data from the Childhood Cancer Survivor Study and St. Jude Lifetime Cohort. We computed 99 externally-validated primary breast cancer polygenic risk scores (PRS). Using deep-coverage WGS, ClinVar-annotated pathogenic/likely pathogenic (P/LP) variants in breast cancer susceptibility genes were identified. Cox proportional hazards models assessed associations with SBC risk, adjusting for treatments and genetic ancestry. RESULTS: Among 5388 female survivors (genetic ancestry, European: N&#x2009;=&#x2009;4,752; African: N&#x2009;=&#x2009;444; East Asian: N&#x2009;=&#x2009;192), 319 developed SBC. Most (90.9%) PRSs were nominally associated with SBC risk (P&#x2009;<&#x2009;0.05), but effect sizes varied substantially. PRSs with superior discriminatory ability had greater genome-wide coverage (e.g., 6.4 million-variant PRS, HR per SD&#x2009;=&#x2009;1.71, 95% CI&#x2009;=&#x2009;1.43 to 2.05; P&#x2009;=&#x2009;4.2x10-9) and 7.7-fold higher odds (P&#x2009;=&#x2009;7.0x10-4) of including variants in multiple DNA damage repair pathways compared with PRSs with weaker risk associations. Among survivors with WGS, 1.6% carried P/LP variants in clinical testing panel genes, which was associated with a 7.4-fold greater risk (95% CI&#x2009;=&#x2009;3.16 to 17.19). Including genetic factors improved SBC risk prediction by age 40 (P&#x2009;<&#x2009;0.001) compared to treatment exposures alone. CONCLUSIONS: Externally-validated primary breast cancer genetic susceptibility predictors are relevant for SBC risk prediction and should be prioritized for risk stratification in survivors.

Journal Article

Recruitment issues, health habits, and the decision to participate in a health promotion program.

To understand the external validity of experimental studies, it is important to estimate the extent to which the participants are representative of the general population. This paper describes recruitment methods and considers the representativeness of participants in the San Diego Family Health Project. The study was designed to experimentally evaluate the effectiveness of a family-based behavior change intervention in Anglo and Mexican-American families. Initial contact with the families was made through a household health survey that was sent home with all fifth- and sixth-grade children in 12 participating elementary schools. The survey asked about a variety of demographic characteristics, dietary habits, and physical activity habits. Parents were also asked if they were interested in participating in the project. Respondents were classified by level of participation into one of three groups: not interested, expressed initial interest but did not attend the recruitment meeting, and volunteered to participate. Level of participation was the independent variable in the analyses. In separate analyses for Anglo and Mexican-American responders, our data suggested many similarities and a few differences among participant groups. The differences that were observed suggest that participants may already have healthier diets than nonparticipants, although only one of four dietary variables differed by participation status in each ethnic group. The external validity of these data and general recruitment issues are discussed.

Adolescent

An external construct validity study of Rorschach personality variables.

This study examined (a) hypothesized relationships between Rorschach variables and self-report test measures relating to nominally similar aspects of personality functioning and (b) interrelationships among Rorschach variables. Sixty-two undergraduates were administered the Rorschach, Barron Ego Strength Scale, Kaplan Self-Derogation Scale, Eagly Self-Esteem Scale, Multiple Affective Adjective Checklist (MAACL), Marlowe-Crowne Social Desirability Scale, and the Rotter Locus of Control Scale. Only a few of the predictions received confirmation: inanimate movement (m) correlated, as expected, with MAACL anxiety and hostility, the egocentricity index (3r + 2)/R (R = total responses) correlated significantly with self-esteem, and human movement with minus form level (M-) correlated (inversely) with ego strength. Among the unpredicted findings were some that appear inconsistent with standard Rorschach interpretation. Rorschach variables human movement (M), and experience actual (EA), generally interpreted as reflecting coping resources, related significantly with self-report measures of poor coping and of dysphoric affect. In general, the Rorschach appears better at identifying weaknesses in the ego rather than strengths.

Adolescent

Multimodal features and prognostic risk assessment in locally advanced gastric cancer patients following neoadjuvant therapy based on machine learning algorithms: a multicenter study.

BACKGROUND: Neoadjuvant therapy (NAT) is recommended for locally advanced gastric cancer (LAGC), but some patients respond poorly. We aimed to construct a multimodal model integrating CT images, transcriptomic sequencing, and clinicopathological data to assess prognosis in LAGC patients receiving NAT. MATERIALS AND METHODS: This multicenter study included 505 LAGC patients who underwent NAT. Radiomic features were extracted from preoperative CT images of 505 patients. RNA-seq was performed on 277 post-NAT specimens, with additional data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases (n&#x2009;=&#x2009;804). Patients were divided into training (168 cases), internal validation (72 cases), and external validation cohorts. Machine learning algorithms identified key radiomic, molecular, and clinical features associated with NAT response, which were then integrated into a multimodal model to predict overall survival (OS) and disease-free survival (DFS). RESULTS: Six radiomic and three molecular features significantly associated with NAT response were selected. Radiomic risk (hazard ratio [HR]: 4.0, P&#x2009;<&#x2009;0.001) and molecular risk (HR: 7.1, P&#x2009;<&#x2009;0.001) were independent prognostic factors. By integrating radiomic risk, molecular risk, and clinical characteristics, a multimodal model (MuMo) was constructed.The C-index results (OS, C-index&#x2009;=&#x2009;0.855; DFS, C-index&#x2009;=&#x2009;0.786) demonstrated that MuMo outperformed the single-modality models and ypTNM staging.Mechanistic analysis suggested that the efficacy of neoadjuvant therapy was significantly enriched in immune-inflammatory pathways. CONCLUSIONS: MuMo can effectively predict postoperative survival risk in LAGC patients receiving NAT, serving as a powerful tool for optimizing prognostic assessment.

Humans

Noninvasive detection and differentiation of gastric malignancy using cell-free DNA biomarkers.

INTRODUCTION: Gastric cancer remains a major global health burden, with high mortality driven by late-stage diagnoses that limit treatment options and reduce survival. Current diagnostic methods such as endoscopy and biopsy are invasive, resource-intensive, and impractical for large-scale early detection. OBJECTIVES: This study aimed to develop and validate an ensemble machine learning model integrating four cell-free DNA (cfDNA) fragmentomic feature classes derived from 5&#xa0;&#xd7;&#xa0;whole genome sequencing (WGS) data to non-invasively differentiate malignant gastric cancer from benign gastric lesions in high-risk or symptomatic patients. METHODS: A total of 681 plasma samples were prospectively collected, comprising 329 from patients with gastric cancer or high-grade intraepithelial neoplasia (HGIN) and 352 from individuals with benign gastric conditions. The dataset was divided into a training cohort (n&#xa0;=&#xa0;333) and a temporally independent validation cohort (n&#xa0;=&#xa0;348). An external validation cohort of 305 participants was also included. RESULTS: The ensemble model achieved an AUROC of 0.920 in cross-validation testing on the training cohort, 0.912 in the independent validation cohort, and 0.896 (95% CI 0.860-0.932) in the external cohort. At a pre-specified prediction threshold of 0.402, the model demonstrated 93.3% sensitivity and 71.9% specificity in the validation cohort, yielding a PPV of 71.3% and an NPV of 93.5%. In the external cohort, sensitivity and specificity were 91.7% and 69.1%, respectively (PPV 75.7%, NPV 88.8%). Model scores correlated with clinical stage, tumor grade, and histopathological subtype. Approximately 71% of non-cancer patients could have been spared unnecessary endoscopy. CONCLUSIONS: The cfDNA fragmentomics-based ensemble model enables accurate, non-invasive differentiation between gastric cancer and benign gastric lesions in high-risk or symptomatic patients. This approach demonstrates strong potential as a pre-endoscopy triage tool, supporting earlier detection and more efficient use of diagnostic resources.

Humans

A principal-components analysis of the Narcissistic Personality Inventory and further evidence of its construct validity.

We examined the internal and external validity of the Narcissistic Personality Inventory (NPI). Study 1 explored the internal structure of the NPI responses of 1,018 subjects. Using principal-components analysis, we analyzed the tetrachoric correlations among the NPI item responses and found evidence for a general construct of narcissism as well as seven first-order components, identified as Authority, Exhibitionism, Superiority, Vanity, Exploitativeness, Entitlement, and Self-Sufficiency. Study 2 explored the NPI's construct validity with respect to a variety of indexes derived from observational and self-report data in a sample of 57 subjects. Study 3 investigated the NPI's construct validity with respect to 128 subject's self and ideal self-descriptions, and their congruency, on the Leary Interpersonal Check List. The results from Studies 2 and 3 tend to support the construct validity of the full-scale NPI and its component scales.

Adolescent

Risk Prognostication After Hypomethylating Agents Combined With Venetoclax in AML: The PRISM Risk Model.

PURPOSE: As risk stratification for patients with AML treated with lower-intensity venetoclax-based therapy remains suboptimal, we developed and validated a prognostic model integrating clinical, cytogenetic, and molecular features. METHODS: We assembled a multinational data set comprising 2,092 adults with newly diagnosed AML treated with hypomethylating agents plus venetoclax (HMA + VEN). One thousand nine hundred eighteen patients with complete data were randomly divided into training (70%) and internal validation (30%) cohorts. Two independent external validation cohorts were assembled (n = 500 and n = 222). Modeling overall survival (OS), Elastic Net regression was applied in 1,000 bootstrap samples from the training cohort to select variables for a Ridge regression, which generated a continuous Prognostic Risk Integration for Survival Modeling (PRISM) score and risk categories based on tertiles (PRISM-3: low, moderate, high). These PRISM indices were then computed for the validation cohorts and compared with the 4-gene classifier (based on mutations in FLT3-ITD, N/KRAS, and TP53). RESULTS: PRISM integrated 17 clinical and genomic variables and demonstrated a linear association with OS. PRISM-3 stratified survival consistently across all cohorts (median OS: 25.1-28.8 months for low risk, 12.5-14.7 months for moderate risk, and 5.8-6.7 months for high risk; P < .001). Compared with the 4-gene classifier, PRISM-3 reassigned approximately 40% of patients (and >50% of those with favorable risk) and demonstrated significantly better discrimination in validation cohorts (C-index 0.63-0.65 v 0.59-0.61; P < .05). CONCLUSION: PRISM is a validated prognostic model for patients with AML receiving HMA + VEN that improves survival risk stratification beyond current standard tools and supports individualized, risk-adapted clinical decision making. The model, the PRISM-AML Risk Calculator, is publicly available.

Humans

Evaluation and measurement: some dilemmas for health education.

Seven dilemmas of evaluation and measurement posed by the nature of health education are presented, together with suggestions for their resolution. These include the dilemmas of : 1) rigor of experimental design vs significance or program adaptability; 2) internal validity or "true" effectiveness vs external validity or feasibility; 3) experimental vs placebo effectsl 4) effectiveness vs economy of scale; 5) risk vs payoff; 6) measurement of long-term vs short-term out-comon. Emphasis is placed on the need to develop a more cumulative data base through standardization of measures, replication of experiments in different settings, and better documentation, reporting, and diffusion of experiences in practice.

Cost-Benefit Analysis

The potential of clustering methods for pre-test triage in sleep medicine: A systematic review.

Sleep disorders exhibit substantial heterogeneity, and traditional classifications may not fully capture clinically relevant subtypes. Clustering techniques can identify patient subgroups that improve phenotypic characterization and may support personalized management. This systematic review evaluated the application of clustering in sleep medicine, with particular focus on its potential use as a pre-test triage tool prior to formal sleep testing. PubMed/MEDLINE, Embase, Web of Science, and Scopus were searched to February 2025. Eligible studies applied clustering to classify sleep disorders in adults. Two reviewers independently conducted screening, data extraction, and risk-of-bias assessment using QUADAS-2. The protocol was registered on PROSPERO. Fifty-one studies (1983-2025) were included, predominantly focused on obstructive sleep apnea (OSA) (n&#x202f;=&#x202f;38, 74%). Hierarchical clustering (n&#x202f;=&#x202f;20) and K-means clustering (n&#x202f;=&#x202f;14) were the most frequently used techniques. Internal validation was reported in only 18% of studies, and external validation was reported in only 1 study. Seven studies relied exclusively on baseline clinical, demographic, or questionnaire data, representing pre-test scenarios, whereas most incorporated polysomnography-derived variables, limiting their applicability to early clinical stratification. Hierarchical clustering was the most commonly applied method; however, the overall lack of validation limits confidence in the robustness and clinical applicability of identified phenotypes. The potential role of clustering as a pre-test triage strategy remains largely unexplored, as most studies focused on post-diagnostic phenotyping and were affected by incorporation bias. Future research should prioritize pre-test clinical variables, rigorously validate internally and externally, and adopt standardized methodological and reporting practices to facilitate clinical translation.

Humans

Integrated Genomic and Tumor Microenvironment Subtyping Improved Risk Stratification in Primary Central Nervous System Lymphoma.

Current prognostic models fail to capture the biological complexity of primary central nervous system lymphoma (PCNSL). We integrated whole-genome sequencing and multiplex immunofluorescence in 68 treatment-na&#xef;ve patients to define four genomic subtypes (C1, C2, C3, and C4) with divergent survival (C4 worst: median overall survival [OS], 26&#x2009;months). In parallel, a novel tumor microenvironment (TME) classification based on CD8+T/M2 macrophage ratio stratified patients into High (>&#x2009;1.5), Intermediate (0.8-1.5), and Low (<&#x2009;0.8) groups. Unexpectedly, the Intermediate TME group showed the poorest outcomes (5-year OS: 10%). Integration revealed a lethal subgroup (C4&#x2009;+&#x2009;Intermediate TME; 9.8% of cohort) with a median OS of 3.0&#x2009;months (hazard ratio&#x2009;=&#x2009;7.24, p&#x2009;=&#x2009;0.006). Prognostic nomograms incorporating these subtypes showed promising discriminative performance in internal validation (C-index >&#x2009;0.78), but external validation is needed. Together, these findings identify a high-risk biological subset and provide a hypothesis-generating framework for future biomarker-driven risk stratification and therapeutic discovery in PCNSL.

Humans

Age-based construct validation using structural equation modeling.

In this paper we describe some mathematical and statistical models based on structural equation modeling (SEM) using computer programs like LISREL. We focus on SEM methodology for the simultaneous examination of the internal validity of psychological constructs and the external validity represented by age relations. To illustrate these ideas we use a latent variable path model to examine the organization of intellectual abilities measured by the WAIS-R in the standardization sample. We also examine different ways in which age can be used to structure this organization. This is primarily a methodological paper, but we try to integrate conceptual principles of modeling with some substantive issues of research on the psychology of aging.

Aging

Confidence intervals versus p-values for interpretation of clinical trial results: introduction.

The following three papers summarize the presentations at a Society for Clinical Trials annual meeting session on the relative merits of estimation versus testing for analysis of randomized clinical trials. By design, randomized clinical trials have internal validity. Whether they also possess quantitative external validity--generalizability of effect size to some population represented by the trial subjects--is one of the main points of disagreement among the three authors. It may be unrealistic to expect a resolution that applies across the wide variety of therapeutic areas and clinical trial goals. Extrapolation from clinical trial to clinical practice is often endorsed in connection with large trials having loose entry criteria and focusing on an objective, clearly meaningful clinical endpoint. By contrast, the relevance of estimates of effect size is less clear in the case of many clinical trials conducted in the course of drug development.

Clinical Trials as Topic