PubMed HealthSearch

SEARCH · PubMed Health

Results for “SHAP interpretability”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Graph neural network-based risk stratification of prostate cancer using gene expression and SHAP interpretability.

Accurate risk stratification is essential for guiding treatment decisions and preventing over treatment of prostate cancer, which remains one of the most prevalent cancers among adult men. While the Gleason score, obtained from prostate biopsies, is routinely used to assess tumor aggressiveness, the biopsy procedure carries risks such as pain, infection, and, in some cases, serious complications such as sepsis. In this study, we proposed an artificial intelligence-based framework that integrates mRNA expression profiles with functional interaction networks to classify prostate cancer patients into low-, medium-, and high-risk groups defined by Gleason scores. The pipeline comprised five steps: (1) data collection from The Cancer Genome Atlas (TCGA), (2) preprocessing of gene expression data, (3) two-stage feature selection to identify informative biomarkers, (4) risk classification using a dual-branch graph neural network (GNN) that combines gene-gene interaction graphs with sample-level expression features, and (5) model interpretation using SHAP to quantify feature contributions. Differentially expressed genes were identified in the High (ASPN, GMNN, PEBP4, C2, KNCK17), Medium (C2, IGSF1, ASPN, CDKN3, AMH), and Low (TNMD, VWA5B2, ST6GALNAC5, CYP3A5, PHGR1) risk groups, underscoring the molecular heterogeneity of disease progression. On an independent held-out test set, the model achieved AUCs of 0.86, 0.88, and 0.95 for the low-, medium-, and high-risk groups, respectively, with an overall accuracy of 80%. These results suggest that combining GNN-based modeling with explainable AI can capture both global and local molecular patterns relevant to tumor aggressiveness. However, as the model was developed and evaluated solely on the TCGA cohort, the findings should be regarded as exploratory, and external validation will be required to establish generalizability. Within these limitations, the proposed framework highlights the potential of molecular profiling and graph-based deep learning to support more precise, potentially less invasive, risk assessment and individualized treatment planning in prostate cancer.

Prostatic Neoplasms

Attribution of PM2.5-Induced Transcriptomic Perturbation to Toxic Components.

Ambient fine particulate matter (PM2.5) is a chemically complex mixture whose health impacts are not fully captured by particle mass. Here, we developed an interpretable chemotranscriptomic framework to attribute PM2.5-induced molecular perturbations to toxicity-relevant components. PM2.5 collected from urban roadside and coastal environments was separated into whole, extractable, and unextractable fractions, characterized by LC/GC × GC-HRMS-based nontarget analysis and inductively coupled plasma mass spectrometry (ICP-MS), and evaluated using cytotoxicity testing and transcriptomic profiling in human bronchial epithelial cells. Urban PM2.5 exhibited greater cytotoxic potency per unit mass than coastal PM2.5, with extractable fractions accounting for most cytotoxic and pathway-level responses. Transcriptomics revealed distinct site-specific modes of action: urban PM2.5 preferentially induced oxidative stress, xenobiotic metabolism, and cell cycle suppression, consistent with acute, nonapoptotic injury, whereas coastal PM2.5 elicited weaker cytotoxicity but stronger interferon-mediated immune and apoptosis-related signaling. Integrating chemical abundance with pathway activity using random forest regression, SHAP interpretation, and mechanistic corroboration reduced 5,033 detected features to 444 pathway-linked candidate drivers. Fewer than 5% of features explained ∼95% of cumulative model contribution. Standard-confirmed contributors included plasticizer-related compounds, aromatic and heteroaromatic combustion products, and copper for urban PM2.5 and secondary/aged organics and nickel for coastal PM2.5. These findings support mechanism-informed prioritization of hazardous PM2.5 components beyond mass-based assessment.

Particulate Matter

Pan-cancer multi-omics machine learning defines a lactylation-associated immune-excluded tumor state with proteomic and experimental corroboration.

BACKGROUND: Histone lactylation links lactate metabolism to chromatin regulation, but whether lactylation-program-associated transcriptional patterns delineate recurrent pan-cancer tumor states remains unclear. METHODS: We integrated mRNA, lncRNA, and miRNA profiles from 9712 TCGA tumors across 33 cancer types with GTEx references, six GEO cohorts, IMvigor210, and an institutional clear-cell renal cell carcinoma (ccRCC) cohort used for exploratory DIA-NN proteomic corroboration. Random-effects co-expression meta-analysis, multi-omics consensus clustering, regulon inference, immune deconvolution, TIDE, oncoPredict, and SHAP-based machine learning were applied. hsa-miR-431-5p was functionally evaluated as a proof-of-concept CS2-associated miRNA in bladder cancer models. RESULTS: LacCoEx-Atlas comprised 398,491 lactylation-related co-expression pairs across 24,667 RNA features under a random-effects framework (median I² = 88.6%). Consensus clustering identified two subtypes: CS2 showed glycolytic-mesenchymal-immune-excluded features, M2 macrophage enrichment, CD8⁺ T-cell depletion, elevated HDAC4/NSD3/KDM6B activity, and worse survival, whereas CS1 showed oxidative, sirtuin-active programs. CS2 had fewer predicted ICI responders (18.3% vs. 52.0%) and a lower observed ORR in IMvigor210 (15.3% vs. 24.0%). oncoPredict identified NU7441 as a hypothesis-generating CS2-associated sensitivity signal (Hedges' g = 1.17). DIA-NN proteomics in 50 ccRCC specimens provided exploratory support for CS2-associated hypoxia, ECM degradation, and metastasis programs. The 10-feature mRNA LARItools model achieved an apparent AUC of 0.9413, while a separate multi-omics model achieved 0.971; neither was independently validated. LARItools reproduced prognostic separation across six GEO cohorts. miR-431-5p promoted malignant phenotypes and EMT in bladder cancer cells, with concordant CMU4h expression findings. CONCLUSIONS: Lactylation-program-associated transcriptional patterns delineate a recurrent immune-excluded pan-cancer tumor state associated with adverse prognosis, reduced predicted immunotherapy responsiveness, exploratory single-cancer protein-level support, and testable DNA damage response-targeting hypotheses. LacCoEx-Atlas and LARItools provide open resources for lactylation-program-associated tumor-state stratification and future translational research.

Humans

Diagnosing the undiagnosed: AI-enhanced multimodal modeling for placental mesenchymal dysplasia in high-risk pregnancies.

Placental mesenchymal dysplasia (PMD) is a rare vascular placental disorder that mimics molar pregnancy but often coexists with a viable fetus, making its misdiagnosis potentially devastating. In high-risk pregnancies, artificial intelligence (AI)-enhanced multimodal modeling - incorporating imaging, genomics, proteomics, and clinical features - offers a transformative diagnostic strategy. Leveraging Bayesian hyperparameter optimization for model refinement, this approach improves diagnostic accuracy while reducing uncertainty and clinician hesitation. Recent clinical studies support its efficacy and interpretability through SHAP and LIME models, while real-time surgical enhancements using Bayesian methods highlight its broader clinical utility. Despite current challenges such as data heterogeneity and integration barriers, multimodal AI provides unprecedented resolution in placental analysis, enabling precise differentiation between PMD and similar fetopathies. Ultimately, this advancement supports timely, non-invasive diagnosis, personalized management, and emotionally informed decision-making aligned with ethical AI implementation standards.

Bayesian optimization

Distinct immune-metabolic phenotypes underlie poor coronary collateral circulation.

BACKGROUND: Coronary collateral circulation (CCC) significantly impacts myocardial perfusion and clinical outcomes in coronary artery disease patients, yet the underlying molecular heterogeneity remains inadequately characterized. OBJECTIVE: To identify distinct molecular phenotypes in patients with poor CCC, validate these phenotypes using clinical parameters, and evaluate their prognostic implications. METHODS: This study enrolled 149 patients (80 with good CCC and 69 with poor CCC) for high-throughput proteomic profiling. Unsupervised consensus clustering identified molecular subtypes within poor CCC patients, followed by differential expression analysis and KEGG pathway enrichment. Boruta feature selection was implemented, and multiple machine learning algorithms were tested on clinical data, with XGBoost optimization (accuracy 80.0%, F1-score 80.31%) and SHAP value interpretation. External validation was performed using the MIMIC database. Kaplan-Meier analysis and Cox regression models assessed major adverse cardiovascular events (MACE). RESULTS: Two distinct phenotypes emerged among poor CCC patients: Cluster 1 (n&#x2009;=&#x2009;39, Complement-Driven Vascular Remodeling [CDVR]) and Cluster 2 (n&#x2009;=&#x2009;30, Immuno-Thrombotic Myocardial Dysfunction [ITMD]). An XGBoost model incorporating fasting glucose, eosinophil percentage, and HbA1c achieved excellent discrimination (AUC&#x2009;>&#x2009;0.91). External validation confirmed the phenotype-specific clinical patterns. Notably, Cluster 2 demonstrated significantly higher MACE incidence compared to Cluster 1 (Log-rank p&#x2009;<&#x2009;0.05), with KEGG analysis revealing significant upregulation of platelet activation, diabetic cardiomyopathy, and metabolic pathways in the ITMD phenotype. CONCLUSION: Poor CCC encompasses distinct immune-metabolic phenotypes that can be accurately classified using integrated proteomic-clinical modeling. This classification enables more precise risk stratification and may guide personalized therapeutic strategies for coronary artery disease patients with inadequate collateralization.

Humans

SimpleMicrobiome: An integrated web-based platform for streamlined microbiome data analysis and visualization.

Microbiome studies require multiple analytical steps after initial sequence processing. These steps commonly include data harmonization, preprocessing, taxonomic profiling, diversity analysis, differential abundance testing, predictive modeling, network inference, and preparation of publication-ready outputs. Although robust packages are available for many of these tasks, routine use often depends on command-line workflows, repeated data reformatting, and method-specific scripting. These requirements can limit accessibility for experimental researchers and complicate consistent analysis across interdisciplinary teams. We developed SimpleMicrobiome, a web-based R Shiny platform that integrates established microbiome analysis methods into a single interactive downstream workflow. The application accepts standard abundance, taxonomy, and metadata tables, supports interactive preprocessing and sample filtering, and provides modules for taxa profile visualization, alpha and beta diversity analysis, ANCOM-BC2 and MaAsLin2 differential abundance testing, Random Forest modeling with SHAP-based interpretation, microbial association network inference using SparCC and SPIEC-EASI through NetCoMi, correlation heatmaps, and dbRDA/CAP-style association biplots. The platform is implemented as a modular Shiny application so that preprocessing choices are propagated across downstream analyses, results can be exported as figures and tables, and the same application can be run through the public server, source-code installation, or a Docker image. SimpleMicrobiome consolidates major downstream microbiome analysis tasks in an accessible browser-based environment while retaining links to established analytical frameworks. The platform may reduce technical barriers for non-programming users, improve consistency across exploratory and reporting-oriented analyses, and support collaborative microbiome research. The public application is available at https://simplemicrobiome.mglab.org, the source code is available at https://github.com/yjcho2252/SimpleMicrobiome, and a Docker image for local deployment is available at https://hub.docker.com/r/mglab2252/simplemicrobiome.

differential abundance

Identification and Classification of Expressed Orphan Genes, Spurious Orphan Genes, and Conserved Genes in the Human Gut Microbiome.

Orphan genes (OGs)-genes lacking detectable homologs outside a species-are widespread in microbial genomes and are thought to contribute to their adaptation and molecular innovation. However, not all predicted OGs may represent novel functional coding sequences. False positive OGs, also called spurious OGs, can arise from gene prediction errors. We reason that OGs lacking detectable expression are more likely to be spurious. To test this, we combined large-scale metatranscriptomic profiling of the human gut microbiome with machine learning to distinguish expressed OGs from spurious ones and compare them with conserved genes (CGs) found in multiple species. Using nearly 5,000 metatranscriptome libraries, we identified &#x223c;218,000 OGs supported by expression evidence, while &#x223c;330,000 predicted OGs lacked detectable expression and were classified as spurious. We extracted 154 features for sequence, structural, and evolutionary properties for each gene and trained XGBoost classifiers while accounting for genomic representation. The models achieved an area under the receiver operating characteristic curve (AUC) of 0.82 in distinguishing expressed OGs from spurious OGs and an AUC of 0.93 in distinguishing expressed OGs from CGs. Interpretation based on SHAP (SHapley Additive exPlanations) revealed clear biological signals. Particularly, expressed orphans were present in more genomes than spurious ones, and expressed OGs were shorter than CGs. This work improves OG discovery and suggests that expressed OGs differ systematically from CGs and spurious OGs in sequence composition, structural constraints, and evolutionary signals.

Humans

Comprehensive Evaluation and Explainable Interpretation of Peptide-HLA Binding Prediction Tools.

Accurate prediction of peptide binding to human leukocyte antigen class I (HLA-I) molecules is critical for advancing immunological research, particularly in vaccine design and immunotherapy. However, limitations in model performance, interpretability, and dataset quality impede the widespread adoption of existing predictive tools. Here, we present a comprehensive evaluation of 17 HLA-I peptide binding prediction models, utilizing a meticulously curated dataset comprising over 290,000 peptides spanning 44 HLA-I alleles. We assessed model accuracy, robustness, and interpretability, employing explainability techniques such as SHAP (SHapley Additive exPlanations) and LIME (Local Interpretable Model-agnostic Explanations) to elucidate underlying prediction mechanisms. Our results reveal substantial performance disparities, with self-attention-based models, including STMHCpan and BigMHC, exhibiting superior accuracy. Notably, the capsule network model CapsNet-MHC_AN demonstrated robust performance. Models trained on eluted ligand datasets outperformed those relying on binding affinity data, underscoring the critical role of high-quality training data. Ensemble and multi-algorithm approaches further improved prediction reliability. These findings highlight the need for ongoing innovation in model architecture, integration of diverse and high-quality datasets, and incorporation of structural predictors to develop more accurate, interpretable, and clinically applicable HLA-I peptide binding prediction tools.

HLA-I binding

Machine learning-based clinical prediction model and multi-omics integration for assessing pancreatic cancer risk in new-onset diabetes.

BACKGROUND: Given that pancreatic cancer (PC) is typically diagnosed at an advanced stage but is often preceded by new-onset diabetes mellitus (NODM), providing a window for early detection, we sought to develop and validate an interpretable machine-learning model integrated with multi-omics profiling to identify early biomarkers of NODM-associated PC. METHODS: In a population-based cohort, individuals with NODM-associated PC and NODM without PC were identified and randomly divided (70:30) into training and validation sets after feature selection. Eight machine learning (ML) classifiers were compared using fivefold cross-validation, and model performance was evaluated in terms of discrimination, calibration, and decision curve&#x2013;based clinical utility. We evaluated interpretability using the Shapley additive explanations (SHAP) analyses. Mechanistically, Olink proteomic profiling and metabolomics were analyzed through clinical classifications and model-defined risk strata. RESULTS: Categorical boosting achieved the best performance in the independent validation set (AUROC&#x2009;=&#x2009;0.844). The NODM cohort was stratified into high- (n&#x2009;=&#x2009;2,362) and low-risk (n&#x2009;=&#x2009;5,030) groups, and internal validation together with SHAP analyses demonstrated consistent model performance and identified clinically interpretable predictors. Proteomic and metabolomic analyses under clinical and risk-based grouping identified 39 overlapping differentially expressed proteins and 145 overlapping metabolites with enriched across 11 shared KEGG pathways. Cross-platform validation highlighted PLTP, CRTAC1, and ITGAV as serum biomarkers with a strong potential for early NODM-PC detection. CONCLUSIONS: We developed an interpretable ML framework centered on NODM enables practical risk stratification for early PC detection by multi-omics and provides a pathway of ML-based triage followed by biomarker confirmation for earlier detection and diagnosis.

Humans

Rapid glycomic analysis of serum EVs reveals altered N-glycosylation patterns in ASD.

Objective laboratory diagnostics for autism spectrum disorder (ASD) are lacking, necessitating rapid clinical screening tools. Because serum extracellular vesicle (EV) N-glycosylation captures critical neurodevelopmental signatures, we developed a fast, biologically interpretable diagnostic strategy. EVs from ASD patients with language impairment and neurotypical controls were isolated using a rapid extra-polyethylene glycol precipitation/filtration (EPF) workflow, benchmarked against ultracentrifugation. Following MALDI-TOF/MS profiling, machine learning was re-evaluated using repeated nested cross-validation to reduce optimistic bias and potential information leakage. Among five classifiers, Random Forest (RF) showed the best overall balance across discrimination, calibration, and classification metrics. RF-based SHAP analysis provided transparent interpretation, highlighting key discriminative glycans, including H4N3S1F1, H5N5S1F1, and H3N5F1. To elucidate molecular mechanisms, we integrated public EV transcriptomic data. This revealed significant dysregulation of N-glycosylation machinery genes (e.g., MAN1A1, NEU1, OSTC, RPN2), whose expression directionally aligned with observed glycan shifts in synaptic pathways. Collectively, this rapid serum EV N-glycomic workflow, combined with leakage-controlled RF-based interpretation, provides a promising foundation for non-invasive ASD biomarker discovery and future multicenter validation.

Humans

Linking MRI radiomics to transcriptomics-based radiosensitivity in lower-grade glioma: A radiogenomic framework.

BACKGROUND: RSI is a transcriptomics-based biomarker associated with radiotherapy outcomes, but its clinical application is constrained by the requirement for tumor tissue and RNA sequencing. This study investigates whether MRI-derived radiomic features can reflect RSI-defined intrinsic radiosensitivity in lower-grade glioma.This addresses a critical gap arising from the limited availability of matched imaging and genomic data in routine clinical practice. METHODS: MRI-derived radiomic features were extracted from FLAIR images of lower-grade glioma patients obtained from TCIA and matched with transcriptomic data from TCGA. A total of 107 patients with both MRI and RNA sequencing data were included in the radiogenomic analysis. Radiomic features were ranked using a Borda-based ensemble feature selection strategy. Five supervised machine-learning classifiers were trained to predict RSI-based radiosensitivity classification, and model interpretability was assessed using SHAP within radiogenomic framework. RESULTS: Classification performance increased with feature number and stabilized at compact subset of 13 radiomic features. Logistic regression showed stable performance with an AUC of 0.82 (95&#xa0;% CI: 0.71-0.93). SHAP analysis indicated that heterogeneity-related texture features were dominant contributors to model predictions, with many associated with the RR phenotype, while others were linked to the RS phenotype. CONCLUSION: An MRI-based radiomic signature enables non-invasive prediction of RSI-defined radiosensitivity in lower-grade glioma. Rather than offering an immediately deployable clinical tool, this study establishes a proof-of-concept radiogenomic framework demonstrating that intrinsic radiosensitivity, traditionally assessed through invasive molecular assays, can be approximated using quantitative imaging features. These findings highlight the potential of imaging-based radiosensitivity assessment and provide a foundation for future radiogenomic investigations.

Lower-grade glioma

Assessing individual genetic susceptibility to metabolic syndrome: interpretable machine learning method.

BACKGROUND: Genome-wide association studies have provided profound insights into the genetic aetiology of metabolic syndrome (MetS). However, there is a lack of machine-learning (ML)-based predictive models to assess individual genetic susceptibility to MetS. This study utilized single-nucleotide polymorphisms (SNPs) as variables and employed ML-based genetic risk score (GRS) models to predict the occurrence of MetS, bringing it closer to clinical application. METHODS: Feature selection was performed using Least Absolute Shrinkage and Selection Operator. Six ML algorithms were employed to construct GRS models. A fivefold cross-validation was utilized to aid in the internal validation of models. The receiver operating characteristic (ROC) curve was used to select the better-performing GRS model. The SHapley Additive exPlanations (SHAP) was then applied to interpret the model. After extracting GRS, stratified analysis of BMI, age and gender was performed. Finally, these conventional risk factors and GRS were integrated through multivariate logistic regression to establish a combined model. RESULTS: A total of 17 SNPs were selected for analysis. Among the GRS models, the extreme gradient boosting (XGBoost) model demonstrated superior discriminative performance (AUC = 0.837). The XGBoost's optimal robustness was also validated through five-fold cross-validation (mean ROC-AUC = 0.706). The XGBoost-based SHAP algorithm not only elucidated the global effects of 17 SNPs across all samples, but also described the interaction between SNPs, providing a visual representation of how SNPs impact the prediction of MetS in an individual. There was a strong correlation between GRS and MetS risk, particularly observed among young individuals, males and overweight individuals. Furthermore, the model combining conventional risk factors and GRS exhibited excellent discriminative performance (AUC = 0.962) and outstanding robustness (mean ROC-AUC = 0.959). CONCLUSION: This study established a reliable XGBoost-based GRS model and a GRS prediction platform (https://metabolicsyndromeapps.shinyapps.io/geneticriskscore/) to assess individual genetic susceptibility to MetS. This model has high interpretability and can provide personalized reference for determining the necessity of primary prevention measures for MetS. Additionally, there may be interactions between traditional risk factors and GRS, and the integration of both in a comprehensive model is useful in the prediction of MetS occurrence.

Humans

Machine learning-based prediction of unplanned readmission and construction of an online calculator for elderly patients with mild ischemic stroke.

OBJECTIVE: To screen for independent risk factors for unplanned readmission in elderly patients with mild ischemic stroke, and to construct and validate an online risk prediction calculator based on an interpretable machine learning model, thereby providing a promising practical tool for accurate clinical assessment of 30&#x2011;day all&#x2011;cause unplanned readmission risk in this population. METHODS: A prospective cohort study was conducted, including 1050 patients aged&#xa0;&#x2265;&#xa0;60&#xa0;years with mild ischemic stroke admitted between August 2023 and September 2024. Participants were randomly divided into a training set (840 cases) and a test set (210 cases) at a ratio of 8:2. Risk factors were screened by univariate analysis and multivariable Logistic regression. Four machine learning models, namely LightGBM, XGBoost, Random Forest, and K&#x2011;Nearest Neighbors (KNN), were developed and their performance was evaluated using AUC, accuracy, sensitivity, and specificity as metrics. The SHAP framework was used for interpretability analysis, and an online calculator was subsequently developed based on the optimal model. RESULTS: Univariate analysis showed significant differences (P&#xa0;<&#xa0;0.05) in 13 factors including age, smoking, AIP, TyG index, HALP score, etc. Multivariable Logistic regression identified age (OR&#xa0;=&#xa0;9.752), smoking (OR&#xa0;=&#xa0;5.171), AIP (OR&#xa0;=&#xa0;6.691), TyG index (OR&#xa0;=&#xa0;4.393), HALP score (OR&#xa0;=&#xa0;2.831), and&#xa0;&#x2265;&#xa0;2 comorbidities (OR&#xa0;=&#xa0;3.664) as independent risk factors. All four machine learning models demonstrated good predictive performance. Based on a comprehensive evaluation of multiple metrics and computational efficiency, the LightGBM model exhibited the best predictive performance (AUC&#xa0;=&#xa0;0.884, accuracy&#xa0;=&#xa0;0.829, sensitivity&#xa0;=&#xa0;0.812, specificity&#xa0;=&#xa0;0.875). SHAP analysis showed that age, AIP, TyG index, smoking, and HALP score were key predictors. An online calculator developed based on this model enables individualized risk predictions. CONCLUSION: Key risk factors associated with 30&#x2011;day unplanned readmission in elderly patients with mild ischemic stroke were identified. The LightGBM model demonstrated high predictive accuracy, and together with the interpretability analysis and online calculator, offers a practical tool to support clinical risk assessment. However, this tool requires future external validation.

Humans

A machine learning-based predictive model for radiosensitivity in nasopharyngeal carcinoma utilizing serum proteomics.

BACKGROUND: Nasopharyngeal carcinoma (NPC) remains highly sensitive to radiotherapy; however, radioresistance in a subset of patients leads to local recurrence and distant metastasis. Serum proteomics provides a minimally invasive approach to capturing dynamic physiological changes, and machine learning enables efficient construction of predictive models. This study aimed to develop and validate a serum proteomics&#x2013;based machine-learning model for predicting radiotherapy sensitivity in nasopharyngeal carcinoma (NPC). METHODS: Pretreatment serum samples from newly diagnosed NPC patients were analyzed using SELDI-TOF-MS. Differentially expressed proteins between radiosensitive and radioresistant groups were identified using limma. GO and KEGG analyses were performed to explore functional enrichment. Twelve machine-learning algorithms were used to construct predictive models, and the top-performing models were optimized through feature selection. A Random Forest model with seven features was identified as the optimal model. External validation was performed using an independent cohort with ELISA-quantified protein levels. Model performance was assessed using Receiver operating characteristic curve (ROC), calibration analysis, decision curve analysis (DCA), and 10-fold cross-validation. SHapley Additive exPlanations (SHAP) analysis was applied for model interpretability, and the final model was deployed via a ShinyAPP. RESULTS: A total of 96 differentially expressed proteins were identified, which involved multiple function and signaling pathways. The Random Forest model demonstrated the best predictive performance, achieving an area under the curve (AUC) of 0.963 in the training set and 0.975 in the validation set. Cross-validation yielded an average AUC of 0.965. DCA indicated high clinical utility across a broad threshold range, and calibration curves showed good model agreement. Seven proteins (PLXND1, GSR, PGD, PTPRC, OR2T29, ACTG2, CHAD) were selected as final features. SHAP analysis provided global and individual-level interpretability. A web-based tool was developed to facilitate clinical application. CONCLUSION: This study establishes a robust serum proteomics&#x2013;based machine-learning model capable of accurately predicting radiotherapy sensitivity in NPC. The model offers clinical interpretability and practical implementation, supporting personalized radiotherapy decision-making.

Humans

Plasma Proteomic Profiles Predict Individual Future Osteoarthritis Risk.

OBJECTIVE: Osteoarthritis (OA) is a widespread degenerative joint disease that causes a considerable socioeconomic burden. Despite progress in genetic and environmental insights, early diagnosis is still limited by the lack of evident symptoms during the initial phases and accurate biomarkers. This study aims to identify plasma proteins associated with future risk of OA and develop a predictive model. METHODS: We conducted a large-scale proteomic analysis of 45,307 participants from the UK Biobank, excluding those with baseline OA. Plasma samples were assayed using the Olink Explore Proximity Extension Assay targeting 1,463 unique proteins. Clinical variables and OA outcomes were extracted and linked to electronic health records. A predictive model was constructed using the LightGBM machine learning method, and SHapley Additive exPlanations (SHAP) were applied to evaluate the importance of variables. RESULTS: We identified a panel of proteins significantly associated with the risk of developing OA. Notably, after adjusting for multiple confounders, collagen type IX alpha 1 chain (COL9A1) and cartilage acidic protein 1 (CRTAC1) were the most significant predictors of incident OA, with hazard ratios of 1.54 (95% confidence interval [CI] 1.48-1.61) and 1.65 (95% CI 1.54-1.78), respectively. SHAP analysis allowed a profound interpretation of the contribution of each protein and clinical variable to the model, revealing the multifactorial nature of OA risk prediction. The temporal trajectories of plasma proteins indicated that the levels of COL9A1 and CRTAC1 began to deviate from normal for more than a decade before OA onset, suggesting their potential use in early detection strategies. The predictive model, developed using the LightGBM algorithm, integrated proteins with clinical covariates and demonstrated an area under the curve (AUC) of 0.729 for 5-year OA prediction, 0.721 for 10-year prediction, and 0.723 for all incident OA. The predictive accuracy of the model was further enhanced for hip and knee OA, achieving AUCs of 0.820 and 0.803 for 5-year predictions. CONCLUSION: Our study identified the role of plasma proteomics in predicting future OA risk, which could contribute to preemptive measures. The innovative model, which integrates proteomic biomarkers with clinical data, offers a potential tool for risk assessment, potentially optimizing OA management strategies and enhancing prevention efforts.

Humans

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal&#x2011;Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans

Multi-Omics Integration Identifies a Five-Gene Metabolic Signature With Experimental Validation in Clear Cell Renal Cell Carcinoma.

BACKGROUND: Clear cell renal cell carcinoma (ccRCC) is hallmarked by profound metabolic reprogramming; however, its intricate crosstalk with the tumor immune microenvironment (TIME) and its clinical ramifications remain inadequately elucidated. This study aims to systematically decipher the metabolic-immune interplay in ccRCC through multi-omics integration, with the goal of identifying robust prognostic biomarkers and actionable therapeutic vulnerabilities. AIMS: This study aims to systematically decipher the metabolic-immune interplay in clear cell renal cell carcinoma (ccRCC) through multi&#x2011;omics integration, and to identify robust prognostic biomarkers and actionable therapeutic vulnerabilities that can inform precision risk stratification and individualized treatment strategies. METHODS: We integrated bulk transcriptomic, genomic, and clinical data from multiple ccRCC cohorts. Differential expression and functional enrichment analyses were performed to characterize metabolic pathway alterations. Mendelian randomization (MR) was employed to infer causal relationships between metabolic disorders and ccRCC risk. A machine learning-based prognostic framework, incorporating SHAP (SHapley Additive exPlanations) for feature interpretability, was constructed and rigorously validated. TIME heterogeneity was dissected using deconvolution algorithms, while drug sensitivity, tumor mutation burden (TMB), and TIDE scores were utilized to assess therapeutic responses and immune evasion. Candidate gene function was evaluated through in&#xa0;vitro gain- and loss-of-function assays, with expression validated via TCGA, HPA, western blot, and qRT-PCR. RESULTS: Enrichment analysis identified coordinated dysregulation in lipid metabolism, energy homeostasis, and hypoxia response pathways. MR analysis confirmed lipid metabolism disorders as a causal risk factor for ccRCC. Our machine-learning model, centered on five core SHAP-identified features (SUCLA2, ACAT1, PC, SUCLG1, and HMGCS2), demonstrated superior predictive accuracy over conventional clinical staging. Immune profiling unveiled dichotomous TIME states: the low-risk group retained active immune surveillance, whereas the high-risk group was enriched with immunosuppressive subsets. Drug sensitivity screening pinpointed LY2109761 and carmustine as high-risk-specific candidate agents. Furthermore, TMB and TIDE analyses stratified high-risk patients displaying genomic instability and immune evasion phenotypes. Functionally, SUCLA2 knockdown significantly enhanced ccRCC cell proliferation and invasion, while its overexpression suppressed these malignant phenotypes, corroborating its tumor-suppressive role. Expression patterns of the hub genes were consistently validated across multi-level datasets and experimental assays. CONCLUSION: This study establishes a precision oncology framework for ccRCC by functionally linking metabolic biomarkers, immunophenotypes, and stratified therapeutic strategies. Importantly, we identify SUCLA2 as a potential functional tumor suppressor and a promising target for further mechanistic and translational investigation.

Humans

Topologically distinct intratumoral heterogeneity scores for predicting high-risk pathological grades in invasive lung adenocarcinoma: A multicenter study across four institutions.

High-risk subtypes of invasive lung adenocarcinoma (IAC), particularly micropapillary- or solid-predominant patterns, are closely associated with poor prognosis. This multicenter retrospective study developed and validated a predictive model for the preoperative identification of these high-risk subtypes using topologically distinct intratumoral heterogeneity (ITH) scores derived from CT images. The study included 1,051 patients with IAC. Two complementary ITH scores were developed: a two-dimensional ITH score, which integrated local radiomics features with global pixel distribution patterns on the largest cross-sectional CT slice, and a three-dimensional ITH score, which extended this quantification across the entire tumor volume. Clinicoradiological features and ITH scores were incorporated as model inputs to construct six base machine learning classifiers and a final stacking ensemble classifier. Model interpretability and robustness were evaluated using SHapley Additive exPlanations (SHAP)-based ablation analyses. An independent dataset from The Cancer Imaging Archive (TCIA) was used for external validation to investigate associations between ITH scores and pathological characteristics, genomic features, recurrence-free survival, and overall survival. The stacking ensemble classifier achieved the best predictive performance, with an area under the receiver operating characteristic curve of 0.875, outperforming models based solely on radiomics features (0.834) or clinicoradiological features (0.792). SHAP analysis identified the 3D ITH score as the most influential contributor to model output, and TCIA validation showed that higher 3D ITH scores were associated with more aggressive tumor biology and poorer survival outcomes. The topologically distinct 3D ITH score may provide a clinically meaningful imaging biomarker for preoperative risk stratification in IAC.

Journal Article