PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Machine learning-enabled multi-omics discovery of prognostic biomarkers and signaling targets in pancreatic cancer.

Pancreatic ductal adenocarcinoma (PDAC) remains difficult to subtype using single omics layers. We conducted an exploratory investigation integrating reverse-phase protein array (RPPA) and DNA methylation data from the cancer genome atlas (TCGA)- pancreatic adenocarcinoma (PAAD) to assess the feasibility of multi-omics subtyping, alongside a supervised machine learning analysis of a small gene expression omnibus (GEO) transcriptomic cohort (n = 26) to identify candidate diagnostic genes. RPPA-based K-means clustering suggested a weak, possible two-subtype structure (silhouette ≈ 0.16) that remained unassociated with overall survival (log-rank p = 0.113) and lacked independent prognostic value. An independently performed similarity network fusion (SNF) analysis integrating RPPA and methylation data showed low concordance with RPPA-derived subtypes (Adjusted Rand Index (ARI) = 0.014), indicating limited convergence between molecular modalities. Supervised machine learning analysis of the GEO cohort using a fully nested leave-one-out cross-validation pipeline achieved a mean (area under the curve) AUC of 0.896 across four classifiers and identified four-fold-stable candidate genes (ESCO2, COL17A1, BCL2L14, and SOWAHB). However, this gene panel demonstrated limited external validity across two independent PDAC cohorts (log-rank p = 0.438 for both GSE62452 and GSE28735), indicating limited generalizability despite robust internal performance. Collectively, these findings provide limited evidence for a robust, prognostically significant multi-omics subtype or a validated diagnostic gene signature; instead, this study serves as a hypothesis-generating resource and highlights the importance of rigorous cross-validation and independent external validation in small-sample transcriptomic biomarker discovery.

Humans

Screening of core targets for Di(2-ethylhexyl) Phthalate-related gastric cancer based on machine learning, molecular docking, and SHAP analysis.

PURPOSE: Given the existing uncertainties regarding the link between Di(2-ethylhexyl) phthalate (DEHP) exposure and gastric cancer (GC) progression, this study aimed to clarify their association, identify the toxic targets of DEHP, and elucidate the underlying molecular mechanisms. METHODS: Multiple integrated approaches were employed, including Gene Expression Omnibus (GEO) data analysis, network toxicology, molecular docking, and machine learning. STRING and Cytoscape tools were utilized to identify key targets, while Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analyses were performed to explore the functional enrichment of intersecting targets. Machine learning and SHAP analysis were applied to screen core targets in GC. Molecular docking was performed to evaluate the binding affinity of DEHP toward core targets, and 200 ns molecular dynamics simulations were further conducted for representative complexes to validate their dynamic stability. RESULTS: A total of 18 key targets were identified using STRING and Cytoscape. GO and KEGG enrichment analyses demonstrated that these intersecting targets were primarily enriched in the extracellular region, as well as the Calcium signaling pathway and cAMP signaling pathway. Through machine learning analyses, 7 key genes (ADRB2, ESRRG, GRIA4, IL13RA2, NR3C2, PLA2G1B, and SULT2A1) were identified as core targets in GC through machine learning analyses. Molecular docking simulations revealed strong binding specificity between DEHP and the target proteins. Among them, NR3C2 and ADRB2 exhibited relatively high predictive importance in the machine learning models. DEHP showed favorable binding affinity toward these core targets, and molecular dynamics simulations further confirmed that ADRB2-DEHP and NR3C2-DEHP complexes maintained stable conformations throughout the simulation. CONCLUSIONS: Our findings identified GC associated genes that were computationally predicted as potential targets of DEHP. These results indicated structural compatibility between DEHP and its target proteins but did not prove that DEHP exposure accounts for the gene expression changes in GC.

Molecular Docking Simulation

AI-driven diagnostic and prognostic models for metabolic dysfunction-associated steatotic liver disease: insights from clinical, imaging, and multi-omics studies-a scoping review.

Metabolic dysfunction-associated steatotic liver disease (MASLD), formerly known as non-alcoholic fatty liver disease (NAFLD), is the most common chronic liver disease around the world, affecting 33.6% of the adult population (95% CI: 28.1%-39.5%; I 2 = 99.9%), or roughly one in three. The extent of the liver damage is variable, from simple steatosis to metabolic dysfunction-associated steatohepatitis (MASH, formerly NASH), cirrhosis and hepatocellular carcinoma (HCC). Early diagnosis is essential to prevent serious liver damage. Traditional diagnostic techniques such as liver biopsy, imaging, and biomarker testing are all invasive, costly, reduced sensitive to early-stage disease, and they also have variability among observers. Modern diagnostic and prognostic approaches based on the principles of Artificial Intelligence (AI) and specifically on machine learning (ML) and deep learning (DL) have enabled multimodal approaches integrating clinical, imaging and molecular data. This scoping review conducted per PRISMA-ScR guidelines, synthesizes findings from 73 studies (search window 2020-2026) across three dimensions: clinical data driven models, imaging-based classifiers (ultrasound, CT and MRI), and multi-omics (genomics, transcriptomics and proteomics) techniques. Moreover, emergence of models such as U-Net and LiverNet 2.x, classification models like DeepLiverNet and BiLSTM models, as well as transformer frameworks and the identification of biomarkers models are also described. This study also investigates challenges such as data heterogeneity, data interpretability, fairness and real-world clinical application. Finally, important areas of research opportunities and future directions are highlighted to present a developing clinically applicable, explainable and ethical AI solutions to manage MASLD.

MASLD

A machine learning-derived and functionally validated circadian rhythm signature predicts clinical outcomes and in silico drug sensitivity in colorectal cancer.

BACKGROUND: Colorectal cancer (CRC) displays considerable heterogeneity in clinical outcomes, highlighting the need for reliable prognostic biomarkers. While the aberrant expression of circadian rhythm-related genes has been implicated in cancer pathogenesis, its comprehensive role in CRC progression and predicted therapeutic vulnerabilities remains inadequately characterized. METHODS: Bulk and single-cell RNA-sequencing data were integrated from multiple CRC cohorts. A circadian rhythm signature (CRS) was developed through machine learning algorithms and validated for prognostic value. Comprehensive analyses of tumor microenvironment, genomic alterations, and drug sensitivity were performed. Furthermore, the biological function of the core gene, BHLHE40, was validated in CRC cell lines through CCK-8, EdU, and wound healing assays. RESULTS: Single-cell analysis demonstrated an elevated expression signature of circadian rhythm-related genes in dendritic cells. The optimized CRS, comprising 14 circadian rhythm-related genes, successfully categorized patients into high- and low-risk groups. Patients with a high CRS showed markedly poorer overall survival and computationally inferred immunosuppressive features, including reduced CD8+ T cell infiltration and increased M2 macrophage polarization. Genomic analysis revealed enhanced mutation burden in TP53 and alterations in RTK-RAS/WNT pathways. Notably, in vitro assays confirmed that BHLHE40 is significantly overexpressed in CRC cells. Knockdown of BHLHE40 markedly inhibited tumor cell proliferation and migration. Drug sensitivity profiling identified bexarotene and SMER-3 as potential therapeutic options for high-CRS patients. A nomogram integrating CRS with clinical parameters demonstrated superior predictive accuracy for 1-, 3-, and 5-year survival. CONCLUSIONS: The CRS represents a promising prognostic biomarker that reflects tumor immune status and genomic features, providing valuable insights for personalized treatment strategies in CRC.

Circadian rhythm

Proteomic profiling of bone for the estimation of post-mortem interval and post-mortem submersion interval: a systematic review.

Accurate estimation of the Post-Mortem Interval (PMI) and Post-Mortem Submersion Interval (PMSI) remains a persistent challenge in forensic science, especially when traditional morphological and entomological methods fail due to advanced decomposition or in aquatic environments. Proteomic profiling of bone tissues has recently emerged as a promising approach, leveraging the predictable degradation patterns of bone proteins to estimate time since death more reliably. This systematic review, conducted in accordance with PRISMA guidelines, analyzed 24 peer-reviewed studies focusing on the application of proteomic techniques to bone tissue for PMI and PMSI estimation. The included studies were evaluated based on sample type, analytical techniques used, identified biomarkers, environmental conditions assessed, and the overall reliability and reproducibility of the findings. The review found that specific bone proteins, particularly collagen, osteocalcin, fetuin-A, etc. exhibited consistent degradation patterns that correlated strongly with elapsed post-mortem time. Cortical bone was identified as a more stable and informative matrix compared to trabecular bone. Mass spectrometry, especially LC-MS/MS, emerged as the predominant analytical technique due to its high sensitivity and accuracy in detecting low-abundance proteins over extended PMIs and PMSIs. However, protein degradation rates were significantly influenced by environmental variables such as temperature, humidity, soil pH, and microbial activity. This review also emphasizes the transformative role of bone proteomics in advancing forensic science while identifying key gaps that must be addressed to achieve global standardization and practical implementation in diverse forensic contexts. The integration of proteomics with other emerging technologies, such as machine learning algorithms and computational modeling, may further enhance the precision of PMI and PMSI estimation in future applications.

Postmortem Changes

Multi-level Transcriptomic and Machine-learning Analyses Identify MZT1 as a Proliferation-associated Prognostic Marker in Lung Adenocarcinoma.

BACKGROUND/AIM: Lung adenocarcinoma (LUAD) exhibits substantial molecular heterogeneity and variable clinical outcomes, highlighting the need for biomarkers that reflect core tumor biological processes. Centrosome-associated proteins regulate mitotic fidelity and genome stability, yet their roles in LUAD remain incompletely defined. In this study, we systematically characterized mitotic spindle organizing protein 1 (MOZART1; MZT1) and related family members in LUAD. MATERIALS AND METHODS: We performed integrated analyses combining bulk transcriptomic datasets, survival modeling, gene set enrichment, immune deconvolution, machine-learning based prognostic modeling, and single-cell RNA sequencing. Expression patterns and clinical associations of MZT family genes were evaluated across pan-cancer and LUAD cohorts. RESULTS: MZT family genes were consistently upregulated in tumor tissues, with MZT1 showing the most robust expression pattern. Elevated MZT1 expression was significantly associated with reduced overall survival. Functional analyses revealed coordinated activation of proliferative and genome maintenance pathways, including G2/M checkpoint regulation, E2F and MYC signaling, and DNA repair. A multivariable analysis indicated that the prognostic association of MZT1 was reduced after adjusting for canonical proliferation markers, suggesting partial overlap with established proliferation signals. The LASSO-based Cox model demonstrated stable time-dependent predictive performance at 1-, 3-, and 5-year survival. Immune analyses indicated associations between MZT1 expression and tumor microenvironmental features. Single-cell analysis showed that MZT1 expression was predominantly enriched in malignant epithelial cells and associated with proliferative cellular states. Protein-level validation supported concordance with transcriptomic findings. CONCLUSION: MZT1 is a proliferation-associated marker that integrates clinical risk, transcriptional programs, cellular heterogeneity, and predictive modeling in LUAD, providing a potential framework for biomarker development and risk stratification.

Humans

STRUMP-I: Structure-based machine learning approach to pMHC-I binding prediction using force field energy features.

The adaptive immune system monitors cellular integrity by recognizing short peptides from intracellular proteins presented on Major Histocompatibility Complex class I (MHC-I) molecules, collectively termed peptide-MHC complexes (pMHC), enabling detection of foreign or mutated proteins. With the rising importance of immunotherapies targeting neoantigens in cancers, the ability to accurately predict which peptides will bind to the diverse population of MHC alleles is critically important. Current computational methods for pMHC-I prediction fall broadly into sequence-based methods, which rely heavily on large training datasets, and structure-based methods that leverage structural modeling and energetics of pMHC binding. While sequence-based methods have been popularly used, their performance is dependent on the size and quality of training data. On the other hands, while structure-based approaches can generalize better across diverse MHC alleles, they traditionally depend on identifying a single global minimum energy conformation, an assumption that often fails due to the inherent binding promiscuity of MHC-I molecules. To address these limitations, we developed a STRUMP-I (STRUcture-based pMHC Prediction (for class I)), a novel pMHC binding prediction tool that directly leverages a broad set of force-field-derived energy terms as machine-learning features. STRUMP-I achieves performance comparable to state-of-the-art sequence-based models while significantly outperforming them on MHC alleles with limited representation in training data. Furthermore, STRUMP-I demonstrates strong synergy when integrated with sequence-based methods, notably enhancing prediction precision. The robustness and generalizability of STRUMP-I were confirmed by evaluating its predictive performance on independent, previously unseen datasets, including an experimentally validated cancer neoantigen dataset. This combined approach advances our capability to reliably identify clinically relevant neoantigen targets. The source code and trained models are available at https://github.com/yoonjoolab/STRUMP-I.

energy optimization

Beyond data and technology: the need for new thinking to enable the era of precision prevention.

BACKGROUND: Global flagship initiatives increasingly advocate for proactive health maintenance to alleviate the growing burden on reactive, disease-focused healthcare systems. Precision prevention is conceived as the targeted modulation of causal pathways across the disease continuum, from latent risk and pre-disease states to clinical manifestation, surpassing conventional public health prevention strategies that prioritise managing population-level risk factors. Traditional discovery and implementation models, however, remain poorly aligned with the pace and breadth of scientific and technological advances. This review outlines key barriers to scaling precision prevention and argues for the integration of conceptual, methodological, and policy perspectives into a single implementation‑oriented framework. MAIN: Individualised risk stratification lies at the core of precision prevention. Genomics serves as a stable substrate for lifetime susceptibility assessment, while meaningful prediction in multifactorial chronic disease requires additional risk monitoring using dynamic intermediate molecular markers and high-resolution exposomic data. Machine learning and other artificial intelligence (AI) methods are increasingly helpful tools for integrating large, heterogeneous and temporally structured real-world data to generate personalised predictions of health trajectories. Trustworthy AI-enabled risk prediction or decision-support systems are expected to provide transparency about model logic, assumptions and performance. In discovery, existing diagnostic classifications and conventional case-control designs can obscure mechanistic heterogeneity. Shifting toward precision phenotyping and biologically grounded disease redefinition could reveal a new layer of molecular understanding. Evidence generation strategies that reflect the temporal change of disease, including high‑risk enrichment, surrogate endpoints, and adaptive, trajectory-based monitoring, are particularly important for common conditions with prolonged latency periods (e.g., cancer, cardiovascular disease). Features often dismissed as "noise", such as stochastic molecular variation and minimal exposures, may in fact encode meaningful individual-level signals and thus merit investigation. CONCLUSION: To shift healthcare from reactive treatment toward proactive health maintenance requires coordinated action from stakeholders to reshape the pillars of discovery, reform outcome assessments and modernise implementation strategies.

Humans

Exposome influences: a multi-omics perspective on the combined toxic effects of pharmaceuticals and personal care products in Alzheimer's disease.

According to WHO data, approximately 57 million people worldwide were affected by dementia in 2021, with prevalence projected to rise. Alzheimer's disease (AD), responsible for 60%-80% of dementia cases, continues to be a leading cause of mortality, with current treatments offering limited efficacy and disease-modifying therapies lacking widespread adoption or conclusive safety evidence, shifting the focus toward prevention and risk modification. Risk factors for AD include both non-modifiable elements, such as age, genetics, and gender, and modifiable factors, like environmental pollution, health status, and diet. While age remains the primary non-modifiable risk factor, early-onset dementia represents only up to 9% of cases. Addressing modifiable factors is essential, as it could prevent or delay almost half of dementia cases, with interventions-such as increased physical activity, smoking cessation, alcohol limitation, and overall health management-being significantly associated with a reduced risk. In this context, the exposome approach offers a comprehensive, integrative framework in which both modifiable and non-modifiable risk factors interact to influence individual susceptibility. Within the neural exposome, chronic low-dose exposure to xenobiotics-such as industrial chemicals, pesticides, metals, pharmaceuticals and personal care products (PPCPs), and air pollutants-may induce neurodegeneration via mechanisms including oxidative stress, neuroinflammation, proteinopathies, and epigenetic modifications, although establishing causality remains challenging. Integration of genomics, transcriptomics, proteomics, metabolomics, and lipidomics, combined with artificial intelligence (AI) techniques such as machine learning (ML) and deep learning (DL), provides promising avenues for biomarker discovery, enhanced preventive strategies, early non-invasive diagnosis, and therapeutic target identification by integrating multi-layered biological data with exposure profiles. This review highlights emerging AD risk factors-including PPCPs-underscoring complex, multifactorial nature of AD and exposome, and the requirement for an interdisciplinary research approach, while also addressing several critical research gaps and methodological limitations.

Alzheimer’s disease

Multimodal features and prognostic risk assessment in locally advanced gastric cancer patients following neoadjuvant therapy based on machine learning algorithms: a multicenter study.

BACKGROUND: Neoadjuvant therapy (NAT) is recommended for locally advanced gastric cancer (LAGC), but some patients respond poorly. We aimed to construct a multimodal model integrating CT images, transcriptomic sequencing, and clinicopathological data to assess prognosis in LAGC patients receiving NAT. MATERIALS AND METHODS: This multicenter study included 505 LAGC patients who underwent NAT. Radiomic features were extracted from preoperative CT images of 505 patients. RNA-seq was performed on 277 post-NAT specimens, with additional data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases (n&#x2009;=&#x2009;804). Patients were divided into training (168 cases), internal validation (72 cases), and external validation cohorts. Machine learning algorithms identified key radiomic, molecular, and clinical features associated with NAT response, which were then integrated into a multimodal model to predict overall survival (OS) and disease-free survival (DFS). RESULTS: Six radiomic and three molecular features significantly associated with NAT response were selected. Radiomic risk (hazard ratio [HR]: 4.0, P&#x2009;<&#x2009;0.001) and molecular risk (HR: 7.1, P&#x2009;<&#x2009;0.001) were independent prognostic factors. By integrating radiomic risk, molecular risk, and clinical characteristics, a multimodal model (MuMo) was constructed.The C-index results (OS, C-index&#x2009;=&#x2009;0.855; DFS, C-index&#x2009;=&#x2009;0.786) demonstrated that MuMo outperformed the single-modality models and ypTNM staging.Mechanistic analysis suggested that the efficacy of neoadjuvant therapy was significantly enriched in immune-inflammatory pathways. CONCLUSIONS: MuMo can effectively predict postoperative survival risk in LAGC patients receiving NAT, serving as a powerful tool for optimizing prognostic assessment.

Humans

Spatial Multiomics Reveal Insights Into ADC Efficacy.

Antibody-drug conjugates (ADCs) have transformed the therapeutic landscape of solid tumors; however, responses remain heterogeneous and complex to predict. In addition, a growing number of multiple ADC targets are either approved or in late-stage clinical development, such as NECTIN-4, HER2, or TROP2 for metastatic urothelial cancer. Spatial multiomics-representing next-generation methods that couple high-plex RNA sequencing and multiplex protein imaging with precise x-y-z coordinates within tissues-offer a direct way to correlate (ADC) antigen expression, cell state information, and micro-anatomical context with patient treatment outcomes. In this review, we highlight suitability and technological advancements in current spatial transcriptomics and proteomics approaches to decode modes of action and resistance to ADCs and extract biological insights, particularly in metastatic urothelial cancer-and propose an integrative framework that combines spatial readouts with machine and/or deep learning-driven analytics to stratify patients, forecast on- and off-target toxicities, and guide next-generation linker-payload designs or combination therapies.

Humans

Predicting enhancer-promoter interactions using a stacking-based ensemble strategy.

MOTIVATION: Enhancer-promoter interactions (EPIs) are essential for gene regulation and disease progression. Recent studies have shown that distal enhancers can regulate target genes through interactions with nearby promoters, providing important insights into transcriptional regulation mechanisms. Although high-throughput experimental techniques have enabled large-scale identification of EPIs, these methods are often costly and time-consuming. In addition, existing computational approaches still face challenges in effectively integrating heterogeneous feature representations from different cell lines. RESULTS: We propose a stacked ensemble framework for EPI prediction that integrates feature representations from diverse cell line datasets using multiple machine learning algorithms. The extracted complementary patterns are further combined by an XGBoost classifier to improve robustness against overfitting. Experiments on six independent datasets show that the proposed method achieves superior accuracy and generalization compared with existing EPI prediction models, with an average AUROC of 0.909 while maintaining computational efficiency. AVAILABILITY: The source code and its archived release are available at GitHub and Zenodo. The Zenodo archive provides a versioned snapshot of the repository: https://zenodo.org/records/19952998.

Promoter Regions, Genetic

AI-Driven Precision Medicine in Alzheimer's Disease: Drug Repurposing, Digital Therapeutics and Clinical Decision Support.

Alzheimer's Disease (AD) is a neurodegenerative disease that causes significant clinical, social, and economic burden worldwide. Despite improvements in understanding its multifaceted pathogenesis, current treatments are mostly symptomatic and ineffective across varied patient populations. To overcome these constraints, AI-driven precision medicine allows tailored risk assessment, treatment selection, and disease monitoring. This review covers AI's role in AD precision medicine, focusing on drug repurposing, digital therapies and clinical decision support systems. Machine and deep learning models are used to predict medication response, integrate heterogeneous data sources such as genomics, transcriptomics, neuroimaging and electronic health records, and uncover pharmacogenomic treatment success factors. The paper covers AIenabled precision pharmacology, including tailored dosing algorithms, adaptive therapeutic monitoring, and adverse drug reaction prediction. Bioinformatics-based target identification, network pharmacology, graphbased AI models, virtual screening, and real-world and clinical data validation are emphasized in AI-driven medication repurposing. AI-powered digital treatments like personalized cognitive training platforms, wearable- derived digital biomarkers, virtual and mixed reality interventions, adherence monitoring, and digital twins for therapy optimization have been discussed. AI-based clinical decision support systems are also thoroughly assessed for clinical value, accuracy, and explainability in disease subtyping, trajectory prediction, and risk stratification in preclinical and prodromal AD. Despite these promises, data heterogeneity, algorithmic bias, legal barriers, and privacy concerns exist. Federated learning enables safe multi-center collaboration and hybrid AI-human approaches, and it represents the future. AI's ability to alter AD care opens the door to precision medicine paradigms that use repurposed medications, digital tools and intelligent decision-making to improve patient outcomes.

Alzheimer&#x2019;s disease

Unlocking the Circulating Proteome: Toward Clinical Translation.

Blood-based proteomics is approaching a translational inflection point. Driven by advances in measurement technologies, rapid expansion of analytical capabilities, and growing adoption across research and medical communities, there is increasing demand for clinically actionable biomarkers. As the field transitions away from purely large-scale discovery-oriented studies toward more informed, targeted, application-driven analyses, the generation of proteomic data is no longer the bottleneck. Instead, the central challenge is to translate these measurements into robust, reproducible, and clinically meaningful insights. In this Review, we assess recent technological and methodological developments, evaluate persistent preanalytical and interpretative limitations, and outline the key steps required for clinical translation. We focus on three deeply interconnected dimensions: the capabilities and constraints of current measurement platforms, the role of computational and machine learning approaches in extracting biological and clinical signals, and the emergence of large-scale population studies that create new opportunities for validation and generalization. Finally, we discuss a forward-looking vision in which proteomics plays a central role in dynamic, multilayered omics frameworks, where integration with genomics, temporal profiling, and imaging can deepen our understanding of health, disease, and therapeutic response.

Humans

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)&#x2500;a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC &#x2265; 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median &#x3c1; &#x223c; 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals

Pan-cancer multi-omics machine learning defines a lactylation-associated immune-excluded tumor state with proteomic and experimental corroboration.

BACKGROUND: Histone lactylation links lactate metabolism to chromatin regulation, but whether lactylation-program-associated transcriptional patterns delineate recurrent pan-cancer tumor states remains unclear. METHODS: We integrated mRNA, lncRNA, and miRNA profiles from 9712 TCGA tumors across 33 cancer types with GTEx references, six GEO cohorts, IMvigor210, and an institutional clear-cell renal cell carcinoma (ccRCC) cohort used for exploratory DIA-NN proteomic corroboration. Random-effects co-expression meta-analysis, multi-omics consensus clustering, regulon inference, immune deconvolution, TIDE, oncoPredict, and SHAP-based machine learning were applied. hsa-miR-431-5p was functionally evaluated as a proof-of-concept CS2-associated miRNA in bladder cancer models. RESULTS: LacCoEx-Atlas comprised 398,491 lactylation-related co-expression pairs across 24,667 RNA features under a random-effects framework (median I&#xb2; = 88.6%). Consensus clustering identified two subtypes: CS2 showed glycolytic-mesenchymal-immune-excluded features, M2 macrophage enrichment, CD8&#x207a; T-cell depletion, elevated HDAC4/NSD3/KDM6B activity, and worse survival, whereas CS1 showed oxidative, sirtuin-active programs. CS2 had fewer predicted ICI responders (18.3% vs. 52.0%) and a lower observed ORR in IMvigor210 (15.3% vs. 24.0%). oncoPredict identified NU7441 as a hypothesis-generating CS2-associated sensitivity signal (Hedges' g = 1.17). DIA-NN proteomics in 50 ccRCC specimens provided exploratory support for CS2-associated hypoxia, ECM degradation, and metastasis programs. The 10-feature mRNA LARItools model achieved an apparent AUC of 0.9413, while a separate multi-omics model achieved 0.971; neither was independently validated. LARItools reproduced prognostic separation across six GEO cohorts. miR-431-5p promoted malignant phenotypes and EMT in bladder cancer cells, with concordant CMU4h expression findings. CONCLUSIONS: Lactylation-program-associated transcriptional patterns delineate a recurrent immune-excluded pan-cancer tumor state associated with adverse prognosis, reduced predicted immunotherapy responsiveness, exploratory single-cancer protein-level support, and testable DNA damage response-targeting hypotheses. LacCoEx-Atlas and LARItools provide open resources for lactylation-program-associated tumor-state stratification and future translational research.

Humans

Radiogenomics predicts immune microenvironment heterogeneity and response to combination immunotherapy in hepatocellular carcinoma.

BACKGROUND: The combination of immune checkpoint inhibitors (ICIs) with anti-angiogenic agents is the preferred first-line therapy option for patients with advanced hepatocellular carcinoma (HCC), yet only a subset of patients responds, urging the quest for prediction biomarkers. We aimed to integrate genomics with radiology to propose an immune-derived radiogenomics biomarker of response to such combination immunotherapy and evaluate its added value in clinical context. METHODS: We integrated bulk RNA sequencing (RNA-seq) and proteomics data of 994 HCC patients with single-cell RNA-seq data of 11 samples across multiple datasets to identify an immune-related signature (IRS) that may influence sensitivity or resistance to such combined immunotherapy strategy, followed by verification of selected marker genes using immunohistochemistry and cytological experiments. We then trained/validated a cross-modality radiogenomics biomarker using machine learning based on TCIA database that was further tested in multi-scale independent cohorts covering 754 HCC patients. RESULTS: Integrative multi-omics analysis identifed a parsimonious 2-gene prognostic signature including KPNA2 and SMG5 that was significantly associated with immune heterogeneity and response to combination immunotherapy. Machine-learning pipeline exported the optimal 4-feature radiogenomics biomarker using support vector machine that significantly discriminated prognosis (hazard ratio 1.415&#x2013;1.890; p&#x2009;<&#x2009;0.05 for all) and modestly predicted response to ICI plus anti-angiogenic therapy (area under the curve 0.720&#x2013;0.829) in independent retrospective series across major imaging modalities (computed tomography/magnetic resonance imaging). In a prospective neoadjuvant cohort, this biomarker also showed favorable performance for predicting pathological response and tumor recurrence, accompanied by biological validation through single-cell RNA-seq analysis of pre-treatment biopsies. CONCLUSIONS: Our study provides a cross-device-cross-modal radiogenomics biomarker that can improve patient selection for emerging ICI plus anti-angiogenic therapy with novel potential therapeutic targets in HCC.

Humans

Integrative multi-omics analyses suggest a candidate microbial metabolite-associated host gene network in ulcerative colitis.

Ulcerative colitis (UC) is associated with gut microbial dysbiosis, but the host molecular alterations potentially linked to microbially derived metabolites remain incompletely understood. We integrated Mendelian randomization (MR), microbial metabolite annotation, computational target prediction, colonic transcriptomics, network analysis, and machine learning. MiBioGen microbiome GWAS data were used as exposures and FinnGen Release 12 ULCERENTER as the outcome. Metabolites linked to MR-prioritized taxa were retrieved from GutMGene, and human targets were predicted using SwissTargetPrediction and SEA. UC-related genes were defined by integrating differential expression analysis and WGCNA and then intersected with predicted metabolite targets. MR prioritized one family and eight genera showing nominal genetically supported associations with UC, but none remained significant after Benjamini-Hochberg FDR correction. Three prioritized genera were linked to 15 microbe-metabolite records, corresponding to 13 unique metabolites; nine were retained for target prediction, yielding 277 unique predicted human targets. Transcriptomic analysis identified 1,530 DEGs and a 312-gene MEgrey60 module, with 273 overlapping genes, producing 1,569 unique UC-related genes. Their intersection with the 277 predicted targets yielded 47 candidate genes. Enrichment analyses highlighted mainly metabolic and lipid-related processes. Random Forest showed the highest mean AUC across the two independent external benchmarking cohorts, and SHAP prioritized EPHX1, HSD17B2, IGFBP5, and MMP10. IBDome analysis showed inflammation-associated expression differences in these genes. This study provides a genomics-informed, hypothesis-generating framework that prioritizes candidate microbe-metabolite-host relationships in UC for future experimental validation.

Humans