PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans

Integrating machine learning and GWAS for variant prioritization in the INCIPE cohort highlights ABC transporter genes in chronic kidney disease.

INTRODUCTION: Chronic kidney disease (CKD) is a major public health challenge, affecting approximately 674 million people worldwide and representing one of the fastest-growing causes of mortality. Since CKD is frequently asymptomatic in its early stages, the identification of novel genetic biomarkers may improve early detection and risk stratification. Genome-Wide Association Studies (GWAS) have identified numerous genetic loci associated with CKD and related traits; however, their performance is often limited in small and imbalanced cohorts, where reduced statistical power increases both false-positive and false-negative findings. Machine learning (ML) approaches can complement conventional GWAS by prioritizing biologically relevant genetic signals from high-dimensional genomic data. METHODS: In this study, we implemented a nested ensemble (NCBC) model composed of an undersampler and a CatBoostClassifier (CBC) to prioritize candidate genetic variants associated with CKD in the INCIPE cohort. Prioritized variants were functionally annotated and evaluated through enrichment analyses, GTEx gene expression profiling, and protein-protein interaction network analyses. Genes identified by the CKDGen Consortium were analysed as an external reference set and used to validate the biological relevance of the prioritized results. RESULTS: The NCBC model outperformed conventional ML classifiers, achieving a ROC AUC score of 87.77%, compared to 50%-53% for the other evaluated models. Among the prioritized genes, 56.25% showed protein-protein interactions with genes previously reported by the CKDGen Consortium, whereas only 1.9% of randomly generated gene sets showed interactions. DISCUSSION: Our study demonstrates that the NCBC model improves the prioritization of biologically plausible candidate variants in a small and imbalanced CKD cohort. Functional analyses suggested ABC transporter-related genes, including ABCA13, ABCA4, and ABCC4 genes, as promising candidate for future validation, with ABCA4 showing substantial expression in kidney tissues. Overall, these findings support the integration of ML with GWAS to prioritize candidate genes and investigate the genetic architecture of complex diseases.

SNP prioritization

Integrative machine learning and transcriptomic analysis reveals molecular mechanisms underlying low survival rate in larval Chinese Bahaba (Bahaba taipingensis).

Chinese Bahaba (Bahaba taipingensis) is a Class I protected marine fish endemic to China. Low larvae survival during artificial breeding severely hinder population recovery. To investigate the molecular mechanism of high mortality in larval fish, this study performed RNA-seq on liver from naturally deceased (ND) and mass-dead (MD) individuals, combined with least absolute shrinkage and selection operator (LASSO) regression and random forest (RF) algorithms to screen for core signature genes. A total of 873 differentially expressed genes (DEGs) were identified, including 112 upregulated and 761 downregulated genes. GO and KEGG enrichment analyses revealed significant enrichment in amino acid metabolism disorders, one‑carbon folate pool impairment, PPAR signaling abnormalities, ECM-receptor interaction, focal adhesion pathway, indicating widespread metabolic suppression accompanied by extracellular matrix remodeling and signaling disturbances in the livers of MD fish. MAD pre-filtering combined with dual machine learning algorithms yielded 18 robust core signature genes, among which SLC38A4, MMP1, FADD, FKBP5, and APOB were consistently identified as high-frequency core genes by both algorithms. SLC38A4 exhibited the highest importance score in the RF model and was significantly downregulated, making it the primary molecule distinguishing ND from MD phenotypes. ROC curve analysis showed that both models achieved an AUC of 1.000 (95% CI lower bound: 0.610), confirming the precise discriminatory ability of the core genes. GSEA further demonstrated significant enrichment of this core gene set in ND samples. This study provides the first systematic elucidation of the molecular mechanisms underlying liver dysfunction in low survival rate B. taipingensis, characterized by amino acid transport impairment, metabolic reprogramming, and structural remodeling, offering theoretical foundations for health assessment, early mortality risk warning, and artificial breeding conservation of this species.

Animals

Integrative machine learning models to unravel gut microbial dysbiosis and functional disruption in polycystic ovary syndrome.

OBJECTIVE: To study gut microbial diversity and metabolic pathway disruptions in women with PolyCystic Ovary Syndrome (PCOS) compared with healthy controls, and to evaluate the diagnostic potential of microbiome-driven machine learning models. DESIGN: Case-controlled metagenomic data analysis SUBJECTS: Gut metagenomic data from women diagnosed with PCOS and age-matched healthy female controls EXPOSURE: Presence of PCOS MAIN OUTCOME MEASURES: The primary outcome measures will include gut microbial alpha and beta diversity indices, microbial taxon abundance, functional pathway profiles, predicted metabolite levels, microbe-functional pathway-metabolite interaction networks, and the diagnostic accuracy of microbiome-based machine learning models. RESULTS: Alpha and beta diversity analyses revealed marked gut microbial dysbiosis in women with PCOS, despite comparable species richness to healthy controls. Differential abundance analysis identified 41 significantly altered microbial species, including enrichment of proinflammatory taxa, such as Bacteroides vulgatus and Ruminococcus gnavus, and depletion of beneficial commensals, including Roseburia hominis and Prevotella copri. These compositional shifts indicate a proinflammatory microbial community structure in PCOS. Functional profiling demonstrated the upregulation of pathways involved in nucleotide turnover, lipid and carbohydrate metabolism, and neurotransmitter synthesis, potentially contributing to metabolic and neuroendocrine disruption. Network analysis revealed fragmented and unstable microbial-metabolite associations in PCOS compared with cohesive networks in controls. Microbiome-based machine learning models achieved a diagnostic accuracy of 84.25% (area under the curve 0.93), underscoring their predictive potential. CONCLUSION: The gut microbiome in PCOS is characterized by a proinflammatory community structure and disrupted metabolic pathways. These findings demonstrate the diagnostic potential of microbiome-based models and underscore the gut microbiome as a promising target for therapeutic interventions in the management of PCOS.

Polycystic Ovary Syndrome

Machine learning-integrated multi-omics risk prediction for pulmonary fungal infection in COPD and lung cancer: a transcriptomic and immune profiling study.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) and lung cancer are major risk factors for invasive pulmonary fungal infection (IPFI), carrying an attributable mortality of 30%-80%. Their coexistence further amplifies immunosuppression, while current diagnostic criteria remain inadequate for early risk identification. METHODS: Transcriptomic data from the GEO dataset GSE296912 (scRNA-seq; 12,078 cells from normal and COPD lung tissue) and The Cancer Genome Atlas (TCGA)-lung adenocarcinoma (LUAD) bulk RNA-seq cohort (539 tumor and 59 normal samples) underwent differential expression and cross-omics integration analysis. Five machine learning models were constructed: logistic regression, SVM, random forest, XGBoost, and LASSO. Candidate genes were validated by qRT-PCR in A549 cells and THP-1-derived macrophages stimulated with heat-inactivated Aspergillus fumigatus conidia, a protocol selected to ensure BSL-2 biosafety compliance and isolate PAMP-mediated innate immune signaling. Model performance was evaluated using 5-fold stratified cross-validation with AUC, calibration curves, and decision curve analysis. RESULTS: Single-cell transcriptomic analysis of 12,078 cells identified 14 distinct cell populations, with marked myeloid expansion and immune dysregulation in COPD lung tissue. Cross-omics integration with TCGA-LUAD data identified 1,145 shared genes (79 immune-related), converging on NF-κB, TLR4, and cytokine receptor signaling. The random forest model achieved excellent discriminative performance (5-fold CV AUC = 0.988), with Treg infiltration, TLR4, and MMP9 as the top predictors. qRT-PCR confirmed significant upregulation of all five candidate genes (DEFB4A, S100A8, IL-8, MMP9, and TLR4) in both A549 and THP-1 cells following fungal stimulation. CONCLUSION: This multi-omics machine learning model integrating scRNA-seq and TCGA transcriptomic data demonstrates excellent discriminative performance (AUC = 0.988), with mechanistic convergence of NF-κB, TLR4, and oncogenic signaling pathways identified across shared immune gene signatures. In vitro qRT-PCR validation confirms the biological relevance of five key antifungal immune genes, providing a transcriptomic foundation for future prospective IPFI risk stratification in patients with COPD and lung cancer.

TLR4

In silico screening of anti-atherosclerotic compounds from Morus alba leaves by machine learning and network pharmacology.

OBJECTIVE: This study integrates machine learning with network pharmacology, molecular docking, and molecular dynamics simulations to screen bioactive compounds from Mulberry leaves and elucidate their potential mechanisms against atherosclerosis (AS). METHODS: A training dataset of anti-AS active compounds was compiled and encoded as Morgan fingerprints. Three machine learning classifiers, specifically Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XG-Boost), were constructed and evaluated using multiple performance metrics. Potential active components from Mulberry leaves and AS-related targets were retrieved, followed by protein-protein interaction network construction and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Molecular docking was then performed to evaluate binding affinities between core targets and candidate compounds, and the most stable complex was subjected to molecular dynamics simulations using GROMACS (2025). RESULTS: The RF model achieved superior performance (accuracy= 0.8354, F1 = 0.8408, AUC = 0.9119) with 100% external validation accuracy. Thirteen anti-AS candidates were prioritized from mulberry leaves, four of which have been previously documented. Network pharmacology revealed AKT1 and IL6 as core targets, enriched in pathways such as endocrine resistance. Molecular docking and dynamics simulations confirmed strong binding between oxysanguinarine and AKT1, with the complex exhibiting high stability. CONCLUSION: The RF model provides a reliable computational tool for prioritizing anti-AS compounds from Mulberry leaves. The integrated analysis reveals that Mulberry leaves exert anti-atherosclerotic effects through multi-target (e.g., AKT1, IL6) and multi-pathway (e.g., PI3K-Akt) mechanisms, offering a framework for further experimental validation.

Morus

A CFH- and SPINT2-based prognostic signature for cholangiocarcinoma.

BACKGROUND: Cholangiocarcinoma (CCA) is a highly malignant tumor with a poor prognosis, and reliable biomarkers for postoperative risk stratification remain limited. This study aimed to develop and validate a CFH- and SPINT2-based prognostic signature to support postoperative risk stratification and inform adjuvant therapy selection in CCA through integrative machine learning and single-cell transcriptomics. METHODS: Differentially expressed genes were screened from GSE26566. Integrative machine learning (least absolute shrinkage and selection operator-Cox, random forest, and univariate Cox regression) was performed in the training cohort (GSE89749; n=115) to construct a risk model, which was externally validated in two independent cohorts: cohort 1 (E-MTAB-6389; n=75) and cohort 2 [The Cancer Genome Atlas Cholangiocarcinoma (TCGA-CHOL) data set; n=36]. Systematic analysis was conducted and included examinations of immune infiltration [via single-sample gene set enrichment analysis (ssGSEA)], pathway enrichment (via hallmark GSEA), cellular localization (via single-cell RNA sequencing), and drug sensitivity (via the Genomics of Drug Sensitivity in Cancer 2 database). RESULTS: Two genes, CFH and SPINT2, were identified and incorporated into a prognostic risk score. High-risk patients in the training cohort had a significantly worse overall survival (log-rank P=0.02). External validation was performed in two independent cohorts. In validation cohort 1, the risk group was an independent prognostic factor [hazard ratio =2.27, 95% confidence interval (CI): 1.18-4.37; P=0.01]. In validation cohort 2, the model demonstrated acceptable discriminative ability (concordance index =0.721; 3-year area under the curve =0.692). The high-risk group exhibited an immunosuppressive microenvironment characterized by increased infiltration of macrophages and myeloid-derived suppressor cells, along with the activation of epithelial-mesenchymal transition, inflammatory response, and NF-κB signaling pathways. Single-cell analysis revealed a cell-type-specific expression pattern: CFH was predominantly expressed in fibroblasts, while SPINT2 was mainly expressed in malignant cells. Drug sensitivity analysis demonstrated that the high-risk group was more sensitive to gemcitabine, cisplatin, poly(ADP-ribose) polymerase (PARP) inhibitors, and mammalian target of rapamycin (mTOR) inhibitors, whereas the low-risk group was more sensitive to lapatinib. CONCLUSIONS: The CFH- and SPINT2-based prognostic signature may serve as an independent biomarker for postoperative risk stratification in CCA. High-risk patients, characterized by fibroblast-derived CFH enrichment and malignant-cell SPINT2 loss, exhibit an immunosuppressive microenvironment and may be more suitable for gemcitabine-based chemotherapy or PARP/mTOR inhibitors, whereas low-risk patients may benefit from less intensive adjuvant strategies or HER2/EGFR-targeted lapatinib. Prospective validation is warranted before clinical implementation.

Cholangiocarcinoma (CCA)

Multi-cohort integration and machine learning identify CPVL as a novel oncogenic driver in gastric cancer.

BACKGROUND: Gastric cancer (GC) remains a leading cause of cancer-related mortality worldwide, and the prognosis of advanced GC remains poor. Systematic identification of robust biomarkers through multi-cohort integration and computational prioritization may facilitate the discovery of novel therapeutic targets. AIM: To identify key genes associated with gastric cancer progression through integrative multi-omics analysis and to elucidate the biological functions and molecular mechanisms of the top-prioritized candidate gene. METHODS: Comprehensive bioinformatics analyses integrating The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), and Gene Expression Omnibus (GEO) datasets were performed using differential expression analysis, weighted gene co-expression network analysis (WGCNA), Cox regression, and eight machine-learning algorithms to systematically identify and prioritize GC-associated hub genes. Among the identified candidates, CPVL was selected for further validation based on its diagnostic and prognostic performance. CPVL expression and clinical relevance were validated by independent datasets and immunohistochemistry. Lentiviral constructs were used to overexpress or silence CPVL in GC cell lines. Functional assays were performed, including CCK-8, colony formation, EdU incorporation, and flow cytometry, to assess cell proliferation and cell-cycle distribution. Western blotting and JAK2 inhibitor (AZD1480) rescue experiments were performed to elucidate the underlying mechanisms, and a nude mouse xenograft model was used to evaluate tumorigenicity in vivo. RESULTS: Multi-cohort screening identified five hub genes (CPVL, AADAC, BCAT1, CPXM1, and FBN1). Among them, CPVL exhibited the highest diagnostic accuracy (AUC = 0.895) and the strongest correlation with poor overall survival, and was therefore selected for mechanistic investigation. CPVL expression was markedly upregulated in GC tissues and cell lines. Functional assays demonstrated that CPVL promotes GC cell proliferation and accelerates G1/S-phase transition. Mechanistically, CPVL activated the JAK2/STAT3 signaling pathway, upregulating Cyclin D1 and CDK4 while downregulating p27. Treatment with the JAK2 inhibitor AZD1480 partially reversed these effects. In vivo, CPVL knockdown significantly inhibited tumor growth. CONCLUSION: Through systematic multi-cohort integration and machine-learning prioritization, CPVL was identified as a novel oncogenic driver in gastric cancer. CPVL promotes tumor growth via activation of the JAK2/STAT3 pathway and regulation of the Cyclin D1/CDK4/p27 axis, highlighting its potential as a diagnostic biomarker and therapeutic target.

Biomarker

seq2ribo: structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context. AVAILABILITY: seq2ribo is available at https://github.com/Kingsford-Group/seq2ribo.

Machine Learning

Inflammatory pathways and immune dysregulation in pediatric postoperative septic shock: A study integrating transcriptomics, machine learning and molecular docking.

This study elucidates the molecular and immune regulatory mechanisms of pediatric postoperative septic shock. Transcriptomic data were obtained from the Gene Expression Omnibus database. Differentially expressed genes were identified using the limma package, and gene co-expression modules were constructed using Weighted Gene Co-expression Network Analysis. Functional enrichment was performed via gene set enrichment analysis, Gene Ontology, and Kyoto Encyclopedia of Genes and Genomes analyses. Immune cell infiltration was assessed using ESTIMATE and CIBERSORT. Mendelian randomization was applied to explore causal relationships between gene expression and septic shock. Feature genes were selected using machine learning algorithms, and a diagnostic nomogram model was constructed. Finally, molecular docking analysis was performed to screen and evaluate the binding affinity of traditional Chinese medicine monomers to core target proteins. A total of 1331 differentially expressed genes were identified, and the turquoise module was strongly correlated with septic shock. Enrichment analysis revealed significant activation of IL-6/JAK/STAT3, TNF-α/NF-κB, and PI3K/Akt/mTOR pathways. Immune infiltration analysis indicated suppressed immune scores and imbalances in neutrophils, macrophages, T cells, and B cells. Mendelian randomization confirmed causal associations for 6 genes, including PIM3. The predictive model based on feature genes demonstrated high diagnostic performance. Molecular docking suggested that quercetin and astramembrannin I could stably bind PIM3. This study systematically identified core genes, dysregulated immune pathways, and candidate small-molecule interventions in pediatric septic shock, providing novel insights for early diagnosis and targeted therapy.

Humans

seq2ribo: Structure-aware integration of machine learning and simulation to predict ribosome location profiles from RNA sequences.

MOTIVATION: Ribosome dynamics are vital in the process of protein expression. Current methods rely on ribosome profiling (Ribo-seq), RNA-seq profiles, and full genomic context. This restricts their use in de novo sequence design, like messenger RNA (mRNA) vaccines. Simulation-only approaches like the Totally Asymmetric Simple Exclusion Process (TASEP) oversimplify translation by focusing solely on codon elongation times. RESULTS: We present seq2ribo, a hybrid simulation and machine learning framework that predicts ribosome A-site locations using only an mRNA sequence as input. Our method first employs a novel structure-aware TASEP (sTASEP), which models translation using a comprehensive set of fitted parameters that include codon wait times and structural features, such as local angles, base-pairing, and discrete positional buckets. The ribosome locations generated by sTASEP are then processed by a polisher model, which learns to refine the simulated ribosome distributions. seq2ribo provides high-fidelity predictions of ribosome locations across diverse cell types (iPSC, HEK293, LCL, and RPE-1), significantly outperforming baselines. seq2ribo is the first method to achieve meaningful positional correlation with observed ribosome profiles from sequence alone, reaching transcript-level Pearson correlations up to 0.920 and within-transcript shape correlations up to 0.186, where all baselines yield near-zero values on these metrics. seq2ribo also reduces elementwise error by up to 37.7% relative to the sequence-only Translatomer baseline. By adding a task-specific head, seq2ribo achieves Pearson correlations up to 0.732 with experimental translation efficiency (TE) across several cell lines, and up to 0.903 with measured protein expression. By operating from sequence alone, seq2ribo provides a new tool for synthetic biology, enabling the rational design and optimization of mRNA sequences without the need for expression-level data or genomic context.

Journal Article

Modeling unknowns: A vision for uncertainty-aware machine learning in healthcare.

The integration of machine learning (ML) into healthcare is accelerating, driven by the proliferation of biomedical data and the promise of data-driven clinical support. A key challenge in this context is managing the pervasive uncertainty inherent in medical reasoning and decision-making. Despite its recognized importance, uncertainty is often underrepresented in the design and evaluation of clinical AI systems. Here we report an editorial overview of a special issue dedicated to uncertainty modeling in medical AI, which gathers theoretical, methodological, and practical contributions addressing this critical gap. Across these works, authors reveal that fewer than 4% of studies address uncertainty explicitly, and propose alternative design principles-such as optimizing for clinical net benefit or embedding explainability with confidence estimates. Notable contributions include the RelAI system for real-time prediction reliability, empirical findings on how uncertainty communication shapes clinical interpretation, and benchmarks for out-of-distribution detection in tabular data. Furthermore, this issue highlights the use of causal reasoning and anomaly detection to enhance system robustness and accountability. Together, these studies argue that representing, communicating, and operationalizing uncertainty are essential not only for clinical safety but also for building trust in AI-driven care. This special issue thus repositions uncertainty from a limitation to a foundational asset in the responsible deployment of ML in healthcare.

Machine Learning

A machine learning-derived intratumoral heterogeneity-related signature predicts the prognosis for and therapeutic response in patients with skin cutaneous melanoma.

BACKGROUND: Reliable biomarkers for predicting prognosis and therapeutic response in skin cutaneous melanoma (SKCM) remain limited. This study aimed to develop an intratumoral heterogeneity (ITH)-related prognostic signature for SKCM using integrative machine learning. METHODS: RNA sequencing (RNA-seq) data from 472 SKCM patients in The Cancer Genome Atlas (TCGA) and 214 patients in the GSE65904 cohort were analyzed. ITH scores were calculated using the DEPTH2 algorithm. Differentially expressed genes (DEGs) were identified between high- and low-ITH groups [|log2fold change (FC)| &#x2265;1, false discovery rate (FDR) <0.05]. Based on 38 prognostic DEGs identified by univariate Cox regression, we employed an integrative framework of 101 machine learning algorithm combinations to construct prognostic models in the TCGA training cohort. The model with the highest average concordance index (C-index) was validated in the GSE65904 cohort and selected as the prognostic ITH-related signature (PIRS). Associations of the PIRS risk score with tumor mutational burden (TMB), immune cell infiltration, immune checkpoint gene expression, and drug sensitivity were systematically evaluated. Model performance was assessed using receiver operating characteristic (ROC) curves and Cox regression analyses. RESULTS: A 38-gene PIRS was constructed using the plsRcox algorithm. Patients with high PIRS risk scores exhibited significantly poorer overall survival (OS) in both the TCGA and Gene Expression Omnibus (GEO) cohorts. The PIRS was identified as an independent prognostic factor, with area under the curve (AUC) values of 0.779, 0.734, and 0.756 for 1-, 3-, and 5-year survival, respectively. High-risk samples displayed significantly lower TMB (P<0.05), reduced immune and stromal cell infiltration (P<0.001), downregulated immune function, and decreased expression of immune checkpoint genes. Additionally, high- and low-PIRS risk score groups exhibited distinct sensitivity patterns to different classes of targeted agents. CONCLUSIONS: The machine learning-derived PIRS robustly predicts prognosis in SKCM patients. Its clinical application is promising for optimizing patient risk stratification and treatment decisions, though further prospective validation is warranted.

Skin cutaneous melanoma (SKCM)

Multi-omics dynamic profiling reveals predictive biomarkers for first-line immunochemotherapy in extensive-stage small-cell lung cancer.

BACKGROUND: Extensive-stage small-cell lung cancer (ES-SCLC) is associated with a poor prognosis. Although first-line immunochemotherapy improves clinical outcomes, robust prognostic biomarkers for this treatment modality remain unavailable. The aim of this study was to identify non-invasive, easily accessible, and dynamically monitored biomarkers of ES-SCLC by machine learning integrating serum metabolomics, lipidomics, and proteomics at multiple time points. METHODS: A total of 816 serum samples were collected from ES-SCLC patients receiving first-line immunotherapy combined with chemotherapy or first-line chemotherapy for metabolomics, lipidomics, and proteomics analysis. The immunochemotherapy cohort was randomly divided into training and validation subsets at a 6:4 ratio. Biomarkers were identified using machine learning algorithms, and their prognostic significance was evaluated through receiver operating characteristic (ROC) analysis, Kaplan&#x2013;Meier survival analysis, and multivariate Cox regression. Potential metabolic pathways and mechanisms were further explored via integrated multi-omic analysis. RESULTS: The immunochemotherapy exhibited a prolonged median progression-free survival (PFS) and higher objective response rate (ORR) compared to the chemotherapy group. A total of 5 serum metabolites (uric acid, L-aspartate-semialdehyde, dimethisterone, xanthine, L-cysteine), 6 lipids (Cer d18:1/26:0, Cer d18:2/25:0, SM d18:1/20:1, SM d17:1/25:1, DG O-18:1_16:0, PS 18:0_24:0), and 3 proteins (ACIN1, ACSL4, PHGDH) were identified and constructed into independent prognostic models. Among patients receiving immunochemotherapy, those categorized as low-risk based on the model demonstrated significantly longer PFS compared with those in the high-risk group. These prognostic signatures also retained predictive value in patients who underwent second-line treatment with anlotinib plus immunochemotherapy. Integrated analysis revealed that glycine, serine, and threonine metabolism was the commonly enriched pathway across all three omics layers. Notably, PHGDH (protein), L-aspartate-semialdehyde and L-cysteine (metabolites), and PS (18:0_24:0) (lipid), key elements in this pathway, were all incorporated in the predictive model. In addition, models of the composition of these substances after one cycle of treatment can still predict the prognosis of patients. CONCLUSION: In this study, we constructed and validated a set of non-invasive, dynamically monitorable prognostic models (containing 5 metabolites, 6 lipids, and 3 proteins) using machine learning by integrating multiple time point data from the serum metabolome, lipid panel, and proteome to accurately distinguish the prognostic risk of patients with ES-SCLC receiving immunochemotherapy. PFS was significantly prolonged in patients in the low-risk group, and this model remains predictive in the subsequent second-line treatment with anlotinib in combination with immunochemotherapy. Glycine-serine-threonine metabolic pathway may be the key mechanism, of which PHGDH, L-aspartate semialdehyde, L-cysteine and PS (18:0_24:0) are the core predictors. This study provides the first multi-omics dynamic prognostic tool for ES-SCLC immunochemotherapy and reveals potential therapeutic targets.

Humans

Flnc: Machine Learning Improves the Identification of Novel Long Noncoding RNAs from Stand-Alone RNA-Seq Data.

Long noncoding RNAs (lncRNAs) play critical regulatory roles in human development and disease. Although there are over 100,000 samples with available RNA sequencing (RNA-seq) data, many lncRNAs have yet to be annotated. The conventional approach to identifying novel lncRNAs from RNA-seq data is to find transcripts without coding potential but this approach has a false discovery rate of 30-75%. Other existing methods either identify only multi-exon lncRNAs, missing single-exon lncRNAs, or require transcriptional initiation profiling data (such as H3K4me3 ChIP-seq data), which is unavailable for many samples with RNA-seq data. Because of these limitations, current methods cannot accurately identify novel lncRNAs from existing RNA-seq data. To address this problem, we have developed software, Flnc, to accurately identify both novel and annotated full-length lncRNAs, including single-exon lncRNAs, directly from RNA-seq data without requiring transcriptional initiation profiles. Flnc integrates machine learning models built by incorporating four types of features: transcript length, promoter signature, multiple exons, and genomic location. Flnc achieves state-of-the-art prediction power with an AUROC score over 0.92. Flnc significantly improves the prediction accuracy from less than 50% using the conventional approach to over 85%. Flnc is available via GitHub platform.

RNA-seq

Mining metagenomes from extremophiles as a resource for novel glycoside hydrolases for industrial applications.

The exploration of metagenomes from extremophiles has emerged as a promising approach for discovering novel glycoside hydrolases (GHs) with potential industrial applications. Extremophiles, which thrive in harsh conditions such as high salinity, extreme temperatures, and acidic or alkaline environments, produce enzymes naturally adapted to function under these conditions. This unique adaptability makes them highly desirable for industrial processes requiring robust and efficient biocatalysts. These biocatalysts reduce reliance on harsh chemicals and energy-intensive processes, contributing to greener industrial operations. This review underscores the power of metagenomics in bypassing the need to culture large libraries of extremophiles in the lab. High-throughput sequencing and bioinformatics enable the identification of novel GH-encoding genes directly from environmental DNA. While metagenomic mining has yielded promising results, challenges such as the expression of extremophile-derived genes in mesophilic hosts, low activity yields, and scalability remain. Advances in synthetic biology and protein engineering could address these bottlenecks, enabling more efficient utilization of GHs. Additionally, integrating machine learning for predictive functional annotation may accelerate the identification of high-value candidates.

Glycoside Hydrolases

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] >&#x2009;0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans

Machine learning-based integration develops a novel lysosome-related prognostic signature associated with prognosis and immune infiltration landscape in acute myeloid leukemia.

BACKGROUND: Lysosomes are essential for intracellular degradation and recycling, and changes in their function significantly contribute to tumor growth. Nonetheless, the exact role of lysosome-related genes (LRGs) in the pathogenesis of acute myeloid leukemia (AML) is still inadequately comprehended. METHODS: Differentially expressed LRGs (DE-LRGs) between AML and control groups were identified using AML-related data extracted from the Gene Expression Omnibus (GEO). The LRGs-related prognostic genes were identified and the risk model was established using univariate COX regression analysis and machine learning algorithms, based on the data obtained from The Cancer Genome Atlas (TCGA). Subsequently, we performed comprehensive analyses regarding clinical features, functional pathways, immune microenvironment, and chemotherapeutic drugs sensitivity between the high- and low-risk groups. Reverse transcription Quantitative polymerase chain reaction (RT-qPCR) and western blot were adopted to validate the expression of prognostic genes in human bone marrow-derived cell line HS-27&#xa0;A and human AML cell line MOLM-13. RESULTS: Through comprehensive analysis, a risk model was developed utilizing ten LRGs (ATP6V0E2, CALCRL, TMEM165, GZMB, HCK, TCIRG1, CD1D, GPRASP1, ABCA1, and NAGA), and this model was further validated using GEO datasets. Significant differences in clinical characteristics, functional pathways, immune microenvironment characteristics, and chemotherapeutic drug sensitivity were observed between the two risk groups In vitro validation experiment illustrated that the expression trends of ATP6V0E2, TMEM165, and ABCA1 were consistent with our bioinformatics analysis. CONCLUSION: Our study demonstrates that lysosome-associated signature might forecast the prognosis of AML patients and offer guidance for subsequent immunotherapy and chemotherapy strategies.

Acute myeloid leukemia