PubMed HealthSearch

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning

Multi-cohort integration and machine learning identify CPVL as a novel oncogenic driver in gastric cancer.

BACKGROUND: Gastric cancer (GC) remains a leading cause of cancer-related mortality worldwide, and the prognosis of advanced GC remains poor. Systematic identification of robust biomarkers through multi-cohort integration and computational prioritization may facilitate the discovery of novel therapeutic targets. AIM: To identify key genes associated with gastric cancer progression through integrative multi-omics analysis and to elucidate the biological functions and molecular mechanisms of the top-prioritized candidate gene. METHODS: Comprehensive bioinformatics analyses integrating The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), and Gene Expression Omnibus (GEO) datasets were performed using differential expression analysis, weighted gene co-expression network analysis (WGCNA), Cox regression, and eight machine-learning algorithms to systematically identify and prioritize GC-associated hub genes. Among the identified candidates, CPVL was selected for further validation based on its diagnostic and prognostic performance. CPVL expression and clinical relevance were validated by independent datasets and immunohistochemistry. Lentiviral constructs were used to overexpress or silence CPVL in GC cell lines. Functional assays were performed, including CCK-8, colony formation, EdU incorporation, and flow cytometry, to assess cell proliferation and cell-cycle distribution. Western blotting and JAK2 inhibitor (AZD1480) rescue experiments were performed to elucidate the underlying mechanisms, and a nude mouse xenograft model was used to evaluate tumorigenicity in vivo. RESULTS: Multi-cohort screening identified five hub genes (CPVL, AADAC, BCAT1, CPXM1, and FBN1). Among them, CPVL exhibited the highest diagnostic accuracy (AUC = 0.895) and the strongest correlation with poor overall survival, and was therefore selected for mechanistic investigation. CPVL expression was markedly upregulated in GC tissues and cell lines. Functional assays demonstrated that CPVL promotes GC cell proliferation and accelerates G1/S-phase transition. Mechanistically, CPVL activated the JAK2/STAT3 signaling pathway, upregulating Cyclin D1 and CDK4 while downregulating p27. Treatment with the JAK2 inhibitor AZD1480 partially reversed these effects. In vivo, CPVL knockdown significantly inhibited tumor growth. CONCLUSION: Through systematic multi-cohort integration and machine-learning prioritization, CPVL was identified as a novel oncogenic driver in gastric cancer. CPVL promotes tumor growth via activation of the JAK2/STAT3 pathway and regulation of the Cyclin D1/CDK4/p27 axis, highlighting its potential as a diagnostic biomarker and therapeutic target.

Biomarker

Uncovering encrypted antimicrobial peptides in health-associated Lactobacillaceae by large-scale genomics and machine learning.

BACKGROUND: Antimicrobial peptides (AMPs) are well known for their broad-spectrum activity and have shown great promise in addressing the antibiotic-resistant crisis. The Lactobacillaceae family, recognized for its health-promoting effects in humans, represents a valuable source of novel AMPs. However, the global prevalence and distribution of AMPs within Lactobacillaceae remains largely unknown, which limits the efficient discovery and development of novel AMPs. RESULTS: We analyzed all available genomes (10,327 genomes), encompassing 38 genera and 515 species, to investigate the biosynthetic potential (indicated by the number of AMP sequences in the genome) of AMP in the Lactobacillaceae family. We demonstrated Lactobacillaceae species had ubiquitous (69.90%) biosynthetic potential of AMPs. Overall, 9601 AMPs were identified, clustering into 2092 gene cluster families (GCFs), which showed strong interspecies specificity (95.27%), intraspecies heterogeneity (93.31%), and habitat uniqueness (95.83%), that greatly expanded on the AMP sequence landscape. Novelty assessment indicated that 1516 GCFs (72.47%) had no similarity to any known AMPs in existing databases. Machine learning predictions suggested that novel AMPs from Lactobacillaceae possessed strong antimicrobial potential, with 664 GCFs having an additive minimum inhibitory concentration (MIC) below 100&#xa0;&#x3bc;M. We randomly synthesized 16 AMPs (with predicted MIC&#x2009;<&#x2009;100&#xa0;&#x3bc;M) and identified 10 AMPs exhibiting varied-spectrum activity against 11 common pathogens. Finally, we identified one Lactobacillus delbrueckii-originated AMP (delbruin_1) having broad-spectrum (all 11 pathogens) and high antimicrobial activity (average MIC&#x2009;=&#x2009;38.56 &#xb5;M), which proved its potential as a clinically viable antimicrobial agent. CONCLUSIONS: We uncovered the global prevalence of AMPs in Lactobacillaceae and proved that Lactobacillaceae is an untapped and invaluable source of novel AMPs to combat the antibiotic-resistance crisis. Meanwhile, we provided a machine learning-guided framework for AMP discovery, offering a scalable roadmap for identifying novel AMPs not only in Lactobacillaceae but also in other organisms. Video Abstract.

Machine Learning

A genetic-based machine learning system to discover the diagnostic rules for female urinary incontinence.

A machine learning system named Galactica has been developed which uses a genetic algorithm to discover the rules for an expert system from databases. Galactica devised accurate diagnostic rules for female urinary incontinence from difficult heterogeneous data. The percentages of correctly classified stress, mixed and sensory urge incontinence testing cases were 89, 86 and 87%, respectively. However, these rules were rather general, consisting of 4-6 out of 13 conditions available in the data. Diagnostic rules for stress and mixed incontinence extracted from straightforward homogeneous data were highly accurate, classifying 100% of testing cases correctly as well as being specific, having from 10 to 11 conditions. More specific, but less accurate, rules were found from heterogeneous data with a biased fitness function. All of the rules were correct, i.e. every condition in the rules had the expected value specified by the expert. Although, Galactica achieved a slightly better classification than the discriminant analysis, it is argued that the genetic approach is better than the statistical one, due to symbolic rules being comprehensible, whereas understanding a complex mathematical model requires statistical expertise.

Algorithms

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning

Transcriptome Analysis, Machine Learning, and Experimental Identification of CDK7 Affecting the Progression of Pregnancy-induced Hypertension by Influencing Macrophage Polarization.

INTRODUCTION: Pregnancy-induced hypertension (PIH) is a severe pregnancy complication characterized by placental insufficiency, abnormal vascular remodeling, and immune dysregulation, but personalized therapeutic markers remain unclear. This study aimed to identify key genes and explore immune mechanisms in PIH using transcriptome analysis, machine learning, and experimental validation. METHODS: We analyzed the GSE204835 transcriptomic dataset to screen differentially expressed genes (DEGs) and performed Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), Reactome, and Gene Set Enrichment Analysis (GSEA) for functional annotation. Immune infiltration analysis was also performed to examine the immune landscape in PIH. Least Absolute Shrinkage and Selection Operator (LASSO) regression identified key genes, which were validated in a PIH cell model. Flow cytometry and immunofluorescence assays assessed the effect of CDK7 knockdown on macrophage polarization. RESULTS: A total of 1,598 DEGs (1,123 upregulated, 475 downregulated) were identified. Enrichment analyses highlighted associations with embryonic organ development, oxidative phosphorylation, angiogenesis, and oxidative stress. Immune infiltration analysis revealed altered eosinophil and macrophage polarization in PIH. LASSO regression selected 12 key genes, with CDK7 showing the most significant upregulation in the PIH model. CDK7 knockdown promoted macrophage polarization toward the anti-inflammatory M2 phenotype. DISCUSSION: These findings link CDK7 to immune dysregulation in PIH by modulating macrophage polarization, expanding our understanding of PIH's molecular mechanisms. The study's limitations include reliance on public datasets and in vitro models, warranting in vivo validation. CONCLUSION: CDK7 emerges as a potential therapeutic target for PIH, offering new insights into immunoregulatory interventions for this complication.

Female

Screening of the key single nucleotide polymorphisms in type 2 diabetes mellitus complicated with lower extremity arterial disease by machine learning.

OBJECTIVES: Diabetic lower extremity arterial disease (LEAD) is a manifestation of diabetic lower extremity vascular complications. This study aimed to screen the key single nucleotide polymorphism (SNP) gene signature in patients with type 2 diabetes mellitus (T2DM) and LEAD. METHODS: A total of 147 patients with T2DM complicated by LEAD and 144 patients with T2DM without LEAD were enrolled for transcriptome sequencing. The Plink software was used to preprocess the data. Five machine learning methods were adopted to build the SNP diagnosis models. The receiver operating characteristic (ROC) curve was used to quantify the predicted probabilities of the model. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed using the cluster Profiler package. Finally, regression statistical analysis was used to correlate the key SNPs with clinical information and biochemical indicators. RESULTS: A total of 24 SNPs were retained and 10 SNPs were risk allele genes. Nine SNPs (rs7412, rs1800629, rs699947, rs3918242, rs668, rs1800470, rs1800449, rs1800469, and rs1024611) were identified as the key SNPs sites. GO and KEGG pathway analyses revealed that these genes are mainly enriched in fluid shear stress and atherosclerosis. Finally, rs1800449 was associated with low-density lipoprotein cholesterol (LDL-C). With high density lipoprotein cholesterol (HDL-C), related site was rs1024611. The sites associated with total cholesterol (CHOL) were rs1800449 and rs7412.The site associated with apolipoprotein B (APOB) and apolipoprotein A1 (APOA1) were rs1800470 and rs1800469. CONCLUSION: This study authenticated nine SNPs for the diagnosis of T2DM patients with LEAD, which will be of great significance in the development of diagnostic molecular biomarkers for T2DM patients.

Humans

Integrative multi-omics and machine learning identify the SPI1-METTL16-PLIN4 axis as a candidate driver of steatosis in HepG2 cells.

BACKGROUND: Non-alcoholic fatty liver disease (NAFLD) is a prevalent metabolic disorder with limited therapeutic options. This study aimed to identify potential regulators and explore their functional roles in a cellular model of NAFLD. METHODS: WGCNA was performed on the hepatic transcriptomic dataset GSE126848 (31 NAFLD vs. 26 controls), followed by integration with serum proteomic data from 12 NAFLD patients and 12 healthy controls. Hub genes were prioritized using three machine learning algorithms. Functional validation was conducted in a HepG2 cellular steatosis model induced by high fructose (3.2&#x202f;g/L) and oleic acid (400&#x202f;&#x3bc;M) for 48&#x202f;h. Lipid accumulation was assessed by Oil Red O staining and triglyceride/total cholesterol measurement. Inflammation was evaluated by TNF-&#x3b1; and IL-6 secretion (ELISA), and oxidative stress by ROS levels (flow cytometry). The binding interaction between METTL16 and PLIN4 mRNA was validated by RNA immunoprecipitation (RIP)-quantitative PCR. METTL16-mediated m6A modification of PLIN4 was assessed by Methylated RIP (MeRIP)-quantitative PCR. Transcriptional regulation of METTL16 by SPI1 was examined by chromatin immunoprecipitation (ChIP) and dual-luciferase reporter assays. RESULTS: Integrative analysis identified PLIN4 as a core hub gene. PLIN4 was upregulated in the HepG2 steatosis model (P&#x202f;<&#x202f;0.001). PLIN4 knockdown alleviated lipid droplet accumulation (P&#x202f;<&#x202f;0.001), reduced TNF-&#x3b1; and IL-6 secretion (P&#x202f;<&#x202f;0.01), and decreased ROS levels (P&#x202f;<&#x202f;0.001) in fructose/oleic acid-treated HepG2 cells. Mechanistically, METTL16 mediated its m6A modification to enhance PLIN4 mRNA stability. Furthermore, SPI1 was found to transcriptionally activate METTL16 by binding to its promoter (P&#x202f;<&#x202f;0.001). PLIN4 re-expression partially reversed the protective effects of SPI1 knockdown on lipid accumulation (P&#x202f;=&#x202f;0.01), inflammation (P&#x202f;<&#x202f;0.05), and oxidative stress (P&#x202f;<&#x202f;0.001). CONCLUSION: This study identifies the SPI1/METTL16/PLIN4 axis as a potential regulatory mechanism contributing to in vitro steatosis, inflammation, and oxidative stress in steatotic HepG2 cells.

Humans

Machine learning of motor vehicle accident categories from narrative data.

Bayesian inferencing as a machine learning technique was evaluated for identifying pre-crash activity and crash type from accident narratives describing 3,686 motor vehicle crashes. It was hypothesized that a Bayesian model could learn from a computer search for 63 keywords related to accident categories. Learning was described in terms of the ability to accurately classify previously unclassifiable narratives not containing the original keywords. When narratives contained keywords, the results obtained using both the Bayesian model and keyword search corresponded closely to expert ratings (P(detection) > or = 0.9, and P (false positive) < or = 0.05). For narratives not containing keywords, when the threshold used by the Bayesian model was varied between p > 0.5 and p > 0.9, the overall probability of detecting a category assigned by the expert varied between 67% and 12%. False positives correspondingly varied between 32% and 3%. These latter results demonstrated that the Bayesian system learned from the results of the keyword searches.

Accidents, Traffic

Machine Learning in Hyperlipidaemia Research: Screening and Experimental Insights into Lipid Metabolism Modulators.

Hyperlipidemia, characterized by elevated blood lipid levels, represents a major global health concern due to its strong association with cardiovascular disease, diabetes, and metabolic syndrome. While current therapies - such as statins, fibrates, bile acid sequestrants, and PCSK9 inhibitors - are effective in controlling hyperlipidemia, they are often associated with adverse effects, potential drug resistance, and suboptimal efficacy in certain patient populations. All of the above underscore the urgent need for safer and more effective therapeutic alternatives. Among the major molecular targets involved in the regulation of lipid metabolism are HMG-CoA reductase, PCSK9, peroxisome proliferator-activated receptors (PPARs), cholesteryl ester transfer protein (CETP), and nuclear receptors, including the liver X receptor (LXR) and farnesoid X receptor (FXR), which are also targets for future antihyperlipidemic drug development. Recent advancements in artificial intelligence (AI) and machine learning (ML) have significantly transformed and accelerated drug discovery by enabling the processing of vast amounts of genomic, proteomic, and chemical data. Furthermore, ML tools such as quantitative structure-activity relationship (QSAR) modelling, deep learning, random forest, and support vector machines (SVM) have proven predictive and effective in identifying novel lipid metabolism modulators, thereby enhancing the efficacy and accuracy of virtual screening. Meanwhile, molecular docking has become an integral part of structure-based drug design (SBDD), and software such as AutoDock, Glide, and GOLD have proven effective in generating accurate ligand-target docking models. Molecular docking, together with ML-based approaches, enables the identification of potent and selective drug candidates. Overall, the combination of ML and molecular docking offers an efficient and accurate platform for antihyperlipidemic drug discovery, helping to overcome the limitations of currently available therapeutic strategies.

HMG-CoA reductase

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans

Pan-cancer multi-omics machine learning defines a lactylation-associated immune-excluded tumor state with proteomic and experimental corroboration.

BACKGROUND: Histone lactylation links lactate metabolism to chromatin regulation, but whether lactylation-program-associated transcriptional patterns delineate recurrent pan-cancer tumor states remains unclear. METHODS: We integrated mRNA, lncRNA, and miRNA profiles from 9712 TCGA tumors across 33 cancer types with GTEx references, six GEO cohorts, IMvigor210, and an institutional clear-cell renal cell carcinoma (ccRCC) cohort used for exploratory DIA-NN proteomic corroboration. Random-effects co-expression meta-analysis, multi-omics consensus clustering, regulon inference, immune deconvolution, TIDE, oncoPredict, and SHAP-based machine learning were applied. hsa-miR-431-5p was functionally evaluated as a proof-of-concept CS2-associated miRNA in bladder cancer models. RESULTS: LacCoEx-Atlas comprised 398,491 lactylation-related co-expression pairs across 24,667 RNA features under a random-effects framework (median I&#xb2; = 88.6%). Consensus clustering identified two subtypes: CS2 showed glycolytic-mesenchymal-immune-excluded features, M2 macrophage enrichment, CD8&#x207a; T-cell depletion, elevated HDAC4/NSD3/KDM6B activity, and worse survival, whereas CS1 showed oxidative, sirtuin-active programs. CS2 had fewer predicted ICI responders (18.3% vs. 52.0%) and a lower observed ORR in IMvigor210 (15.3% vs. 24.0%). oncoPredict identified NU7441 as a hypothesis-generating CS2-associated sensitivity signal (Hedges' g = 1.17). DIA-NN proteomics in 50 ccRCC specimens provided exploratory support for CS2-associated hypoxia, ECM degradation, and metastasis programs. The 10-feature mRNA LARItools model achieved an apparent AUC of 0.9413, while a separate multi-omics model achieved 0.971; neither was independently validated. LARItools reproduced prognostic separation across six GEO cohorts. miR-431-5p promoted malignant phenotypes and EMT in bladder cancer cells, with concordant CMU4h expression findings. CONCLUSIONS: Lactylation-program-associated transcriptional patterns delineate a recurrent immune-excluded pan-cancer tumor state associated with adverse prognosis, reduced predicted immunotherapy responsiveness, exploratory single-cancer protein-level support, and testable DNA damage response-targeting hypotheses. LacCoEx-Atlas and LARItools provide open resources for lactylation-program-associated tumor-state stratification and future translational research.

Humans

Multimodal features and prognostic risk assessment in locally advanced gastric cancer patients following neoadjuvant therapy based on machine learning algorithms: a multicenter study.

BACKGROUND: Neoadjuvant therapy (NAT) is recommended for locally advanced gastric cancer (LAGC), but some patients respond poorly. We aimed to construct a multimodal model integrating CT images, transcriptomic sequencing, and clinicopathological data to assess prognosis in LAGC patients receiving NAT. MATERIALS AND METHODS: This multicenter study included 505 LAGC patients who underwent NAT. Radiomic features were extracted from preoperative CT images of 505 patients. RNA-seq was performed on 277 post-NAT specimens, with additional data from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases (n&#x2009;=&#x2009;804). Patients were divided into training (168 cases), internal validation (72 cases), and external validation cohorts. Machine learning algorithms identified key radiomic, molecular, and clinical features associated with NAT response, which were then integrated into a multimodal model to predict overall survival (OS) and disease-free survival (DFS). RESULTS: Six radiomic and three molecular features significantly associated with NAT response were selected. Radiomic risk (hazard ratio [HR]: 4.0, P&#x2009;<&#x2009;0.001) and molecular risk (HR: 7.1, P&#x2009;<&#x2009;0.001) were independent prognostic factors. By integrating radiomic risk, molecular risk, and clinical characteristics, a multimodal model (MuMo) was constructed.The C-index results (OS, C-index&#x2009;=&#x2009;0.855; DFS, C-index&#x2009;=&#x2009;0.786) demonstrated that MuMo outperformed the single-modality models and ypTNM staging.Mechanistic analysis suggested that the efficacy of neoadjuvant therapy was significantly enriched in immune-inflammatory pathways. CONCLUSIONS: MuMo can effectively predict postoperative survival risk in LAGC patients receiving NAT, serving as a powerful tool for optimizing prognostic assessment.

Humans

Multi-level Transcriptomic and Machine-learning Analyses Identify MZT1 as a Proliferation-associated Prognostic Marker in Lung Adenocarcinoma.

BACKGROUND/AIM: Lung adenocarcinoma (LUAD) exhibits substantial molecular heterogeneity and variable clinical outcomes, highlighting the need for biomarkers that reflect core tumor biological processes. Centrosome-associated proteins regulate mitotic fidelity and genome stability, yet their roles in LUAD remain incompletely defined. In this study, we systematically characterized mitotic spindle organizing protein 1 (MOZART1; MZT1) and related family members in LUAD. MATERIALS AND METHODS: We performed integrated analyses combining bulk transcriptomic datasets, survival modeling, gene set enrichment, immune deconvolution, machine-learning based prognostic modeling, and single-cell RNA sequencing. Expression patterns and clinical associations of MZT family genes were evaluated across pan-cancer and LUAD cohorts. RESULTS: MZT family genes were consistently upregulated in tumor tissues, with MZT1 showing the most robust expression pattern. Elevated MZT1 expression was significantly associated with reduced overall survival. Functional analyses revealed coordinated activation of proliferative and genome maintenance pathways, including G2/M checkpoint regulation, E2F and MYC signaling, and DNA repair. A multivariable analysis indicated that the prognostic association of MZT1 was reduced after adjusting for canonical proliferation markers, suggesting partial overlap with established proliferation signals. The LASSO-based Cox model demonstrated stable time-dependent predictive performance at 1-, 3-, and 5-year survival. Immune analyses indicated associations between MZT1 expression and tumor microenvironmental features. Single-cell analysis showed that MZT1 expression was predominantly enriched in malignant epithelial cells and associated with proliferative cellular states. Protein-level validation supported concordance with transcriptomic findings. CONCLUSION: MZT1 is a proliferation-associated marker that integrates clinical risk, transcriptional programs, cellular heterogeneity, and predictive modeling in LUAD, providing a potential framework for biomarker development and risk stratification.

Humans

Machine learning approaches for cancer prognosis and diagnosis via non-coding RNA: a comprehensive review.

Non-coding RNAs (ncRNAs), once considered genomic dark matter, are now established as key regulators of gene expression with widespread roles in cellular homeostasis and disease. In cancer, ncRNA expression is frequently and systematically dysregulated, and many of these molecules circulate in stable, protected form within biofluids, offering a compelling basis for non-invasive or minimally invasive diagnostic strategies. However, their clinical translation remains substantially hindered to date due to biological complexity, technical noise, and high dimensionality inherent to ncRNA expression datasets. In this context, machine learning (ML) has emerged as a powerful analytical tool to address these challenges, enabling the identification of subtle, reproducible ncRNA signatures predictive of diverse malignancies. This review critically evaluates ML-driven frameworks for cancer diagnosis and prognosis across four ncRNA subclasses, namely miRNAs, lncRNAs, circRNAs, and piRNAs, while also acknowledging the biophysical and thermodynamic models that reinforce ncRNA bioinformatics. Despite substantial methodological progress in ML-based cancer diagnosis and prognosis, key challenges persist, including tumor biological heterogeneity, limited multicenter validation, and the lack of widely adopted standardized protocols for preprocessing, normalization, and reporting workflows. Furthermore, many current ML models lack interpretability in biological or clinical context, constraining their translational utility. By synthesizing recent advances and identifying unresolved barriers, this review charts a roadmap for developing a robust, clinically actionable ncRNA biomarker platform for cancer detection. With global cancer incidence projected to exceed 35 million annual cases by 2050, validated ncRNA-ML-driven frameworks hold potential to revolutionize early-stage detection and personalized therapeutic strategies, thereby reducing the escalating socio-economic burden of cancer worldwide.

Humans

Exploring the Genetic Link between Irritable Bowel Syndrome and Polycystic Ovary Syndrome: Bidirectional Mendelian Randomization and Machine Learning Approaches.

BACKGROUND: Research has shown a certain correlation between polycystic ovary syndrome (PCOS) and irritable bowel syndrome (IBS). The study aims to determine the directionality and underlying biological processes influencing the relationship between these two disorders. METHODS: We explored the causal relationship between IBS and PCOS by conducting a comprehensive bidirectional Mendelian randomization (MR) analysis using five different methods and conducted robustness assessments. We extracted differentially expressed genes from the IBS and PCOS datasets for Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment analysis. Additionally, we developed a protein-protein interaction (PPI) network and applied the Least Absolute Shrinkage and Selection Operator (LASSO) and Support Vector Machine (SVM) methodologies to pinpoint key diagnostic markers. Diagnostic efficacy was further assessed through Receiver Operating Characteristic (ROC) curve analysis for selected genes. Finally, single-sample gene set enrichment analysis (ssGSEA) was carried out to examine immune cell infiltration in IBS and PCOS. RESULTS: MR analysis identified a causal effect of PCOS on IBS (IVW, OR = 1.034, 95% CI: 1.003-1.065, P = 0.029). Conversely, no relationship between IBS and PCOS was observed in the reverse analysis. Furthermore, integrative bioinformatics and machine learning analyses identified CD14 and CASP1 as key diagnostic biomarkers for both IBS and PCOS, which were significantly associated with immune cell infiltration. CONCLUSION: MR analysis has demonstrated a significant positive causal relationship between PCOS and IBS, though the reverse causality from IBS to PCOS appeared non-significant. The genes CD14 and CASP1 emerged as potential shared diagnostic markers between these two conditions.

Polycystic Ovary Syndrome

A machine learning approach to identify key epigenetic transcripts for ageing research in human blood (Epitage).

DNA methylation is an established biomarker of human ageing and is used by a variety of tools to identify meaningful epigenetic signals. We investigated whether analysing CpGs grouped by transcript as functional units could generate a ranked list of transcripts most correlated with age that might otherwise be overlooked in genome-wide CpG-based studies. Here we present Epitage ( https://github.com/a00s/epitage ), a continuously updated ranked list of transcripts built from the GSE87571 dataset (714 whole-blood samples, ages 14-94 years) through intensive testing with machine-learning models. To support reproducible analyses, we developed ugPlot ( https://github.com/a00s/ugplot ), an open-source R/Shiny tool with a graphical user interface that automates model training, testing, and comparison. Initially, we identified 48 transcripts across 13 genes, with some transcripts from the genes OBSCN, PRRT1, and SPTBN4 showing better predictive performance when multiple associated CpGs were analysed together rather than individually. In contrast, for the majority of transcripts, a dominant individual CpG still showed a higher Spearman correlation with age, as seen in established ageing genes such as ELOVL2, FHL2, and TRIM59. Epitage is a transcript-ranking list based on the methylation patterns observed in the analysed dataset. It provides a reproducible framework for prioritising transcripts associated with human ageing and for guiding future epigenetic studies.

Humans