PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Artificial Intelligence and Machine Learning Applications in Fibromuscular Dysplasia: Transforming Diagnosis, Risk Stratification, and Clinical Decision-Making.

Fibromuscular dysplasia (FMD) is a non-atherosclerotic vascular disorder with heterogeneous presentations, making diagnosis and management highly dependent on imaging and clinical expertise. This narrative review examines how artificial intelligence (AI) and machine learning (ML) are transforming FMD care. AI-enhanced imaging, particularly convolutional neural network-based analysis, improves detection of the characteristic "string-of-beads" pattern on CT angiography, magnetic resonance angiography, and ultrasound, although FMD-specific validation remains limited. ML models facilitate risk stratification, prediction of disease progression, and early identification of complications such as aneurysms and stroke by integrating clinical, imaging, and genomic data. AI-driven clinical decision support systems further enable personalized treatment selection through pharmacogenomic insights and robot-assisted interventions. Despite promising real-world applications, challenges persist, including limited large-scale datasets, workflow integration, regulatory barriers, and algorithmic bias affecting underrepresented populations. Future advances in explainable AI, federated learning, and digital health integration may enable a shift toward predictive, patient-centered FMD management.

Humans

Diagnostic Performance of Machine Learning for Systemic Lupus Erythematosus: Systematic Review and Meta-Analysis.

BACKGROUND: Early and accurate diagnosis of systemic lupus erythematosus (SLE) and its organ involvement is essential. Previous reviews of machine learning (ML) in SLE combined heterogeneous tasks and validation strategies and may have overinterpreted model performance. OBJECTIVE: This study evaluated the diagnostic performance of ML and deep learning (DL) models for 3 clinically distinct SLE-related tasks: SLE classification or diagnosis, lupus nephritis (LN) diagnosis, and neuropsychiatric systemic lupus erythematosus (NPSLE) discrimination. We also assessed methodological quality and certainty of evidence. METHODS: PubMed, Embase, Cochrane Library, Web of Science, and IEEE Xplore were searched from January 2014 to April 2026. Eligible peer-reviewed diagnostic accuracy studies developed or validated ML or DL models for 1 of the 3 prespecified tasks, used an accepted reference standard, and provided data for a 2×2 contingency table. Bivariate random-effects meta-analyses with the Hartung-Knapp-Sidik-Jonkman adjustment were used to pool sensitivity and specificity. We reported 95% prediction intervals (PIs), assessed risk of bias using the Quality Assessment of Diagnostic Accuracy Studies for Artificial Intelligence tool (QUADAS-AI; Viknesh Sounderajah [Imperial College London]), and evaluated certainty of evidence using the Grading of Recommendations Assessment, Development, and Evaluation framework for diagnostic test accuracy. RESULTS: Twenty-nine studies were included: 17 for SLE classification, 5 for LN diagnosis, and 7 for NPSLE discrimination. In the primary task-stratified analysis, pooled sensitivity was 0.91 (95% CI 0.86-0.94; 95% PI 0.56-0.99), and pooled specificity was 0.94 (95% CI 0.91-0.96; 95% PI 0.69-0.99), with low heterogeneity (I²=23.9% and 22.9%, respectively). DL models showed a sensitivity of 0.93 and specificity of 0.95, compared with 0.88 and 0.94 for traditional ML models. Certainty of evidence was high for most analyses but low for LN diagnosis because of inconsistency and imprecision. All studies were retrospective, and only 9 of 29 (31%) performed independent external validation. Overall risk of bias was high or unclear in 22 of 29 (75.9%) studies. No study reported model calibration, decision-curve analysis, or net clinical benefit. CONCLUSIONS: ML models showed promising diagnostic accuracy across 3 distinct SLE-related tasks, but wide PIs, limited external validation, and pervasive risk of bias restrict conclusions about real-world generalizability. Prospective multicenter studies with standardized tasks and reference standards, independent external validation, and formal assessment of calibration and clinical utility are required before clinical implementation.

Humans

Diagnostic performance of machine learning models for malignant and non-malignant pleural effusion: Systematic review and meta-analysis.

BACKGROUND: Accurately distinguishing malignant pleural effusion (MPE) from non-malignant pleural effusion is clinically important, but the generalisability and methodological quality of machine-learning (ML) models remain uncertain. METHODS: We searched eight databases to 23 April 2026. Diagnostic performance was pooled using random-effects and Reitsma bivariate models, and study quality was assessed using PROBAST+AI. RESULTS: Forty-two studies were included; 17 contributed to the AUC meta-analysis and 14 to the bivariate analysis. The pooled AUC was 0.90 (95 % CI 0.85-0.94; 95 % prediction interval 0.62-0.98), with sensitivity of 0.80 (95 % CI 0.77-0.83) and specificity of 0.87 (95 % CI 0.79-0.92). Only nine studies reported external, temporal or independent validation. Externally validated studies had a lower pooled AUC than studies without external validation (0.83 vs 0.92), with lower specificity observed in the two externally validated studies contributing sensitivity and specificity data. All 42 development assessments had high overall quality concerns, and all 42 model evaluations were judged at high risk of bias. CONCLUSIONS: ML models showed good apparent accuracy for distinguishing MPE from non-MPE, but the evidence was limited by substantial heterogeneity, high risk of bias and scarce external validation. The pooled estimates reflect the average performance of different selected models rather than the expected accuracy of a single clinical test. ML models should be regarded as adjuncts to existing diagnostic pathways until they are confirmed by rigorous multicentre prospective external validation and clinical-impact studies.

Humans

Diagnostic performance of machine learning models versus established risk stratification for intracranial aneurysm rupture: a systematic review and bivariate meta-analysis.

BACKGROUND: Machine learning (ML) models have been proposed to improve the discrimination of intracranial aneurysm rupture status beyond established clinical risk stratification tools. However, reported performance is heterogeneous and the relative contribution of model architecture and feature dominance remains unclear. METHODS: We performed a Preferred Reporting Items for Systematic Reviews and Meta-Analyses-diagnostic test accuracy systematic review and diagnostic meta-analysis of studies evaluating ML models for intracranial aneurysm rupture discrimination. PubMed, Embase and CENTRAL were searched to February 2026. Sensitivity and specificity were pooled using a bivariate random-effects model, with summary receiver operating characteristic curves generated across training, internal testing and external validation datasets. Models were compared with regression-based approaches and Population, Hypertension, Age, Size of aneurysm, Earlier subarachnoid haemorrhage, Site of aneurysm (PHASES) scores. Subgroup and meta-regression analyses explored associations between algorithm family and feature domain. RESULTS: Sixty-two retrospective cohorts (29 709 patients 209 models) met the inclusion criteria. In training datasets, pooled sensitivity and specificity for ML were 0.81 (95% CI 0.75 to 0.85) and 0.83 (0.80-0.86), with an area under the curve (AUC) of 0.878, exceeding PHASES (AUC 0.667). In testing datasets, ML retained higher discrimination (AUC 0.837) than regression models (0.806) and PHASES (0.646). In external validation, sensitivity was preserved (0.82), but specificity declined (0.66). Deep learning demonstrated the highest AUCs (training and testing). Incorporation of haemodynamic or radiomic features improved pooled discrimination relative to morphology alone. Evidence of small-study effects and mostly unclear Prediction Model Risk Of Bias Assessment Tool ratings were observed. CONCLUSIONS: ML approaches demonstrate higher pooled discrimination for aneurysm rupture status than conventional risk scores in retrospective datasets, but reduced external validation specificity and heterogeneity limit confidence for clinical translation. Prospective, externally validated, calibrated models are required before integration into routine cerebrovascular risk stratification.

Humans

Multimodal artificial intelligence and machine learning in oncology: from data integration to precision cancer care.

Cancer remains a major global health burden, with approximately 20 million new cases and 9.7 million cancer-related deaths reported globally in 2022. While advances in radiological imaging, molecular profiling, and clinical data have enhanced the interpretation of disease progression, the availability of multiple such modalities still does not meet the needs of a large patient population. This narrative review focuses on the role of multimodal artificial intelligence and machine learning in bridging the gap in interpreting heterogeneous modalities to improve risk prediction, prognostic assessment, and treatment decision-making in precision oncology. Multimodal frameworks such as Pathomic Fusion illustrate how complementary histopathological and genomic information can be integrated for cancer diagnosis and prognostic modeling. Multimodal models have demonstrated potential in virtual biopsy, cancer screening, prognostic prediction, radiotherapy planning, intraoperative guidance, and clinical-trial design using digital twins and synthetic control arms. The major limitations of incorporating multimodal artificial intelligence and machine learning in oncology include data heterogeneity, demographic or institutional biases, and reproducibility challenges that hinder translation. Accordingly, appropriate data-governance strategies, fairness audits, and privacy-preserving approaches such as federated learning should be considered where appropriate. Future progress will depend on the development of standardized benchmarking datasets, robust external validation, seamless integration with electronic health records and picture archiving and communication systems, and the implementation of explainable, secure, and clinically validated multimodal artificial intelligence frameworks that support precision oncology in routine clinical practice.

deep learning

The effect of noise and biases on the performance of machine learning algorithms.

This paper describes the results of experiments with a machine learning algorithm for the induction of classification trees. We mainly address the impact of noise on the resulting classification tree and on the classification results obtained with the derived tree. We use the domain of the biochemical assessment of thyroid diseases as an example. Some suggestions for quality assessment are outlined that should be available in tools that assist users in deriving classification trees in noisy domains.

Algorithms

Identification of Biomarkers for Right Ventricular Dysfunction in Idiopathic Dilated Cardiomyopathy Via Urinary Proteomics and Machine Learning.

BACKGROUND: Right ventricular dysfunction (RVD) is a common complication of idiopathic dilated cardiomyopathy linked to poor outcomes. However, reliable noninvasive biomarkers for RVD remain lacking. This study aimed to identify urinary proteomic markers using mass spectrometry and machine learning. METHODS: In this prospective cohort, patients with idiopathic dilated cardiomyopathy were classified by cardiac magnetic resonance imaging into groups with RVD (RV ejection fraction <45%) and without RVD groups. Baseline urine samples were profiled by data-independent acquisition mass spectrometry. Differentially expressed proteins were identified and selected by least absolute shrinkage and selection operator regression to build a diagnostic model, developed in a training set, and validated in a test set. The primary end point was a composite of cardiovascular death, heart failure rehospitalization, left ventricular assist device implantation, or heart transplantation. RESULTS: The study enrolled 147 patients with idiopathic dilated cardiomyopathy (64 with RVD, 83 without), with a median follow-up of 19.3&#x2009;months. Of 3579 quantified urinary proteins, 46 were differentially expressed between groups. A 3-protein panel (RARRES1 [retinoic acid receptor responder protein 1], MVB12B [multivesicular body subunit 12B], GSK3A [glycogen synthase kinase 3 alpha]) was identified and showed excellent diagnostic accuracy (training area under the curve 0.946; validation area under the curve0.935), outperforming both NT-proBNP (N-terminal pro-brain natriuretic peptide) and tricuspid annular plane systolic excursion. The risk score derived from this panel effectively stratified patients, with the high-risk group exhibiting significantly worse outcomes than the low-risk group (hazard ratio, 3.24 [95% CI, 1.56-6.71], P=0.002). CONCLUSIONS: The urinary proteomic panel developed in this study demonstrates diagnostic and prognostic potential for identifying RVD in idiopathic dilated cardiomyopathy, providing a promising noninvasive tool for precise detection and clinical risk stratification.

Humans

Enhanced identification of key bacterial motility genes via a cross-species genomic hybrid feature machine learning approach.

Efficient and accurate identification of functional genes is critical to biological research, yet traditional single-species approaches are often limited by low efficiency. Previously, we established a novel method for identifying key genes using cross-species protein domain features and machine learning. However, the high multiplicity of gene members associated with specific domains creates a substantial workload for subsequent experimental validation. To address this, this study proposes an enhanced approach that integrates EggNOG-based protein sequence annotation with domain analysis. Unannotated sequences are subsequently analyzed for protein domains, generating a comprehensive "direct gene annotation plus domain" hybrid feature matrix. While the hybrid matrix model yielded comparable predictive accuracy, it significantly enhanced feature resolution: the top 50 predicted features were all known motility-related genes or domains. Furthermore, among the top 100 ranked features, 58 are confirmed to be directly related to motility based on experimental evidence. Although strict genus-level control still yielded 51 confirmed features, excessive taxonomic restriction drastically reduces the number of training genomes, which may paradoxically impair identification efficiency. These results demonstrate that the new method effectively reduces the subsequent experimental workload and enables high-throughput identification of functional genes in a single analysis. With accuracy and efficiency far exceeding those of existing single-species identification methods, it provides a highly efficient solution for mining key genes underlying other complex bacterial phenotypes.

Machine Learning

Multi-cohort integration and machine learning identify CPVL as a novel oncogenic driver in gastric cancer.

BACKGROUND: Gastric cancer (GC) remains a leading cause of cancer-related mortality worldwide, and the prognosis of advanced GC remains poor. Systematic identification of robust biomarkers through multi-cohort integration and computational prioritization may facilitate the discovery of novel therapeutic targets. AIM: To identify key genes associated with gastric cancer progression through integrative multi-omics analysis and to elucidate the biological functions and molecular mechanisms of the top-prioritized candidate gene. METHODS: Comprehensive bioinformatics analyses integrating The Cancer Genome Atlas (TCGA), Genotype-Tissue Expression (GTEx), and Gene Expression Omnibus (GEO) datasets were performed using differential expression analysis, weighted gene co-expression network analysis (WGCNA), Cox regression, and eight machine-learning algorithms to systematically identify and prioritize GC-associated hub genes. Among the identified candidates, CPVL was selected for further validation based on its diagnostic and prognostic performance. CPVL expression and clinical relevance were validated by independent datasets and immunohistochemistry. Lentiviral constructs were used to overexpress or silence CPVL in GC cell lines. Functional assays were performed, including CCK-8, colony formation, EdU incorporation, and flow cytometry, to assess cell proliferation and cell-cycle distribution. Western blotting and JAK2 inhibitor (AZD1480) rescue experiments were performed to elucidate the underlying mechanisms, and a nude mouse xenograft model was used to evaluate tumorigenicity in vivo. RESULTS: Multi-cohort screening identified five hub genes (CPVL, AADAC, BCAT1, CPXM1, and FBN1). Among them, CPVL exhibited the highest diagnostic accuracy (AUC&#x2009;=&#x2009;0.895) and the strongest correlation with poor overall survival, and was therefore selected for mechanistic investigation. CPVL expression was markedly upregulated in GC tissues and cell lines. Functional assays demonstrated that CPVL promotes GC cell proliferation and accelerates G1/S-phase transition. Mechanistically, CPVL activated the JAK2/STAT3 signaling pathway, upregulating Cyclin D1 and CDK4 while downregulating p27. Treatment with the JAK2 inhibitor AZD1480 partially reversed these effects. In vivo, CPVL knockdown significantly inhibited tumor growth. CONCLUSION: Through systematic multi-cohort integration and machine-learning prioritization, CPVL was identified as a novel oncogenic driver in gastric cancer. CPVL promotes tumor growth via activation of the JAK2/STAT3 pathway and regulation of the Cyclin D1/CDK4/p27 axis, highlighting its potential as a diagnostic biomarker and therapeutic target.

Biomarker

Uncovering encrypted antimicrobial peptides in health-associated Lactobacillaceae by large-scale genomics and machine learning.

BACKGROUND: Antimicrobial peptides (AMPs) are well known for their broad-spectrum activity and have shown great promise in addressing the antibiotic-resistant crisis. The Lactobacillaceae family, recognized for its health-promoting effects in humans, represents a valuable source of novel AMPs. However, the global prevalence and distribution of AMPs within Lactobacillaceae remains largely unknown, which limits the efficient discovery and development of novel AMPs. RESULTS: We analyzed all available genomes (10,327 genomes), encompassing 38 genera and 515 species, to investigate the biosynthetic potential (indicated by the number of AMP sequences in the genome) of AMP in the Lactobacillaceae family. We demonstrated Lactobacillaceae species had ubiquitous (69.90%) biosynthetic potential of AMPs. Overall, 9601 AMPs were identified, clustering into 2092 gene cluster families (GCFs), which showed strong interspecies specificity (95.27%), intraspecies heterogeneity (93.31%), and habitat uniqueness (95.83%), that greatly expanded on the AMP sequence landscape. Novelty assessment indicated that 1516 GCFs (72.47%) had no similarity to any known AMPs in existing databases. Machine learning predictions suggested that novel AMPs from Lactobacillaceae possessed strong antimicrobial potential, with 664 GCFs having an additive minimum inhibitory concentration (MIC) below 100&#xa0;&#x3bc;M. We randomly synthesized 16 AMPs (with predicted MIC&#x2009;<&#x2009;100&#xa0;&#x3bc;M) and identified 10 AMPs exhibiting varied-spectrum activity against 11 common pathogens. Finally, we identified one Lactobacillus delbrueckii-originated AMP (delbruin_1) having broad-spectrum (all 11 pathogens) and high antimicrobial activity (average MIC&#x2009;=&#x2009;38.56 &#xb5;M), which proved its potential as a clinically viable antimicrobial agent. CONCLUSIONS: We uncovered the global prevalence of AMPs in Lactobacillaceae and proved that Lactobacillaceae is an untapped and invaluable source of novel AMPs to combat the antibiotic-resistance crisis. Meanwhile, we provided a machine learning-guided framework for AMP discovery, offering a scalable roadmap for identifying novel AMPs not only in Lactobacillaceae but also in other organisms. Video Abstract.

Machine Learning

Machine learning for population-level risk prediction of future cholangiocarcinoma.

BACKGROUND: The poor prognosis of cholangiocarcinoma (CCA) is largely driven by rapid, asymptomatic disease progression, which usually results in a late diagnosis in the absence of established screening strategies. An early, cost-effective, and universally applicable risk assessment strategy would therefore be valuable. METHODS: We developed machine learning (ML) models on prospective, multimodal data from 487,495 UK Biobank (UKB) participants, of whom 649 developed CCA during follow-up. Data from England (80%) were utilised for ML development via five-fold cross-validation, and then all models were tested on withheld data from Scotland, Wales, and Newcastle (20%). Iterative ablation studies reduced inputs from >150 features across demographic data, lifestyle, health records, blood parameters, genomics, and metabolomics to models built on five and ten routinely available clinical parameters. These were externally validated in the Penn Medicine Biobank (PMBB; n = 2638; 28 CCA), All of Us Research Program (AOU; n = 330,433; 362 CCA), Japan Medical Data Centre Claims Database (JMDC; n = 8,425,522; 723 CCA) and TriNetX (n = 728,886; 1592 CCA). FINDINGS: We show that ML models integrating biliary-disease associated health records and Gamma glutamyltransferase can stratify risk of future CCA. Evaluation on the UKB test set as well as three independent cohorts revealed robust performance and generalisability across ethnicities. We achieved AUROCs of 0.71 [95% CI: 0.703-0.711], 0.77 [95% CI: 0.764-0.778 ], 0.796 [95% CI: 0.795-0.798] and 0.8 [95% CI: 0.794-0.805] for UKB, PMBB, AOU, and JMDC respectively, with respective AUPRCs of 0.014 [95% CI: 0.009-0.018], 0.042 [95% CI: 0.037-0.048], 0.038 [95% CI: 0.033-0.042] and 0.001 [95% CI: 0.001-0.001]. In AOU, application of the Youden J-optimised threshold yielded a number needed to screen of 79. Separate models for intra- and extrahepatic CCA did not improve performance. In line with the pathophysiology, performance declined for longer intervals between assessment and event. A group-level analysis in the TriNetX cohort revealed hazard ratios of up to 82.5 [95% CI: 26.4-257.96]. We provide extensive interpretability results and release all source codes used to develop the presented models. INTERPRETATION: We provide a comprehensive framework for early CCA risk stratification in the general population, identifying key predictors, and demonstrating the potential of data-driven models in personalised screening for hepatobiliary cancer. FUNDING: German Cancer Aid (grant #70115730), Junior Principal Investigator Fellowship programme of RWTH Aachen Excellence strategy.

Humans

Fishing for a reelGene: evaluating gene models with evolution and machine learning.

Assembled genomes and their associated annotations have transformed our study of gene function. However, each new annotated assembly generates new gene models. Inconsistencies between annotations likely arise from biological and technical causes, including pseudogene misclassification, transposon activity, and intron retention from sequencing of unspliced transcripts. To evaluate gene model predictions, we developed reelGene, a pipeline of machine learning models focused on (1) transcription boundaries, (2) mRNA integrity, and (3) protein structure. The first two models leverage sequence characteristics and evolutionary conservation across related taxa to learn the grammar of conserved transcription boundaries and mRNA sequences, while the third uses the conserved evolutionary grammar of protein sequences to predict whether a gene can produce a protein. Evaluating 1.8 million transcript models in Zea mays ssp. mays (maize), reelGene classified 28% as incorrectly annotated or non-functional. We find that reelGene classifies 92.2% of genes in the maize proteome and 99.2% of genes within the maize classical gene list as functional. reelGene also provides a way to further investigate genome biology- for instance, reelGene indicates that 10.3% of dispensable genes in B73 are functional, and within retained duplicate genes, reelGene identifies a 30% bias toward the retention of the M1 subgenome when one copy is functional and the other is non-functional. As an annotation-evaluating tool, reelGene is directly applicable to species of the Andropogoneae tribe, including other important crops like sorghum and miscanthus. As a community resource, reelGene has been integrated onto MaizeGDB both as a browser track and as an individual Shiny App, allowing researchers to evaluate gene model accuracy and further investigate genome biology.

Machine Learning

Transcriptome Analysis, Machine Learning, and Experimental Identification of CDK7 Affecting the Progression of Pregnancy-induced Hypertension by Influencing Macrophage Polarization.

INTRODUCTION: Pregnancy-induced hypertension (PIH) is a severe pregnancy complication characterized by placental insufficiency, abnormal vascular remodeling, and immune dysregulation, but personalized therapeutic markers remain unclear. This study aimed to identify key genes and explore immune mechanisms in PIH using transcriptome analysis, machine learning, and experimental validation. METHODS: We analyzed the GSE204835 transcriptomic dataset to screen differentially expressed genes (DEGs) and performed Gene Ontology (GO), Kyoto Encyclopedia of Genes and Genomes (KEGG), Reactome, and Gene Set Enrichment Analysis (GSEA) for functional annotation. Immune infiltration analysis was also performed to examine the immune landscape in PIH. Least Absolute Shrinkage and Selection Operator (LASSO) regression identified key genes, which were validated in a PIH cell model. Flow cytometry and immunofluorescence assays assessed the effect of CDK7 knockdown on macrophage polarization. RESULTS: A total of 1,598 DEGs (1,123 upregulated, 475 downregulated) were identified. Enrichment analyses highlighted associations with embryonic organ development, oxidative phosphorylation, angiogenesis, and oxidative stress. Immune infiltration analysis revealed altered eosinophil and macrophage polarization in PIH. LASSO regression selected 12 key genes, with CDK7 showing the most significant upregulation in the PIH model. CDK7 knockdown promoted macrophage polarization toward the anti-inflammatory M2 phenotype. DISCUSSION: These findings link CDK7 to immune dysregulation in PIH by modulating macrophage polarization, expanding our understanding of PIH's molecular mechanisms. The study's limitations include reliance on public datasets and in vitro models, warranting in vivo validation. CONCLUSION: CDK7 emerges as a potential therapeutic target for PIH, offering new insights into immunoregulatory interventions for this complication.

Female

Screening of the key single nucleotide polymorphisms in type 2 diabetes mellitus complicated with lower extremity arterial disease by machine learning.

OBJECTIVES: Diabetic lower extremity arterial disease (LEAD) is a manifestation of diabetic lower extremity vascular complications. This study aimed to screen the key single nucleotide polymorphism (SNP) gene signature in patients with type 2 diabetes mellitus (T2DM) and LEAD. METHODS: A total of 147 patients with T2DM complicated by LEAD and 144 patients with T2DM without LEAD were enrolled for transcriptome sequencing. The Plink software was used to preprocess the data. Five machine learning methods were adopted to build the SNP diagnosis models. The receiver operating characteristic (ROC) curve was used to quantify the predicted probabilities of the model. Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses were performed using the cluster Profiler package. Finally, regression statistical analysis was used to correlate the key SNPs with clinical information and biochemical indicators. RESULTS: A total of 24 SNPs were retained and 10 SNPs were risk allele genes. Nine SNPs (rs7412, rs1800629, rs699947, rs3918242, rs668, rs1800470, rs1800449, rs1800469, and rs1024611) were identified as the key SNPs sites. GO and KEGG pathway analyses revealed that these genes are mainly enriched in fluid shear stress and atherosclerosis. Finally, rs1800449 was associated with low-density lipoprotein cholesterol (LDL-C). With high density lipoprotein cholesterol (HDL-C), related site was rs1024611. The sites associated with total cholesterol (CHOL) were rs1800449 and rs7412.The site associated with apolipoprotein B (APOB) and apolipoprotein A1 (APOA1) were rs1800470 and rs1800469. CONCLUSION: This study authenticated nine SNPs for the diagnosis of T2DM patients with LEAD, which will be of great significance in the development of diagnostic molecular biomarkers for T2DM patients.

Humans

Integrative multi-omics and machine learning identify the SPI1-METTL16-PLIN4 axis as a candidate driver of steatosis in HepG2 cells.

BACKGROUND: Non-alcoholic fatty liver disease (NAFLD) is a prevalent metabolic disorder with limited therapeutic options. This study aimed to identify potential regulators and explore their functional roles in a cellular model of NAFLD. METHODS: WGCNA was performed on the hepatic transcriptomic dataset GSE126848 (31 NAFLD vs. 26 controls), followed by integration with serum proteomic data from 12 NAFLD patients and 12 healthy controls. Hub genes were prioritized using three machine learning algorithms. Functional validation was conducted in a HepG2 cellular steatosis model induced by high fructose (3.2&#x202f;g/L) and oleic acid (400&#x202f;&#x3bc;M) for 48&#x202f;h. Lipid accumulation was assessed by Oil Red O staining and triglyceride/total cholesterol measurement. Inflammation was evaluated by TNF-&#x3b1; and IL-6 secretion (ELISA), and oxidative stress by ROS levels (flow cytometry). The binding interaction between METTL16 and PLIN4 mRNA was validated by RNA immunoprecipitation (RIP)-quantitative PCR. METTL16-mediated m6A modification of PLIN4 was assessed by Methylated RIP (MeRIP)-quantitative PCR. Transcriptional regulation of METTL16 by SPI1 was examined by chromatin immunoprecipitation (ChIP) and dual-luciferase reporter assays. RESULTS: Integrative analysis identified PLIN4 as a core hub gene. PLIN4 was upregulated in the HepG2 steatosis model (P&#x202f;<&#x202f;0.001). PLIN4 knockdown alleviated lipid droplet accumulation (P&#x202f;<&#x202f;0.001), reduced TNF-&#x3b1; and IL-6 secretion (P&#x202f;<&#x202f;0.01), and decreased ROS levels (P&#x202f;<&#x202f;0.001) in fructose/oleic acid-treated HepG2 cells. Mechanistically, METTL16 mediated its m6A modification to enhance PLIN4 mRNA stability. Furthermore, SPI1 was found to transcriptionally activate METTL16 by binding to its promoter (P&#x202f;<&#x202f;0.001). PLIN4 re-expression partially reversed the protective effects of SPI1 knockdown on lipid accumulation (P&#x202f;=&#x202f;0.01), inflammation (P&#x202f;<&#x202f;0.05), and oxidative stress (P&#x202f;<&#x202f;0.001). CONCLUSION: This study identifies the SPI1/METTL16/PLIN4 axis as a potential regulatory mechanism contributing to in vitro steatosis, inflammation, and oxidative stress in steatotic HepG2 cells.

Humans

Machine Learning in Hyperlipidaemia Research: Screening and Experimental Insights into Lipid Metabolism Modulators.

Hyperlipidemia, characterized by elevated blood lipid levels, represents a major global health concern due to its strong association with cardiovascular disease, diabetes, and metabolic syndrome. While current therapies - such as statins, fibrates, bile acid sequestrants, and PCSK9 inhibitors - are effective in controlling hyperlipidemia, they are often associated with adverse effects, potential drug resistance, and suboptimal efficacy in certain patient populations. All of the above underscore the urgent need for safer and more effective therapeutic alternatives. Among the major molecular targets involved in the regulation of lipid metabolism are HMG-CoA reductase, PCSK9, peroxisome proliferator-activated receptors (PPARs), cholesteryl ester transfer protein (CETP), and nuclear receptors, including the liver X receptor (LXR) and farnesoid X receptor (FXR), which are also targets for future antihyperlipidemic drug development. Recent advancements in artificial intelligence (AI) and machine learning (ML) have significantly transformed and accelerated drug discovery by enabling the processing of vast amounts of genomic, proteomic, and chemical data. Furthermore, ML tools such as quantitative structure-activity relationship (QSAR) modelling, deep learning, random forest, and support vector machines (SVM) have proven predictive and effective in identifying novel lipid metabolism modulators, thereby enhancing the efficacy and accuracy of virtual screening. Meanwhile, molecular docking has become an integral part of structure-based drug design (SBDD), and software such as AutoDock, Glide, and GOLD have proven effective in generating accurate ligand-target docking models. Molecular docking, together with ML-based approaches, enables the identification of potent and selective drug candidates. Overall, the combination of ML and molecular docking offers an efficient and accurate platform for antihyperlipidemic drug discovery, helping to overcome the limitations of currently available therapeutic strategies.

HMG-CoA reductase

An Integrated Machine Learning and Genomic Framework for Precise Detection of Gastric Cancer.

This study presents a novel integrative approach for the analysis of high-dimensional gene expression data, leveraging the complementary strengths of unsupervised clustering and supervised classification. Using K-means clustering, the data set is stratified into three distinct clusters, revealing intrinsic biological patterns and relationships. The resulting cluster assignments are subsequently used as pseudolabels to train machine learning models, including support vector machines, random forest, and a stacking ensemble classifier. To validate and enhance the robustness of clustering, complementary methods, such as hierarchical clustering and density-based spatial clustering of applications with noise (DBSCAN), are used, with results visualized through principal component analysis-driven dimensionality reduction. The high predictive accuracy achieved by the classifiers underlines the separability and reliability of the identified clusters. Furthermore, feature importance analysis highlighted key genetic determinants within each cluster, offering actionable insights into potential biomarkers and critical genomic features. This framework bridges the gap between exploratory unsupervised learning and predictive supervised modeling, providing a scalable and interpretable method for analyzing complex genomic data sets. Its applicability extends to biomarker discovery, patient stratification, and other precision medicine applications, emphasizing its utility in advancing genomic research and clinical practice.

Humans

Pan-cancer multi-omics machine learning defines a lactylation-associated immune-excluded tumor state with proteomic and experimental corroboration.

BACKGROUND: Histone lactylation links lactate metabolism to chromatin regulation, but whether lactylation-program-associated transcriptional patterns delineate recurrent pan-cancer tumor states remains unclear. METHODS: We integrated mRNA, lncRNA, and miRNA profiles from 9712 TCGA tumors across 33 cancer types with GTEx references, six GEO cohorts, IMvigor210, and an institutional clear-cell renal cell carcinoma (ccRCC) cohort used for exploratory DIA-NN proteomic corroboration. Random-effects co-expression meta-analysis, multi-omics consensus clustering, regulon inference, immune deconvolution, TIDE, oncoPredict, and SHAP-based machine learning were applied. hsa-miR-431-5p was functionally evaluated as a proof-of-concept CS2-associated miRNA in bladder cancer models. RESULTS: LacCoEx-Atlas comprised 398,491 lactylation-related co-expression pairs across 24,667 RNA features under a random-effects framework (median I&#xb2; = 88.6%). Consensus clustering identified two subtypes: CS2 showed glycolytic-mesenchymal-immune-excluded features, M2 macrophage enrichment, CD8&#x207a; T-cell depletion, elevated HDAC4/NSD3/KDM6B activity, and worse survival, whereas CS1 showed oxidative, sirtuin-active programs. CS2 had fewer predicted ICI responders (18.3% vs. 52.0%) and a lower observed ORR in IMvigor210 (15.3% vs. 24.0%). oncoPredict identified NU7441 as a hypothesis-generating CS2-associated sensitivity signal (Hedges' g = 1.17). DIA-NN proteomics in 50 ccRCC specimens provided exploratory support for CS2-associated hypoxia, ECM degradation, and metastasis programs. The 10-feature mRNA LARItools model achieved an apparent AUC of 0.9413, while a separate multi-omics model achieved 0.971; neither was independently validated. LARItools reproduced prognostic separation across six GEO cohorts. miR-431-5p promoted malignant phenotypes and EMT in bladder cancer cells, with concordant CMU4h expression findings. CONCLUSIONS: Lactylation-program-associated transcriptional patterns delineate a recurrent immune-excluded pan-cancer tumor state associated with adverse prognosis, reduced predicted immunotherapy responsiveness, exploratory single-cancer protein-level support, and testable DNA damage response-targeting hypotheses. LacCoEx-Atlas and LARItools provide open resources for lactylation-program-associated tumor-state stratification and future translational research.

Humans