PubMed HealthSearch

SEARCH · PubMed Health

Results for “machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Clinical Variable-Based Machine Learning for Predicting Early mCRPC Using Exclusively Clinical Variables: Development and Multicenter External Validation.

BACKGROUND AND OBJECTIVE: Metastatic hormone-sensitive prostate cancer (mHSPC) exhibits heterogeneous progression patterns, with early progression to metastatic castration-resistant prostate cancer (mCRPC) within 12 months indicating aggressive tumor biology and poor prognosis. Current risk stratification tools (CHAARTED, LATITUDE) offer limited individualized prediction. Machine learning approaches are increasingly applied to predict prostate cancer progression, but most models show modest performance (AUC 0.68-0.72), limited external validation, or require genomic variables unavailable in routine practice. This study aimed to develop and externally validate a novel RINH algorithm for predicting early mCRPC progression (≤ 12 months) using exclusively clinical variables, positioning it as a superior alternative to conventional ML classifiers. METHODS: This multicenter study enrolled 412 patients with de novo mHSPC from seven Spanish academic centers using mixed retrospective-prospective data collection. Twenty clinical variables were recorded, including demographics, PSA, ISUP grade, metastatic localization, CHAARTED/LATITUDE classifications, and treatment modalities. Following RINH-based outlier exclusion (55 patients), 357 patients (29 with early progression, 8.1%) were used to train six ML algorithms: RINH, Logistic Regression, Linear Discriminant, Support Vector Machine, Random Forest, and Subspace Discriminant. A two-tiered validation strategy integrated stratified fivefold cross-validation across all centers and formal external validation using center 1 (n = 121, 19 events) for training and centers 2-7 (n = 207, 10 events) for independent testing. Performance metrics included AUC, sensitivity, specificity, accuracy, and F1-score. KEY FINDINGS AND LIMITATIONS: Artificial intelligence and machine learning (ML) are transforming oncology, promising personalized risk stratification beyond traditional clinical criteria. In metastatic hormone-sensitive prostate cancer (mHSPC), early progression to castration resistance (mCRPC) within 12 months signals aggressive biology and poor prognosis, yet current tools (CHAARTED, LATITUDE) offer limited individualized prediction. Multiple ML models have been proposed with variable success: most achieve modest performance (AUC 0.68-0.72), lack robust external validation, or rely on genomic variables inaccessible in routine practice. We propose a novel approach using the Rivality Index Neighborhood (RINH) algorithm, demonstrating superior predictive capacity in an initial multicenter validation with exclusively clinical variables. This study provides rigorous multicenter external validation, advancing toward implementable precision oncology tools. CONCLUSIONS AND CLINICAL IMPLICATIONS: The RINH algorithm achieves superior predictive performance for early mCRPC progression using exclusively clinical variables, representing a significant advance toward implementable risk stratification. However, low reliability scores in external validation underscore that excellent performance metrics alone do not guarantee stability. Before clinical deployment, validation in substantially larger cohorts with higher progression events is essential. If validated, this model could enable personalized, risk-adapted therapeutic strategies, refining patient selection for treatment intensification or de-escalation.

Humans

Machine learning detection of heteroresistance in Escherichia coli.

BACKGROUND: Heteroresistance (HR) is a significant type of antibiotic resistance observed for several bacterial species and antibiotic classes where a susceptible main population contains small subpopulations of resistant cells. Mathematical models, animal experiments and clinical studies associate HR with treatment failure. Currently used susceptibility tests do not detect heteroresistance reliably, which can result in misclassification of heteroresistant isolates as susceptible which might lead to treatment failure. Here we examined if whole genome sequence (WGS) data and machine learning (ML) can be used to detect bacterial HR. METHODS: We classified 467 Escherichia coli clinical isolates as HR or non-HR to the often used β-lactam/inhibitor combination piperacillin-tazobactam using pre-screening and Population Analysis Profiling tests. We sequenced the isolates, assembled the whole genomes and created a set of predictors based on current knowledge of HR mechanisms. Then we trained several machine learning models on 80% of this data set aiming to detect HR isolates. We compared performance of the best ML models on the remaining 20% of the data set with a baseline model based solely on the presence of β-lactamase genes. Furthermore, we sequenced the resistant sub-populations in order to analyse the genetic mechanisms underlying HR. FINDINGS: The best ML model achieved 100% sensitivity and 84.6% specificity, outperforming the baseline model. The strongest predictors of HR were the total number of β-lactamase genes, β-lactamase gene variants and presence of IS elements flanking them. Genetic analysis of HR strains confirmed that HR is caused by an increased copy number of resistance genes via gene amplification or plasmid copy number increase. This aligns with the ML model's findings, reinforcing the hypothesis that this mechanism underlies HR in Gram-negative bacteria. INTERPRETATION: We demonstrate that a combination of WGS and ML can identify HR in bacteria with perfect sensitivity and high specificity. This improved detection would allow for better-informed treatment decisions and potentially reduce the occurrence of treatment failures associated with HR. FUNDING: Funding provided to DIA from the Swedish Research Council (2021-02091) and NIH (1U19AI158080-01).

Machine Learning

Identifying JAK2 and ANXA5 as Key Genes Linking Obstructive Sleep Apnea and Oxidative Stress via Machine Learning and Multilayer Transcriptomic Integration With Functional Validation.

Obstructive sleep apnea (OSA) is a common and severe sleep disorder closely associated with oxidative stress (OS). This study aims to identify and validate potential OS-related genes associated with OSA through bioinformatics methods. We successfully identified OS-related differentially expressed genes (OS-DEGs) by combining the limma test, weighted correlation network analysis (WGCNA), and OS-related genes from the GeneCards database. Key genes and potential biological roles were further identified using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG), enrichment analysis, protein-protein interaction (PPI) network analysis, Lasso regression analysis, random forest algorithm, and support vector machine recursive feature elimination (SVM-RFE) method. Evaluate and validate the accuracy of key genes through receiver operating characteristic (ROC) curve analysis. The human single-cell RNA sequencing (scRNA-seq) dataset is used for cell classification annotation, analysis of key gene single-cell expression profiles, and virtual gene knockout experiments based on the scTenifoldKnk algorithm. Integrating scRNA-seq sequencing, pseudotime trajectory inference, cell-cell communication analysis, and bulk immune infiltration deconvolution reveals monocyte subtype remodeling in OSA. Finally, the expression levels of key genes in clinical samples were validated using real-time quantitative PCR (RT-qPCR) and Western blotting. A total of 57 common DEGs, indicating significant enrichment in OS, inflammation, and tumor pathways, particularly prominent in the immunometabolism pathway. By integrating DEGs, WGCNA, PPI results, and machine learning methods, key genes Janus kinase 2 (JAK2) and ANXA5 were screened out. JAK2 was significantly upregulated under disease conditions, while ANXA5 was significantly downregulated. ROC curve exhibited high accuracy (area under the curve [AUC] > 0.85). Human scRNA-seq analysis revealed that key genes were predominantly highly expressed in monocytes. Virtual knockout experiments demonstrated that these key genes play a crucial role in regulating immune responses and inflammatory reactions. PPI networks and enrichment analysis verified that downstream genes S100P, ALOX5AP, PROK2, and PADI4 may collaboratively participate in immune response and inflammation regulation. Finally, clinical sample experiment further validated the results of bioinformatics analysis. This study provides new research insights for the diagnosis, mechanism research, and treatment development of OSA in the future by integrating multilayer transcriptomic and machine learning techniques.

Humans

Opportunities for machine learning to predict cross-neutralization in FMDV serotype O.

Accurately estimating cross-neutralization between serotype O foot-and-mouth disease viruses (FMDVs) is critical for guiding vaccine selection and disease management. In this study, we developed a machine learning approach to estimate r1 values-an established measure of antigenic similarity-using VP1 sequence data and published virus neutralization titer (VNT) results. Our dataset comprised 108 serum-virus pairs representing 73 distinct FMDV strains. We applied Boruta feature selection and random forest classifiers, optimizing model performance through tenfold cross-validation and sub-sampling to address class imbalance. Predictors included pairwise amino acid distances, site-specific polymorphisms, and differences in potential N-glycosylation sites. Using a 0.3 r1 threshold to define cross-neutralization, the final model achieved high accuracy (0.96), sensitivity (0.93), and specificity (0.96) in training, and performed robustly on independent test sets - accuracy was 0.75 (95% CI 0.60 and 0.90), F1 score 0.86% and PPV 0.77. Importantly, key VP1 residues-positions 48, 100, 135, 150, and 151-emerged as strong predictors of antigenic relationships. Our results demonstrate the utility of integrating routinely generated genomic data with machine learning to inform vaccine candidate selection and anticipate immune interactions among circulating FMDV strains. This approach offers a practical tool for accelerating vaccine decision-making and can be adapted to other FMDV serotypes. The latest version of the r1 predictive model is available for access via a Shiny dashboard (https://dmakau.shinyapps.io/PredImmune-FMD/).

Foot-and-Mouth Disease Virus

Drug design by machine learning: the use of inductive logic programming to model the structure-activity relationships of trimethoprim analogues binding to dihydrofolate reductase.

The machine learning program GOLEM from the field of inductive logic programming was applied to the drug design problem of modeling structure-activity relationships. The training data for the program were 44 trimethoprim analogues and their observed inhibition of Escherichia coli dihydrofolate reductase. A further 11 compounds were used as unseen test data. GOLEM obtained rules that were statistically more accurate on the training data and also better on the test data than a Hansch linear regression model. Importantly machine learning yields understandable rules that characterized the chemistry of favored inhibitors in terms of polarity, flexibility, and hydrogen-bonding character. These rules agree with the stereochemistry of the interaction observed crystallographically.

Artificial Intelligence

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning

Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning.

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

Escherichia coli

Distributed machine learning: scaling up with coarse-grained parallelism.

Machine learning methods are becoming accepted as additions to the biologists data-analysis tool kit. However, scaling these techniques up to large data sets, such as those in biological and medical domains, is problematic in terms of both the required computational search effort and required memory (and the detrimental effects of excessive swapping). Our approach to tackling the problem of scaling up to large datasets is to take advantage of the ubiquitous workstation networks that are generally available in scientific and engineering environments. This paper introduces the notion of the invariant-partitioning property--that for certain evaluation criteria it is possible to partition a data set across multiple processors such that any rule that is satisfactory over the entire data set will also be satisfactory on at least one subset. In addition, by taking advantage of cooperation through interprocess communication, it is possible to build distributed learning algorithms such that only rules that are satisfactory over the entire data set will be learned. We describe a distributed learning system, CorPRL, that takes advantage of the invariant-partitioning property to learn from very large data sets, and present results demonstrating CorPRL's effectiveness in analyzing data from two databases.

Database Management Systems

Enhancing detection of polygenic adaptation: a comparative study of machine learning and statistical approaches using simulated evolve-and-resequence data.

BACKGROUND: Detecting signals of polygenic adaptation remains a significant challenge in population genomics, as traditional methods often struggle to identify the associated subtle, multi-locus allele-frequency shifts. Here, we introduced and tested several novel approaches combining machine learning techniques with traditional statistical tests to detect polygenic adaptation patterns in time-series of allele frequency changes from whole genome data. We implemented a Naive Bayesian Classifier (NBC) and One-Class Support Vector Machines (OCSVM), and compared their performance against the classical Fisher's Exact Test (FET). Furthermore, we combined machine learning and statistical models (OCSVM-FET and NBC-FET), resulting in 5 competing approaches. The framework is mainly designed and validated for evolve-and-resequence (EaR) experimental designs, where defined selection pressures and temporal sampling are feasible, but might be applicable for certain natural experiments as well. RESULTS: Using a simulated dataset based on empirical C. riparius Pool-Seq data, we evaluated methods across evolutionary scenarios varying in generation, selection strength, and number of loci under selection. Our results demonstrate that the combined OCSVM-FET approach consistently outperformed competing methods, achieving the lowest false positive rate, highest area under the curve, and high accuracy. The performance peak aligned with what we term the 'late dynamic phase' of adaptation - the period after initial selection has occurred but before fixation - highlighting the method's sensitivity to ongoing selective processes. CONCLUSIONS: Furthermore, we emphasize the critical role of parameter tuning, balancing biological assumptions with methodological rigor. While broader applicability remains an important direction for future work, the present benchmarking is intentionally scoped to EaR experimental contexts.

Machine Learning

A machine learning model and identification of immune infiltration for chronic obstructive pulmonary disease based on disulfidptosis-related genes.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a chronic and progressive lung disease. Disulfidptosis-related genes (DRGs) may be involved in the pathogenesis of COPD. From the perspective of predictive, preventive, and personalized medicine (PPPM), clarifying the role of disulfidptosis in the development of COPD could provide a opportunity for primary prediction, targeted prevention, and personalized treatment of the disease. METHODS: We analyzed the expression profiles of DRGs and immune cell infiltration in COPD patients by using the GSE38974 dataset. According to the DRGs, molecular clusters and related immune cell infiltration levels were explored in individuals with COPD. Next, co-expression modules and cluster-specific differentially expressed genes were identified by the Weighted Gene Co-expression Network Analysis (WGCNA). Comparing the performance of the random forest (RF), support vector machine (SVM), generalized linear model (GLM), and eXtreme Gradient Boosting (XGB), we constructed the ptimal machine learning model. RESULTS: DE-DRGs, differential immune cells and two clusters were identified. Notable difference in DRGs, immune cell populations, biological processes, and pathway behaviors were noted among the two clusters. Besides, significant differences in DRGs, immune cells, biological functions, and pathway activities were observed between the two clusters.A nomogram was created to aid in the practical application of clinical procedures. The SVM model achieved the best results in differentiating COPD patients across various clusters. Following that, we identified the top five genes as predictor genes via SVM model. These five genes related to the model were strongly linked to traits of the individuals with COPD. CONCLUSION: Our study demonstrated the relationship between disulfidptosis and COPD and established an optimal machine-learning model to evaluate the subtypes and traits of COPD. DRGs serve as a target for future predictive diagnostics, targeted prevention, and individualized therapy in COPD, facilitating the transition from reactive medical services to PPPM in the management of the disease.

Pulmonary Disease, Chronic Obstructive

How advances in machine learning drive early detection and risk prediction of early-onset colorectal cancer.

Early-onset colorectal cancer (EOCRC), defined as colorectal cancer diagnosed before age 50, is rising across high- and middle-income settings whilst organised screening stays anchored to older age thresholds. Blood-based liquid biopsy, combined with machine learning, is the most plausible route to early detection in this group because it does not depend on bowel preparation, endoscopy capacity, or adherence to stool-based testing. The gap is structural: incidence climbs fastest in the population below the age at which any guideline-endorsed modality is offered. The analytical challenge is that early-stage tumour-derived signals in plasma are low in abundance and distributed across heterogeneous molecular layers: circulating tumour DNA mutations, aberrant methylation, cfDNA fragmentomics, and small non-coding RNA. Machine learning converts these into a single calibrated probability. This review examines where artificial intelligence (AI)-driven liquid biopsy genuinely adds diagnostic value in EOCRC, distinguishes components in which learned models are decorative from those in which they are mechanistically necessary, and identifies the validation deficit separating research cohorts from deployable clinical tools. It summarises the first-generation tools used clinically for early detection and post-treatment monitoring, then considers analytes from exosome-bound microRNAs to long-read whole-genome sequencing of circulating plasma DNA, which reads cytosine modification natively, resolves methylation and fragmentation on single molecules, and characterises structural events short reads cannot anchor. Any analyte can feed a learned model, but more diverse input yields better discrimination. The central argument is that approved, guideline-included blood tests were validated in populations aged 45 and above, and their performance in younger patients cannot be assumed.

cfDNA fragmentomics

Causal associations between hormone replacement therapy and brain structure: Evidence from large-scale Mendelian randomization and double machine learning.

BACKGROUND: Hormone replacement therapy (HRT) is widely prescribed for the management of hormone deficiency, particularly during menopause, yet its causal effects on human brain structure remain incompletely understood. Observational studies have reported heterogeneous associations, underscoring the need for robust causal inference. METHODS: We applied an integrated causal framework combining two-sample Mendelian Randomization (MR) and Double Machine Learning (DML) to evaluate the effects of four HRT-related exposures-age at initiation, age at cessation, ever-use of HRT, and a composite medication-based phenotype-on 1366 brain imaging-derived phenotypes from the UK Biobank. Genetic instruments were derived from large-scale GWAS summary statistics, and causal estimates were validated using non-parametric DML models with cross-fitting and performance evaluation. RESULTS: Genetic instruments for age at HRT initiation, age at cessation, and ever-use of HRT were strong (median F-statistics 16.29-36.66). MR analyses identified a causal association between later initiation of HRT and lower orientation dispersion in the right inferior cerebellar peduncle (ubm-a-542; primary finding, no pleiotropy detected). An additional association with the left tapetum FA (ubm-a-243) was identified but exhibited significant directional horizontal pleiotropy (MR-Egger intercept P = 0.001) and is excluded from primary conclusions (Supplementary Note S2). Later cessation of HRT was associated with increased cortical thickness in the left middle occipital gyrus, reduced surface area in the left frontopolar cortex, and increased orientation dispersion in the splenium of the corpus callosum. Ever-use of HRT was causally linked to larger volumes of the right inferior frontal gyrus and right nucleus accumbens. These associations were corroborated by independent DML validation, which provided causally debiased estimates robust to high-dimensional confounding. Results for ukb-b-8080 (median F = 1.45) are provided in Supplementary Note S1 only; weak-instrument bias precludes causal inference. CONCLUSIONS: This study provides genetic-instrument-based and machine-learning-validated evidence for causal associations between HRT exposure-particularly its timing and lifetime use-and specific features of human brain structure, including white-matter microarchitecture, cortical thickness, and regional brain volume. These findings are FDR-controlled within exposures and independently replicated by DML, but require replication in external neuroimaging GWAS cohorts to establish definitive causal conclusions. They highlight the neurobiological relevance of sex steroid exposure and inform future research on brain aging and personalized hormone-based interventions.

Humans

Development and evaluation of a machine learning model for osteoporosis risk prediction in Korean women.

BACKGROUND: The aim of this study was to develop a machine learning (ML) model for classifying osteoporosis in Korean women based on a large-scale population cohort study. This study also aimed to assess ML model performance compared with traditional osteoporosis screening tools. Furthermore, this study aimed to examine the factors influencing the risk of osteoporosis through variable importance. METHODS: Data was collected from 4199 women aged 40-69 years in the baseline survey of the Ansan and Ansung cohort of the Korean Genome and Epidemiology Study. Osteoporosis was set as the dependent variable to develop ML classification models. Independent variables included 122 factors related to osteoporosis risk, such as socio-demographic characteristics, anthropometric parameters, lifestyle factors, reproductive factors, nutrient intakes, diet quality indices, medical history, medication history, family history, biochemical parameters, and genetic factors. The six classification models were developed using ML techniques, including decision tree, random forest, multilayer perceptron, support vector machine, light gradient boosting machine, and extreme gradient boosting (XGBoost). The six ML classification models were compared with two traditional osteoporosis screening tools, including the osteoporosis risk assessment instrument (ORAI) and the osteoporosis self-assessment tool (OST). The ML model performances were evaluated and compared using the confusion matrix and area under the curve (AUC) metrics. Variable importance was assessed using the XGBoost technique to investigate osteoporosis risk factors. RESULTS: The XGBoost model showed the highest performance out of the six ML classification models, with an accuracy of 0.705, precision of 0.664, recall of 0.830, and F1 score of 0.738. Moreover, the XGBoost model showed a higher performance on AUC than ORAI and OST. Variable importance scores were identified for 69 out of the 122 variables associated with osteoporosis risk factors. Age at menopause ranked first in variable importance. Variables of arthritis, physical activities, hypertension, education level, income level; alcohol intake, potassium intake, homeostatic model assessment for insulin resistance; energy intake, vitamin C intake, gout; and dietary inflammatory index ranked in the top 20 out of the 69 variables, using the XGBoost technique. CONCLUSIONS: This study found that an XGBoost model can be utilized to classify osteoporosis in Korean women. Age at menopause is a significant factor in osteoporosis risk, followed by arthritis, physical activities, hypertension, and education level.

Humans

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans

Integrated transcriptome analysis and machine learning to construct a homeostatic model of acetylation for bladder cancer and validate the key gene CES1.

BACKGROUND: Bladder cancer (BLCA) is one of the most common malignant tumors of the urinary system. Protein acetylation (PA) plays a critical role in regulating multiple biological processes (BPs), cellular homeostasis, and cancer-related signaling pathways. This study aimed to construct a homeostatic model of acetylation for BLCA using integrated transcriptome analysis and machine learning and to validate the key gene CES1. METHODS: RNA sequencing (RNA-seq) and clinical data were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases. Acetylation-related differentially expressed genes (DEGs) in BLCA were screened using differential expression analysis (DEA). An acetylation homeostatic model was constructed via univariate, machine learning-based least absolute shrinkage and selection operator (LASSO) and multivariate Cox regression analyses, followed by validation in multiple cohorts. Single-cell RNA-seq analysis was used to explore gene expression patterns in diverse cell types. Enrichment analysis (EA), immune infiltration, and drug sensitivity analysis (DSA) were performed to characterize molecular features of different risk groups. Finally, the biological function of CES1 as the key gene was verified by in vitro knockdown experiments. RESULTS: We established a robust acetylation homeostatic model consisting of five genes, which effectively predicted overall survival (OS) and served as an independent prognostic factor in BLCA. High-risk patients showed significantly poorer prognosis, distinct immune infiltration profiles, and differential drug sensitivity. CES1 was identified and validated as the key gene in this model, which was highly expressed in BLCA and associated with poor prognosis. Knockdown of CES1 markedly suppressed cell proliferation, invasion, and migration, and reduced intracellular coenzyme A (CoA) levels, thereby regulating PA homeostasis. CONCLUSIONS: We developed and validated a novel acetylation homeostatic model for survival stratification and personalized treatment guidance in BLCA, based on integrated transcriptome analysis and machine learning. CES1 is closely associated with intracellular CoA levels and the malignant progression of BLCA. Its potential association with PA homeostasis requires further mechanistic validation, and it may act as a candidate therapeutic biomarker for BLCA.

Bladder cancer (BLCA)

Predicting natural variation in the yeast phenotypic landscape with machine learning.

Most organismal traits result from the complex interplay of many genetic and environmental factors, making their prediction difficult. Here, we used machine learning (ML) models to explore phenotype predictions for 223 traits measured across 1011 genome-sequenced Saccharomyces cerevisiae strains isolated worldwide. We benchmarked a ML pipeline with multiple linear and non-linear models to predict phenotypes from genotypes and gene expression, and determined gradient boosting machines as the best-performing model. Gene function disruption scores and gene presence/absence emerged as best predictors, suggesting a considerable contribution of the accessory genome in controlling phenotypes. The prediction accuracy broadly varied among phenotypes, with stress resistance being easier to predict compared to growth across nutrients. ML identified relevant genomic features linked to phenotypes, including high-impact variants with established relationships to phenotypes, despite these being rare in the population. Near-perfect accuracies were achieved when other phenomics data mostly in similar conditions were used, suggesting that useful information can be conveyed across phenotypes. Overall, our study underscores the power of ML to interpret the functional outcome of genetic variants.

Genetic Variation

Dynamic evolution of chaperone-mediated autophagy is associated with tumor microenvironment remodeling and prognostic stratification in lung adenocarcinoma: insights from single-cell transcriptomics, ensemble machine learning, and experimental validation.

BACKGROUND: Lung adenocarcinoma (LUAD) shows prognostic heterogeneity, and tumor-node-metastasis (TNM) staging is limited for individualized management. Chaperone-mediated autophagy (CMA) maintains proteostasis, but its role during adenocarcinoma in situ (AIS)-minimally invasive adenocarcinoma (MIA)-invasive adenocarcinoma (IAC) progression remains unclear. METHODS: Single-cell RNA sequencing (scRNA-seq) data from GSE189357 and bulk transcriptomes from The Cancer Genome Atlas (TCGA)-LUAD and Gene Expression Omnibus (GEO) cohorts were integrated. CMA activity, cell-cell communication, weighted gene co-expression network analysis (WGCNA), tumor-normal differential expression, machine-learning survival modeling, tumor microenvironment (TME) features, drug sensitivity, and EPC1 function were analyzed. RESULTS: CMA-high tumor epithelial cells increased from AIS (58.1%) to MIA (65.7%) but declined in IAC (44.4%; p < 0.001). CMA-low cells preferentially received fibroblast-derived extracellular matrix cues. A CMA-negatively correlated module identified 69 core genes. Random survival forest (RSF) performed best among 117 machine-learning combinations (mean concordance index > 0.873). High-risk patients had worse survival across cohorts, and the risk score was independently associated with overall survival (hazard ratio = 16.013, 95% confidence interval: 9.579-26.768, p < 0.001). High-risk tumors showed proliferative activation and M0 macrophage enrichment, whereas low-risk tumors showed stronger immune-related signaling. EPC1 overexpression suppressed malignant phenotypes in A549 cells. CONCLUSION: CMA dynamics are associated with stromal and immune remodeling during LUAD progression. A CMA-based model provides robust prognostic stratification and may offer a basis for future TME-guided studies.

Chaperone-mediated autophagy