PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Drug design by machine learning: the use of inductive logic programming to model the structure-activity relationships of trimethoprim analogues binding to dihydrofolate reductase.

The machine learning program GOLEM from the field of inductive logic programming was applied to the drug design problem of modeling structure-activity relationships. The training data for the program were 44 trimethoprim analogues and their observed inhibition of Escherichia coli dihydrofolate reductase. A further 11 compounds were used as unseen test data. GOLEM obtained rules that were statistically more accurate on the training data and also better on the test data than a Hansch linear regression model. Importantly machine learning yields understandable rules that characterized the chemistry of favored inhibitors in terms of polarity, flexibility, and hydrogen-bonding character. These rules agree with the stereochemistry of the interaction observed crystallographically.

Artificial Intelligence

Leveraging structure-informed machine learning for fast steric zipper propensity prediction across whole proteomes.

Predicting the amyloid fold and the propensity of peptide segments to adopt amyloid-like structures remain a challenge. However, recent progress has facilitated structure-based prediction of steric zipper propensity and the use of machine learning to accelerate the calculation of predictive models across many scientific areas. Leveraging these advances, we have developed a new approach for rapid proteome-wide assessment of zipper profiles that is informed by four million steric zipper predictions collected over ten years. This collection is used to build a machine learning model capable of rapidly predicting steric zipper propensity, and allowing for the assessment of zippers at both the protein and proteome level. Our predictions show enrichment for zipper forming segments in proteins involved in cell wall reorganization in yeast, highlighting a potential category of interest for experimental characterization. Overall, our predictive model allows for the exploration of amyloid formation across the tree of life and provides a tool for assessment of both novel and designed sequences for zipper density.

Machine Learning

Genomic signatures associated with epidemiologically defined high-risk pathogenic Escherichia coli isolates identified by interpretable machine learning.

Pathogenic Escherichia coli is a major cause of foodborne illness worldwide and includes strains capable of causing severe disease. To establish a genome-informed framework for foodborne outbreak surveillance, we analyzed 1,029 E. coli isolates from clinical, food, livestock, and environmental sources using whole-genome sequencing. Pathogenic isolates obtained from human clinical cases or linked to documented outbreaks were classified as epidemiologically defined high-risk (EpiHR), whereas the remaining pathogenic isolates were classified as non-EpiHR. Virulence-associated genomic features were extracted using a bioinformatics pipeline, and four machine learning (ML) algorithms, including gradient boosting machine, random forest (RF), and support vector machines with linear and radial basis function kernels, were evaluated. Among them, the RF model showed the best performance, achieving an area under the curve (AUC) of 0.98 and accuracy of 0.93 in 10-fold cross-validation. Additional leave-one-group-out validation showed retained discrimination across held-out sequence types and serotypes, although performance was reduced when isolates were grouped by isolation source. Evaluation using an independent test dataset of 1,908 publicly available pathogenic E. coli genomes showed an AUC of 0.97 and a sensitivity of 0.98. Feature importance analysis using Shapley additive explanations identified influential predictive features, including traT, etpB, and enterotoxin-associated genes. A reduced 10-feature model achieved an AUC of 0.79 in the independent test dataset, supporting its exploratory use for future simplified screening approaches. These results indicate that genome-based ML provides a sensitive framework for surveillance-oriented prioritization of EpiHR pathogenic E. coli isolates, with model predictions interpreted together with epidemiological information.

Escherichia coli

Enhancing detection of polygenic adaptation: a comparative study of machine learning and statistical approaches using simulated evolve-and-resequence data.

BACKGROUND: Detecting signals of polygenic adaptation remains a significant challenge in population genomics, as traditional methods often struggle to identify the associated subtle, multi-locus allele-frequency shifts. Here, we introduced and tested several novel approaches combining machine learning techniques with traditional statistical tests to detect polygenic adaptation patterns in time-series of allele frequency changes from whole genome data. We implemented a Naive Bayesian Classifier (NBC) and One-Class Support Vector Machines (OCSVM), and compared their performance against the classical Fisher's Exact Test (FET). Furthermore, we combined machine learning and statistical models (OCSVM-FET and NBC-FET), resulting in 5 competing approaches. The framework is mainly designed and validated for evolve-and-resequence (EaR) experimental designs, where defined selection pressures and temporal sampling are feasible, but might be applicable for certain natural experiments as well. RESULTS: Using a simulated dataset based on empirical C. riparius Pool-Seq data, we evaluated methods across evolutionary scenarios varying in generation, selection strength, and number of loci under selection. Our results demonstrate that the combined OCSVM-FET approach consistently outperformed competing methods, achieving the lowest false positive rate, highest area under the curve, and high accuracy. The performance peak aligned with what we term the 'late dynamic phase' of adaptation - the period after initial selection has occurred but before fixation - highlighting the method's sensitivity to ongoing selective processes. CONCLUSIONS: Furthermore, we emphasize the critical role of parameter tuning, balancing biological assumptions with methodological rigor. While broader applicability remains an important direction for future work, the present benchmarking is intentionally scoped to EaR experimental contexts.

Machine Learning

A machine learning model and identification of immune infiltration for chronic obstructive pulmonary disease based on disulfidptosis-related genes.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a chronic and progressive lung disease. Disulfidptosis-related genes (DRGs) may be involved in the pathogenesis of COPD. From the perspective of predictive, preventive, and personalized medicine (PPPM), clarifying the role of disulfidptosis in the development of COPD could provide a opportunity for primary prediction, targeted prevention, and personalized treatment of the disease. METHODS: We analyzed the expression profiles of DRGs and immune cell infiltration in COPD patients by using the GSE38974 dataset. According to the DRGs, molecular clusters and related immune cell infiltration levels were explored in individuals with COPD. Next, co-expression modules and cluster-specific differentially expressed genes were identified by the Weighted Gene Co-expression Network Analysis (WGCNA). Comparing the performance of the random forest (RF), support vector machine (SVM), generalized linear model (GLM), and eXtreme Gradient Boosting (XGB), we constructed the ptimal machine learning model. RESULTS: DE-DRGs, differential immune cells and two clusters were identified. Notable difference in DRGs, immune cell populations, biological processes, and pathway behaviors were noted among the two clusters. Besides, significant differences in DRGs, immune cells, biological functions, and pathway activities were observed between the two clusters.A nomogram was created to aid in the practical application of clinical procedures. The SVM model achieved the best results in differentiating COPD patients across various clusters. Following that, we identified the top five genes as predictor genes via SVM model. These five genes related to the model were strongly linked to traits of the individuals with COPD. CONCLUSION: Our study demonstrated the relationship between disulfidptosis and COPD and established an optimal machine-learning model to evaluate the subtypes and traits of COPD. DRGs serve as a target for future predictive diagnostics, targeted prevention, and individualized therapy in COPD, facilitating the transition from reactive medical services to PPPM in the management of the disease.

Pulmonary Disease, Chronic Obstructive

How advances in machine learning drive early detection and risk prediction of early-onset colorectal cancer.

Early-onset colorectal cancer (EOCRC), defined as colorectal cancer diagnosed before age 50, is rising across high- and middle-income settings whilst organised screening stays anchored to older age thresholds. Blood-based liquid biopsy, combined with machine learning, is the most plausible route to early detection in this group because it does not depend on bowel preparation, endoscopy capacity, or adherence to stool-based testing. The gap is structural: incidence climbs fastest in the population below the age at which any guideline-endorsed modality is offered. The analytical challenge is that early-stage tumour-derived signals in plasma are low in abundance and distributed across heterogeneous molecular layers: circulating tumour DNA mutations, aberrant methylation, cfDNA fragmentomics, and small non-coding RNA. Machine learning converts these into a single calibrated probability. This review examines where artificial intelligence (AI)-driven liquid biopsy genuinely adds diagnostic value in EOCRC, distinguishes components in which learned models are decorative from those in which they are mechanistically necessary, and identifies the validation deficit separating research cohorts from deployable clinical tools. It summarises the first-generation tools used clinically for early detection and post-treatment monitoring, then considers analytes from exosome-bound microRNAs to long-read whole-genome sequencing of circulating plasma DNA, which reads cytosine modification natively, resolves methylation and fragmentation on single molecules, and characterises structural events short reads cannot anchor. Any analyte can feed a learned model, but more diverse input yields better discrimination. The central argument is that approved, guideline-included blood tests were validated in populations aged 45 and above, and their performance in younger patients cannot be assumed.

cfDNA fragmentomics

Causal associations between hormone replacement therapy and brain structure: Evidence from large-scale Mendelian randomization and double machine learning.

BACKGROUND: Hormone replacement therapy (HRT) is widely prescribed for the management of hormone deficiency, particularly during menopause, yet its causal effects on human brain structure remain incompletely understood. Observational studies have reported heterogeneous associations, underscoring the need for robust causal inference. METHODS: We applied an integrated causal framework combining two-sample Mendelian Randomization (MR) and Double Machine Learning (DML) to evaluate the effects of four HRT-related exposures-age at initiation, age at cessation, ever-use of HRT, and a composite medication-based phenotype-on 1366 brain imaging-derived phenotypes from the UK Biobank. Genetic instruments were derived from large-scale GWAS summary statistics, and causal estimates were validated using non-parametric DML models with cross-fitting and performance evaluation. RESULTS: Genetic instruments for age at HRT initiation, age at cessation, and ever-use of HRT were strong (median F-statistics 16.29-36.66). MR analyses identified a causal association between later initiation of HRT and lower orientation dispersion in the right inferior cerebellar peduncle (ubm-a-542; primary finding, no pleiotropy detected). An additional association with the left tapetum FA (ubm-a-243) was identified but exhibited significant directional horizontal pleiotropy (MR-Egger intercept P = 0.001) and is excluded from primary conclusions (Supplementary Note S2). Later cessation of HRT was associated with increased cortical thickness in the left middle occipital gyrus, reduced surface area in the left frontopolar cortex, and increased orientation dispersion in the splenium of the corpus callosum. Ever-use of HRT was causally linked to larger volumes of the right inferior frontal gyrus and right nucleus accumbens. These associations were corroborated by independent DML validation, which provided causally debiased estimates robust to high-dimensional confounding. Results for ukb-b-8080 (median F = 1.45) are provided in Supplementary Note S1 only; weak-instrument bias precludes causal inference. CONCLUSIONS: This study provides genetic-instrument-based and machine-learning-validated evidence for causal associations between HRT exposure-particularly its timing and lifetime use-and specific features of human brain structure, including white-matter microarchitecture, cortical thickness, and regional brain volume. These findings are FDR-controlled within exposures and independently replicated by DML, but require replication in external neuroimaging GWAS cohorts to establish definitive causal conclusions. They highlight the neurobiological relevance of sex steroid exposure and inform future research on brain aging and personalized hormone-based interventions.

Humans

Development and evaluation of a machine learning model for osteoporosis risk prediction in Korean women.

BACKGROUND: The aim of this study was to develop a machine learning (ML) model for classifying osteoporosis in Korean women based on a large-scale population cohort study. This study also aimed to assess ML model performance compared with traditional osteoporosis screening tools. Furthermore, this study aimed to examine the factors influencing the risk of osteoporosis through variable importance. METHODS: Data was collected from 4199 women aged 40-69 years in the baseline survey of the Ansan and Ansung cohort of the Korean Genome and Epidemiology Study. Osteoporosis was set as the dependent variable to develop ML classification models. Independent variables included 122 factors related to osteoporosis risk, such as socio-demographic characteristics, anthropometric parameters, lifestyle factors, reproductive factors, nutrient intakes, diet quality indices, medical history, medication history, family history, biochemical parameters, and genetic factors. The six classification models were developed using ML techniques, including decision tree, random forest, multilayer perceptron, support vector machine, light gradient boosting machine, and extreme gradient boosting (XGBoost). The six ML classification models were compared with two traditional osteoporosis screening tools, including the osteoporosis risk assessment instrument (ORAI) and the osteoporosis self-assessment tool (OST). The ML model performances were evaluated and compared using the confusion matrix and area under the curve (AUC) metrics. Variable importance was assessed using the XGBoost technique to investigate osteoporosis risk factors. RESULTS: The XGBoost model showed the highest performance out of the six ML classification models, with an accuracy of 0.705, precision of 0.664, recall of 0.830, and F1 score of 0.738. Moreover, the XGBoost model showed a higher performance on AUC than ORAI and OST. Variable importance scores were identified for 69 out of the 122 variables associated with osteoporosis risk factors. Age at menopause ranked first in variable importance. Variables of arthritis, physical activities, hypertension, education level, income level; alcohol intake, potassium intake, homeostatic model assessment for insulin resistance; energy intake, vitamin C intake, gout; and dietary inflammatory index ranked in the top 20 out of the 69 variables, using the XGBoost technique. CONCLUSIONS: This study found that an XGBoost model can be utilized to classify osteoporosis in Korean women. Age at menopause is a significant factor in osteoporosis risk, followed by arthritis, physical activities, hypertension, and education level.

Humans

Research on identification of key genes and immune-metabolic mechanisms in atrial fibrillation through integrated multi-cohort transcriptomic analysis and machine learning.

This study aimed to integrate multiple datasets for the identification of atrial fibrillation (AF)-related differentially expressed genes (DEGs), analyze their underlying mechanisms through functional enrichment and machine learning, construct diagnostic models, and explore immune-metabolic interactions to provide novel biomarkers and theoretical foundations. Gene expression datasets were integrated and normalized, with batch effects removed using principal component analysis. Differential expression analysis, functional enrichment analysis (Gene Ontology and Kyoto Encyclopedia of Genes and Genomes pathways), and machine learning-based feature gene selection and model construction were performed. Shapley additive explanations analysis was utilized to interpret the constructed models, while gene set enrichment analysis, gene set variation analysis, and immune cell infiltration analysis were conducted to investigate the associations between feature genes and immune infiltration. After integrating and normalizing gene expression data and eliminating batch effects via principal component analysis, 6 DEGs were identified, including 4 upregulated and 2 down-regulated ones. Functional enrichment analysis showed these DEGs were significantly enriched in neuro-related biological processes and pathways, indicating their key roles in AF pathogenesis. Five key feature genes were selected using LASSO, random forest, and support vector machine-recursive feature elimination algorithms. They had significant expression differences between the AF and control groups (P&#x2005;<&#x2005;.001) and were located on distinct chromosomes. The constructed random forest and support vector machine models performed excellently (area under the curve&#x2005;&#x2265;&#x2005;0.85). Shapley additive explanations analysis revealed TNNI1 contributed most to model prediction, with its expression significantly positively correlated with immune cell infiltration. Gene set enrichment analysis and gene set variation analysis analyses further showed feature genes participated in AF pathogenesis by regulating immune modulation, metabolic pathways, and autophagy. Immune cell infiltration analysis found altered proportions of T-cell subsets and M0 macrophages in the AF group, along with complex links between feature gene expression and immune cell function. This study systematically elucidated the unique gene expression patterns and key regulatory pathways associated with AF, clarifying the crucial roles of feature genes in immune regulation, metabolic imbalance, and cellular dysfunction. These findings provide a theoretical basis and potential therapeutic targets for understanding AF pathogenesis and developing targeted treatment strategies.

Atrial Fibrillation

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans

Integrated transcriptome analysis and machine learning to construct a homeostatic model of acetylation for bladder cancer and validate the key gene CES1.

BACKGROUND: Bladder cancer (BLCA) is one of the most common malignant tumors of the urinary system. Protein acetylation (PA) plays a critical role in regulating multiple biological processes (BPs), cellular homeostasis, and cancer-related signaling pathways. This study aimed to construct a homeostatic model of acetylation for BLCA using integrated transcriptome analysis and machine learning and to validate the key gene CES1. METHODS: RNA sequencing (RNA-seq) and clinical data were obtained from The Cancer Genome Atlas (TCGA) and Gene Expression Omnibus (GEO) databases. Acetylation-related differentially expressed genes (DEGs) in BLCA were screened using differential expression analysis (DEA). An acetylation homeostatic model was constructed via univariate, machine learning-based least absolute shrinkage and selection operator (LASSO) and multivariate Cox regression analyses, followed by validation in multiple cohorts. Single-cell RNA-seq analysis was used to explore gene expression patterns in diverse cell types. Enrichment analysis (EA), immune infiltration, and drug sensitivity analysis (DSA) were performed to characterize molecular features of different risk groups. Finally, the biological function of CES1 as the key gene was verified by in vitro knockdown experiments. RESULTS: We established a robust acetylation homeostatic model consisting of five genes, which effectively predicted overall survival (OS) and served as an independent prognostic factor in BLCA. High-risk patients showed significantly poorer prognosis, distinct immune infiltration profiles, and differential drug sensitivity. CES1 was identified and validated as the key gene in this model, which was highly expressed in BLCA and associated with poor prognosis. Knockdown of CES1 markedly suppressed cell proliferation, invasion, and migration, and reduced intracellular coenzyme A (CoA) levels, thereby regulating PA homeostasis. CONCLUSIONS: We developed and validated a novel acetylation homeostatic model for survival stratification and personalized treatment guidance in BLCA, based on integrated transcriptome analysis and machine learning. CES1 is closely associated with intracellular CoA levels and the malignant progression of BLCA. Its potential association with PA homeostasis requires further mechanistic validation, and it may act as a candidate therapeutic biomarker for BLCA.

Bladder cancer (BLCA)

Predicting natural variation in the yeast phenotypic landscape with machine learning.

Most organismal traits result from the complex interplay of many genetic and environmental factors, making their prediction difficult. Here, we used machine learning (ML) models to explore phenotype predictions for 223 traits measured across 1011 genome-sequenced Saccharomyces cerevisiae strains isolated worldwide. We benchmarked a ML pipeline with multiple linear and non-linear models to predict phenotypes from genotypes and gene expression, and determined gradient boosting machines as the best-performing model. Gene function disruption scores and gene presence/absence emerged as best predictors, suggesting a considerable contribution of the accessory genome in controlling phenotypes. The prediction accuracy broadly varied among phenotypes, with stress resistance being easier to predict compared to growth across nutrients. ML identified relevant genomic features linked to phenotypes, including high-impact variants with established relationships to phenotypes, despite these being rare in the population. Near-perfect accuracies were achieved when other phenomics data mostly in similar conditions were used, suggesting that useful information can be conveyed across phenotypes. Overall, our study underscores the power of ML to interpret the functional outcome of genetic variants.

Genetic Variation

Dynamic evolution of chaperone-mediated autophagy is associated with tumor microenvironment remodeling and prognostic stratification in lung adenocarcinoma: insights from single-cell transcriptomics, ensemble machine learning, and experimental validation.

BACKGROUND: Lung adenocarcinoma (LUAD) shows prognostic heterogeneity, and tumor-node-metastasis (TNM) staging is limited for individualized management. Chaperone-mediated autophagy (CMA) maintains proteostasis, but its role during adenocarcinoma in situ (AIS)-minimally invasive adenocarcinoma (MIA)-invasive adenocarcinoma (IAC) progression remains unclear. METHODS: Single-cell RNA sequencing (scRNA-seq) data from GSE189357 and bulk transcriptomes from The Cancer Genome Atlas (TCGA)-LUAD and Gene Expression Omnibus (GEO) cohorts were integrated. CMA activity, cell-cell communication, weighted gene co-expression network analysis (WGCNA), tumor-normal differential expression, machine-learning survival modeling, tumor microenvironment (TME) features, drug sensitivity, and EPC1 function were analyzed. RESULTS: CMA-high tumor epithelial cells increased from AIS (58.1%) to MIA (65.7%) but declined in IAC (44.4%; p < 0.001). CMA-low cells preferentially received fibroblast-derived extracellular matrix cues. A CMA-negatively correlated module identified 69 core genes. Random survival forest (RSF) performed best among 117 machine-learning combinations (mean concordance index > 0.873). High-risk patients had worse survival across cohorts, and the risk score was independently associated with overall survival (hazard ratio = 16.013, 95% confidence interval: 9.579-26.768, p < 0.001). High-risk tumors showed proliferative activation and M0 macrophage enrichment, whereas low-risk tumors showed stronger immune-related signaling. EPC1 overexpression suppressed malignant phenotypes in A549 cells. CONCLUSION: CMA dynamics are associated with stromal and immune remodeling during LUAD progression. A CMA-based model provides robust prognostic stratification and may offer a basis for future TME-guided studies.

Chaperone-mediated autophagy

Investigating the mechanisms of PhIP-induced colorectal cancer through network toxicology, machine learning, and molecular dynamics simulation.

BACKGROUND: Over the past few years, 2-amino-1-methyl-6-phenylimidazo[4,5-b]pyridine (PhIP)- a compound from grilled or processed meats-has emerged as a major player in cancer development, especially colorectal cancer (CRC). This work dives into its potential links to CRC and uncovers the key genes that bridge this connection. METHODS: We tapped into various databases to pinpoint target genes tied to PhIP and CRC, then ran protein-protein interaction (PPI) analyses for visualization. Next, we explored underlying mechanisms through Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) enrichment. To nail down predictions, we tested 107 machine learning pipelines and picked the best one, validating its accuracy and the core genes' prognostic value across datasets. Next, molecular docking and dynamics simulations probed the interactions between these genes and PhIP. Finally, cell proliferation was assessed using Cell Counting Kit-8 (CCK-8) and 5-ethynyl-2'-deoxyuridine (EdU) assays, and polymerase chain reaction (PCR) was performed to validate the expression levels of the hub genes. RESULTS: Our analysis identified 39 overlapping genes, from which a machine learning model (glmBoost + Enet) identified six candidate targets: CDK4, CEBPB, COMT, SOX9, TIMP1, and TOP2A. To prioritize these, a hierarchical screening framework was applied. Molecular docking and dynamics simulations identified CDK4, COMT, and TIMP1 as the most stable interactors with PhIP. Functional assays confirmed that PhIP treatment significantly enhanced the proliferation of CRC cells. Crucially, quantitative PCR (qPCR) validation in multiple CRC cell lines identified TIMP1 as the primary target, showing the most consistent and significant upregulation upon PhIP exposure. CONCLUSIONS: In essence, these genes drive PhIP is role in CRC, offering novel insights into its molecular pathways. This could reshape how we tackle food-related pollutants, paving the way for better prevention and targeted therapies.

Colorectal cancer (CRC)

Integrated single-cell transcriptomics, Mendelian randomization, and machine learning identify CEBPZ as an immune-related biomarker in oral lichen planus.

BACKGROUND: Oral lichen planus (OLP) is a chronic, immune-mediated oral mucosal disease with complex pathophysiology and potential for malignant transformation. Understanding its molecular basis is critical for the development of precise diagnostic and therapeutic strategies. OBJECTIVES: We aimed to identify key immune-related biomarkers and characterize cellular dynamics in OLP, with a particular focus on the role of CEBPZ in disease pathogenesis. MATERIAL AND METHODS: We analyzed single-cell RNA sequencing (scRNA-seq) data from OLP lamina propria samples (GSE211630) to identify disease-specific T-cell subpopulations using high-dimensional weighted gene co-expression network analysis (hdWGCNA) for oxidative stress-related gene modules.-data-based Mendelian randomization (SMR) integrated FinnGen genome-wide association study (GWAS; 342,499 Europeans) data with Genotype-Tissue Expression (GTEx) expression quantitative trait loci (eQTL) data to identify causal genes. Machine learning (ML) models (least absolute shrinkage and selection operator (LASSO) and convolutional neural network (CNN)) were developed using bulk RNA-seq datasets (GSE52130 and GSE38616) for diagnostic purposes. RESULTS: We identified OLP-specific T-cell populations (clusters 0, 3, 5, 7, 13, and 15) with enhanced migration inhibition factor (MIF) pathway signaling toward B cells and monocytes. Two oxidative stress-associated modules contained hub genes, including CEBPZ. Summary-data-based Mendelian randomization analysis identified 231 OLP-associated genes, with CEBPZ uniquely intersecting LASSO-selected markers (odds ratio (OR) = 1.057, 95% confidence interval (95% CI) = 1.013-1.102, p = 0.010). Machine learning models achieved area under the curve (AUC) values ranging from 0.653 to 0.745, with the CNN model reaching a validation accuracy of 0.735. CEBPZ showed elevated expression in OLP T cells and correlated with enhanced MIF-(CD74+CXCR4) signaling. CONCLUSIONS: This integrative approach identifies CEBPZ as a pivotal biomarker linking genetic susceptibility, oxidative stress, and immune dysregulation in OLP. Our diagnostic models offer promising tools for OLP management.

CEBPZ

Machine learning and multi-omics clustering to map cellular rewiring and immune evasion in ccRCC.

Immune checkpoint blockade (ICB) efficacy in clear cell renal cell carcinoma (ccRCC) is limited by tumor microenvironment (TME) heterogeneity. Because traditional bulk-derived models lack spatial resolution, we developed an integrated framework connecting macroscopic survival risks to microscopic TME structures. We applied ten algorithms to establish multi-omics subtypes and evaluated 101 machine-learning combinations across three independent cohorts to generate a Consensus Machine Learning-driven Signature (CMLS). The signature's spatial and cellular origins were decoded using spatial transcriptomics (ST) and a 140,000-cell scRNA-seq atlas. Expression of key genes was experimentally validated via RT-qPCR in 17 paired ccRCC clinical tissues. We identified two molecular subtypes with distinct clinical and epigenetic profiles. SuperPC optimization yielded a 24-gene CMLS serving as an independent prognostic factor. scRNA-seq and ST deconvolution revealed these signals predominantly originate from cancer-associated fibroblasts (CAFs) and malignant epithelial cells, which collaborate to drive spatial immune exclusion. RT-qPCR confirmed significant overexpression of five core CMLS genes in ccRCC versus adjacent normal tissues. Low CMLS scores correlated with enhanced ICB responsiveness, whereas high-CMLS tumors demonstrated specific vulnerability to dasatinib and dabrafenib. The CMLS translates spatial immune-exclusion dynamics into a quantifiable metric, outperforming tumor mutational burden in predicting ICB benefits, providing a robust tool for patient stratification in ccRCC.

Humans

Revealing potential biomarkers and metabolic mechanisms of ovarian aging in hens during late laying period based on machine learning and metabolomics.

Ovarian function decline during the late laying period represents a major bottleneck for the economic efficiency of the global poultry industry. However, the underlying metabolic mechanisms and reliable early-warning biomarkers for ovarian aging remain poorly understood. In this study, we performed the first untargeted LC-MS/MS metabolomics analysis of ovarian tissues from Taihe silky fowls at peak laying (30&#xa0;weeks) and late laying (50&#xa0;weeks) stages, and employed an ensemble machine learning strategy integrating LASSO, random forest, and support vector machine (SVM) algorithms to identify high-confidence core biomarkers of ovarian aging. Gene expression analysis was further conducted to validate the potential molecular mechanisms. Our results showed that the metabolic profiles of ovarian tissues differed significantly between the two groups. A total of 6 core biomarkers were identified, 4 of which were long-chain acylcarnitines. Mechanistic analysis revealed that downregulation of key genes in the carnitine shuttle system led to impaired mitochondrial fatty acid &#x3b2;-oxidation, which in turn triggered excessive oxidative stress and compromised ovarian endocrine function. In conclusion, this study identifies long-chain acylcarnitines as potential metabolic biomarkers for ovarian aging in Taihe silky fowls. These findings provide novel insights into the metabolic basis of poultry ovarian aging and lay a theoretical foundation for the precise regulation of reproductive performance in indigenous poultry breeds.

Animals

MicroRNAs signatures in small extracellular vesicles for psychological resilience in young adults using machine learning.

AIMS: Psychological resilience refers to an individual's capacity to adapt to adverse events. MicroRNAs (miRNAs) play a crucial role in regulating post-transcriptional processes, while small extracellular vesicles (sEVs) act as transport vehicles. This study aimed to employ genome-wide profiling to identify and validate differences in the expression of resilience-associated sEV-miRNAs between low resilience (LR) and high resilience (HR) in young adults. METHODS: Eighty participants were divided into LR or HR based on the Connor - Davidson Resilience Scale (CD-RISC). The expression levels of the target sEV-miRNAs in LR and HR were compared and analyzed. RESULTS: Expression analyses demonstrated significant differences in let-7b, miR-151b, miR-335, and miR-193a between LR and HR (p&#x2009;<&#x2009;0.01), with let-7b showing the highest discriminative ability. The AUC values for each sEV-miRNA ranged from 0.74 to 0.94, based on logistic regression and three machine learning models: random forest, support vector machine, and eXtreme gradient boosting. Based on leave-one-out cross-validation in different models, the combined four sEV-miRNAs demonstrated strong performance for detecting LR (AUC&#x2009;=&#x2009;0.87-0.90). Sex-specific differences were also observed, with female participants showing more pronounced resilience signatures in targeted sEV-miRNAs. CONCLUSIONS: These findings suggest that sEV-miRNAs hold potential as biomarkers for psychological resilience in young adults.

Humans