PubMed HealthSearch

SEARCH · PubMed Health

Results for “Random Forest”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

The mouse gut microbiota responds to predator odor and predicts host behavior.

Chronic stressors can alter the mammalian gut microbiota in ways that mediate host stress responses, but the impacts of acute stressors on these interactions are less well understood. Here, we show that brief exposure of wild-derived mice to predator odor altered gut-microbiota composition, which in turn predicted host behavior. We investigated the individual and combined effects of 15-minute exposures to synthetic fox fecal odor and 30 days of chronic social isolation, an established chronic stressor. Using ethological assays, visceral adipose tissue transcriptomics, and genome-resolved metagenomics, we found that predator-odor exposure significantly affected mouse behavior, gene expression, and gut microbiota. Predator odor-responsive bacteria were associated with the expression of genes involved in anti-microbial defense, and host behavioral responses were predicted by random forest models trained on gut-microbiota profiles. These findings indicate interactions between the gut microbiota and wild-mouse responses to the threat of predation, an ecologically relevant acute stressor.

Journal Article

Gene Specific Pathogenicity Predictor for Chromatin-Remodeling BAF Complex-Associated Neurodevelopmental Disorders.

Advancements in whole genome sequencing have increased the number of variants of uncertain significance (VUS) identified in patient genomes. This has created a diagnostic bottleneck for genetic counselors tasked with sifting through these variants and determining those most likely to be causative for a patient's clinical presentation. Machine learning (ML) tools can aid in identifying pathogenic variants from VUS, but there is a need for gene-specific algorithms that predict pathogenic variants with high accuracy. To address this need, we present a workflow for developing gene-specific, ensemble-learning ML tools, that leverage outputs from other algorithms, locations of variants within the gene, and evolutionary conservation data to make a prediction of pathogenicity. Variants in SMARCA2 and SMARCA4 that are associated with rare neurodevelopmental diseases were used to screen 15 ML algorithms. A random forest learner was tuned to yield a final accuracy of 0.93 on holdout data. Generalizing this predictor to other BAF complex proteins resulted in a sharp decline in performance. We trained a final predictor for all genes in the study to create a predictor that identifies pathogenic variants in these BAF subunits with an accuracy of 0.91 on holdout data. This predictor specific to BAF complex proteins performs with higher accuracy and AUROC than any other predictor. The decline in performance when generalized to other proteins emphasizes the need for the gene-specific calibration of predictors. Our workflow for the development of such models provides a quick, computationally inexpensive route for improving the ML tools available to genetic counselors.

Journal Article

A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.

Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

cross‐taxon transferability

NR3C1 Modulates Wnt Signalling to Influence the Invasiveness and Immune Features of Nonfunctioning Invasive Pituitary Adenomas.

Pituitary adenomas (PAs) are common intracranial tumours, and invasiveness in nonfunctioning invasive pituitary adenomas (NIPAs) predicts poor prognosis. The molecular mechanisms driving this phenotype remain unclear. This study explored the role of nuclear receptor subfamily 3 group C member 1 (NR3C1) in NIPA invasiveness and its regulation of Wnt signalling. mRNA expression profiles of 32 PA samples were generated by RNA-seq, and proteomic data from 19 samples were obtained by mass spectrometry. Immune-related differentially expressed genes (DEGs) were retrieved from GeneCards. Weighted gene coexpression network analysis identified modules and hub genes linked to invasiveness, while machine learning methods (support vector machine, LASSO, random forest) prioritised key genes. Gene set enrichment analysis (GSEA) assessed pathways associated with candidate gene expression. NR3C1 expression and function were validated by immunohistochemistry, Western blotting and invasion assays. Integration of transcriptomic, proteomic and immune-related datasets yielded 11 overlapping genes, with NR3C1 emerging as the top candidate. NR3C1 was significantly upregulated in NIPAs and demonstrated good discriminatory power by ROC analysis. GSEA associated high NR3C1 expression with Wnt pathway activation. Functional experiments confirmed that NR3C1 overexpression enhances the invasive capacity of PA cells. NR3C1 promotes the invasive phenotype of NIPAs by activating Wnt signalling. These findings suggest NR3C1 as a potential biomarker and therapeutic target for invasive pituitary adenomas.

Humans

Integrated Bulk and Single-Cell RNA-Seq Analysis Reveals Transcriptional Activation of PTGS2 by FOS in Progression From T2DM to T2DM-Associated NAFLD.

Type 2 diabetes mellitus (T2DM) and nonalcoholic fatty liver disease (NAFLD) frequently coexist, exacerbating disease burden. However, the molecular mechanisms underlying the progression from T2DM to T2DM-associated NAFLD remain unclear. This study investigated the regulatory function of FOS-mediated PTGS2 activation in this transition. We integrated bulk RNA-seq data from GEO, single-cell transcriptomic data and transcriptomes from patients with T2DM-associated NAFLD. Differentially expressed genes were identified using the limma package, and T2DM-related gene modules were defined by weighted gene co-expression network analysis. LASSO regression and random forest identified 14 candidate genes, with PTGS2 and FOS prioritised. Single-cell analysis showed increased FOS and PTGS2 expression in monocytes, CD8+ T cells and Kupffer cells. Transcription factor prediction and dual-luciferase assays confirmed that FOS directly binds the PTGS2 promoter and drives its transcription. In vitro, FOS silencing decreased PTGS2 expression, cytokine secretion and apoptosis under high-glucose and free fatty acid conditions, whereas PTGS2 overexpression exacerbated inflammation and apoptosis independently of FOS expression. These findings demonstrate that FOS transcriptionally activates PTGS2, contributing to hepatic inflammation and apoptosis during the progression from T2DM to NAFLD. PTGS2 may serve as a promising biomarker and therapeutic target for T2DM-associated NAFLD.

Single-Cell Gene Expression Analysis

Integrative Multi-Omics Analysis of Stem Growth Habit Divergence in Wild Soybean (Glycine soja).

Stem architecture is a major determinant of lodging resistance, biomass accumulation, and harvest efficiency in soybean. However, the molecular features associated with contrasting stem growth habits in wild soybean remain incompletely characterised. Here, we performed an integrated transcriptomic, metabolomic, and epigenomic analysis of stem growth-habit divergence in wild soybean, comparing the wild-type accession ZYD7068 with contrasting vining and erect mutant lines derived from carbon-ion beam mutagenesis. Pairwise transcriptomic comparisons identified between 20 311 and 28 705 differentially expressed genes per contrast, with a core set of 2672 genes consistently altered across the comparisons. Functional enrichment, gene set variation analysis, and gene set enrichment analysis converged on xylem and phloem pattern formation as a prominent molecular pathway associated with growth-habit divergence. Random forest analysis identified BBR-BPC and ARF transcription factor families as major molecular discriminators, while metabolomic profiling revealed distinct metabolic profiles involving amino-acid-derived and lipid-associated metabolites. Whole-genome bisulfite sequencing revealed context-specific DNA methylation differences, including substantial variation in CHG methylation among erect mutant lines. Integrated network and in silico perturbation analyses prioritised four candidate genes associated with vascular development for future functional validation. Together, these results provide a multi-layer molecular resource for investigating stem growth-habit divergence in G. soja and establish testable candidate pathways and genes for subsequent functional studies and soybean improvement.

glycine soja

Spatiotemporal mapping of tertiary lymphoid structure heterogeneity shapes immune niches and clinical outcomes in intrahepatic cholangiocarcinoma.

Intrahepatic cholangiocarcinoma (iCCA) is a highly lethal malignancy with limited therapeutic options. The spatial architecture and functional diversity of tertiary lymphoid structures (TLSs) in iCCA remain unclear. Here, we present a multimodal spatial atlas of TLSs and identified intratumoral TLSs (iTLSs) as independent prognostic markers. Bulk proteomic profiling of 214 discovery and 155 validation cases identified a four-tier TLS-based tumor microenvironment classification system and supported development of a TLS-predictive random forest classifier. Imaging mass cytometry revealed that iTLS+ tumors harbor structured immune architectures, where M1-like tissue-resident macrophages (RTMs), dendritic cells, and CXCL13+ CD4+ T cells colocalize to form antigen-presenting neighborhoods (apc-CNs) spatially coupled to TLS core regions (TLScore-CNs). Single-cell spatial transcriptomics further resolved 61 TLSs into 14 spatial niches and defined a pseudotemporal maturation continuum: aggregated, activated, and postactivated. Intraniche communication, primarily mediated by ifnCAFs, iCAFs, and CXCL12+ macrophages, evolved dynamically with maturation. Single-nucleus RNA sequencing combined with Tangram-based spatial mapping revealed CXCL12+ macrophages and iCAFs forming a peripheral band in aggregated TLSs, whereas ifnCAFs infiltrated TLS interiors during activation. These findings define TLS heterogeneity and provide insights for stroma-directed immunotherapy.

Cholangiocarcinoma

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides

Potential evaluation of SULT1A3 as an early diagnostic marker for nasopharyngeal carcinoma: a study based on serum proteomics screening and ELISA validation.

BACKGROUND: Nasopharyngeal carcinoma (NPC) represents a highly prevalent and aggressive malignancy endemic to Southeast Asia. Early and accurate diagnosis is critical to improving survival outcomes; however, the absence of robust, stage-specific biomarkers remains a key obstacle to clinical implementation of early screening strategies. METHODS: We performed untargeted serum proteomic profiling using mass spectrometry in 15 treatment-na&#xef;ve early-stage NPC patients and 15 VCA-IgA-positive healthy controls. Bioinformatics analyses were conducted to identify differentially expressed proteins (DEPs). Machine learning (random forest combined with recursive feature elimination) was employed to prioritize candidate biomarkers, which were subsequently verified using enzyme-linked immunosorbent assay (ELISA) in independent sample cohorts. RESULTS: In total, 1,428 serum proteins were identified, among which 1,410 were reliably quantified. We observed 31 upregulated and 189 downregulated proteins in NPC patients relative to controls. Spearman correlation analysis revealed significant associations: LTA4H (leukotriene A4 hydrolase) levels correlated with serum cell infiltration (r&#x2009;=&#x2009;0.383, p&#x2009;=&#x2009;0.032) and CD8&#x2009;+&#x2009;T-cell abundance (r&#x2009;=&#x2009;0.408, p&#x2009;=&#x2009;0.021); both SULT1A3 (sulfotransferase family 1&#xa0;A member 3) and FGL1 (fibrinogen-like protein 1) levels were positively associated with M1 macrophage infiltration (r&#x2009;=&#x2009;0.510, p&#x2009;=&#x2009;0.003 and r&#x2009;=&#x2009;0.430, p&#x2009;=&#x2009;0.015, respectively). In a preliminary validation cohort (n&#x2009;=&#x2009;80), ELISA yielded AUC values of 0.631 (95% CI: 0.515-0.736, p&#x2009;=&#x2009;0.04) for LTA4H, 0.787 (95% CI: 0.681-0.871, p&#x2009;<&#x2009;0.001) for SULT1A3, and 0.688 (95% CI: 0.575-0.787, p&#x2009;=&#x2009;0.002) for FGL1. In large-scale independent validation, SULT1A3 achieved an AUC of 0.826 (95% CI: 0.766-0.876; sensitivity&#x2009;=&#x2009;78.89%, specificity&#x2009;=&#x2009;75.47%) in cohort 1 (n&#x2009;=&#x2009;196) and 0.796 (95% CI: 0.723-0.857; sensitivity&#x2009;=&#x2009;76.67%, specificity&#x2009;=&#x2009;76.67%) in cohort 2 (n&#x2009;=&#x2009;150). CONCLUSIONS: Through an integrated workflow combining proteomic screening, machine learning prioritization, and multi-stage ELISA validation, we identified SULT1A3 as a candidate serum-based biomarker for early detection of NPC. Preliminary findings suggest that SULT1A3 may have potential utility in clinical screening, though further validation in independent, multi&#x2011;center cohorts is required.

Humans

Development and evaluation of a machine learning model for osteoporosis risk prediction in Korean women.

BACKGROUND: The aim of this study was to develop a machine learning (ML) model for classifying osteoporosis in Korean women based on a large-scale population cohort study. This study also aimed to assess ML model performance compared with traditional osteoporosis screening tools. Furthermore, this study aimed to examine the factors influencing the risk of osteoporosis through variable importance. METHODS: Data was collected from 4199 women aged 40-69 years in the baseline survey of the Ansan and Ansung cohort of the Korean Genome and Epidemiology Study. Osteoporosis was set as the dependent variable to develop ML classification models. Independent variables included 122 factors related to osteoporosis risk, such as socio-demographic characteristics, anthropometric parameters, lifestyle factors, reproductive factors, nutrient intakes, diet quality indices, medical history, medication history, family history, biochemical parameters, and genetic factors. The six classification models were developed using ML techniques, including decision tree, random forest, multilayer perceptron, support vector machine, light gradient boosting machine, and extreme gradient boosting (XGBoost). The six ML classification models were compared with two traditional osteoporosis screening tools, including the osteoporosis risk assessment instrument (ORAI) and the osteoporosis self-assessment tool (OST). The ML model performances were evaluated and compared using the confusion matrix and area under the curve (AUC) metrics. Variable importance was assessed using the XGBoost technique to investigate osteoporosis risk factors. RESULTS: The XGBoost model showed the highest performance out of the six ML classification models, with an accuracy of 0.705, precision of 0.664, recall of 0.830, and F1 score of 0.738. Moreover, the XGBoost model showed a higher performance on AUC than ORAI and OST. Variable importance scores were identified for 69 out of the 122 variables associated with osteoporosis risk factors. Age at menopause ranked first in variable importance. Variables of arthritis, physical activities, hypertension, education level, income level; alcohol intake, potassium intake, homeostatic model assessment for insulin resistance; energy intake, vitamin C intake, gout; and dietary inflammatory index ranked in the top 20 out of the 69 variables, using the XGBoost technique. CONCLUSIONS: This study found that an XGBoost model can be utilized to classify osteoporosis in Korean women. Age at menopause is a significant factor in osteoporosis risk, followed by arthritis, physical activities, hypertension, and education level.

Humans

Systematic review of machine learning approaches for predicting sickle cell crisis and mortality risk at the climate-health nexus.

BACKGROUND: Sickle cell anemia (SCA) is a severe genetic blood disorder characterized by recurrent vaso-occlusive crises and increased mortality, with the greatest burden occurring in low- and middle-income countries. Climatic and environmental conditions, including temperature variability, humidity, rainfall, air pollution, and seasonal changes, have been associated with disease exacerbation. However, the extent to which these factors have been incorporated into predictive models remains unclear. This study systematically reviews the application of machine learning (ML) models for predicting SCA crises and mortality in relation to climate and environmental factors. METHODOLOGY: The PRISMA guidelines were used, and 34 peer-reviewed studies published between 2005 and 2026 were analyzed to identify the climate variables, ML approaches employed, and predictive performance. The reviewed studies applied a range of ML techniques, including artificial neural networks, random forests, support vector machines, decision trees, logistic regression, and deep learning models. Temperature, humidity, rainfall, wind speed, air quality indicators, and seasonal patterns were the most frequently examined environmental variables. RESULTS: The findings indicate that most existing models rely predominantly on clinical and demographic data, with limited integration of climate information and inadequate representation of high-burden regions, especially Sub-Saharan Africa. Studies incorporating environmental variables reported improved predictive performance and highlighted the potential of climate-informed early warning systems for SCA management. CONCLUSION: The review recommends development of interdisciplinary, climate-aware ML frameworks, expansion of longitudinal environmental datasets, and increased research in underrepresented regions to support climate-resilient and patient-centered SCA care.

Humans

Major depletion of insulin sensitivity-associated taxa in the gut microbiome of persons living with HIV controlled by antiretroviral drugs.

BACKGROUND: Persons living with HIV (PWH) harbor an altered gut microbiome (higher abundance of Prevotella and lower abundance of Bacillota and Ruminococcus lineages) compared to non-infected individuals. Some of these alterations are linked to sexual preference and others to the HIV infection. The relationship between these lineages and metabolic alterations, often present in aging PWH, has been poorly investigated. METHODS: In this study, we compared fecal metagenomes of 25 antiretroviral-treatment (ART)-controlled PWH to three independent control groups of 25 non-infected matched individuals by means of univariate analyses and machine learning methods. Moreover, we used two external datasets to validate predictive models of PWH classification. Next, we searched for associations between clinical and biological metabolic parameters with taxonomic and functional microbiome profiles. Finally, we compare the gut microbiome in 7 PWH after a 17-week ART switch to raltegravir/maraviroc. RESULTS: Three major enterotypes (Prevotella, Bacteroides and Ruminococcaceae) were present in all groups. The first Prevotella enterotype was enriched in PWH, with several of characteristic lineages associated with poor metabolic profiles (low HDL and adiponectin, high insulin resistance (HOMA-IR)). Conversely butyrate-producing lineages were markedly depleted in PWH independently of sexual preference and were associated with a better metabolic profile (higher HDL and adiponectin and lower HOMA-IR). Accordingly with the worst metabolic status of PWH, butyrate production and amino-acid degradation modules were associated with high HDL and adiponectin and low HOMA-IR. Random Forest models trained to classify PWH vs. control on taxonomic abundances displayed high generalization performance on two external holdout datasets (ROC AUC of 80-82%). Finally, no significant alterations in microbiome composition were observed after switching to raltegravir/maraviroc. CONCLUSION: High resolution metagenomic analyses revealed major differences in the gut microbiome of ART-controlled PWH when compared with three independent matched cohorts of controls. The observed marked insulin resistance could result both from enrichment in Prevotella lineages, and from the depletion in species producing butyrate and involved into amino-acid degradation, which depletion is linked with the HIV infection.

Humans

Identification of key immune-related genes and potential therapeutic drugs in diabetic nephropathy based on machine learning algorithms.

BACKGROUND: Diabetic nephropathy (DN) is a major contributor to chronic kidney disease. This study aims to identify immune biomarkers and potential therapeutic drugs in DN. METHODS: We analyzed two DN microarray datasets (GSE96804 and GSE30528) for differentially expressed genes (DEGs) using the Limma package, overlapping them with immune-related genes from ImmPort and InnateDB. LASSO regression, SVM-RFE, and random forest analysis identified four hub genes (EGF, PLTP, RGS2, PTGDS) as proficient predictors of DN. The model achieved an AUC of 0.995 and was validated on GSE142025. Single-cell RNA data (GSE183276) revealed increased hub gene expression in epithelial cells. CIBERSORT analysis showed differences in immune cell proportions between DN patients and controls, with the hub genes correlating positively with neutrophil infiltration. Molecular docking identified potential drugs: cysteamine, eltrombopag, and DMSO. And qPCR and western blot assays were used to confirm the expressions of the four hub genes. RESULTS: Analysis found 95 and 88 distinctively expressed immune genes in the two DN datasets, with 14 consistently differentially expressed immune-related genes. After machine learning algorithms, EGF, PLTP, RGS2, PTGDS were identified as the immune-related hub genes associated with DN. In addition, the mRNA and protein levels of them were obviously elevated in HK-2 cells treated with glucose for 24&#xa0;h, as well as their mRNA expressions in kidney tissues of mice with DN. CONCLUSION: This study identified 4 hub immune-related genes (EGF, PLTP, RGS2, PTGDS), as well as their expression profiles and the correlation with immune cell infiltration in DN.

Diabetic Nephropathies

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans

A machine learning model and identification of immune infiltration for chronic obstructive pulmonary disease based on disulfidptosis-related genes.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a chronic and progressive lung disease. Disulfidptosis-related genes (DRGs) may be involved in the pathogenesis of COPD. From the perspective of predictive, preventive, and personalized medicine (PPPM), clarifying the role of disulfidptosis in the development of COPD could provide a opportunity for primary prediction, targeted prevention, and personalized treatment of the disease. METHODS: We analyzed the expression profiles of DRGs and immune cell infiltration in COPD patients by using the GSE38974 dataset. According to the DRGs, molecular clusters and related immune cell infiltration levels were explored in individuals with COPD. Next, co-expression modules and cluster-specific differentially expressed genes were identified by the Weighted Gene Co-expression Network Analysis (WGCNA). Comparing the performance of the random forest (RF), support vector machine (SVM), generalized linear model (GLM), and eXtreme Gradient Boosting (XGB), we constructed the ptimal machine learning model. RESULTS: DE-DRGs, differential immune cells and two clusters were identified. Notable difference in DRGs, immune cell populations, biological processes, and pathway behaviors were noted among the two clusters. Besides, significant differences in DRGs, immune cells, biological functions, and pathway activities were observed between the two clusters.A nomogram was created to aid in the practical application of clinical procedures. The SVM model achieved the best results in differentiating COPD patients across various clusters. Following that, we identified the top five genes as predictor genes via SVM model. These five genes related to the model were strongly linked to traits of the individuals with COPD. CONCLUSION: Our study demonstrated the relationship between disulfidptosis and COPD and established an optimal machine-learning model to evaluate the subtypes and traits of COPD. DRGs serve as a target for future predictive diagnostics, targeted prevention, and individualized therapy in COPD, facilitating the transition from reactive medical services to PPPM in the management of the disease.

Pulmonary Disease, Chronic Obstructive

Exploring diagnostic m6A regulators in primary open-angle glaucoma: insight from gene signature and possible mechanisms by which key genes function.

PURPOSE: The purpose of this study was to interrogate the potential role of N6-methyladenosine (m6A) regulators in the process of trabecular meshwork (TM) tissue damage in patients with primary open-angle glaucoma (POAG). METHODS: Firstly, the expression profile of m6A regulators in TM tissues of POAG patients was comprehensively analyzed by bioinformatics analysis; Plasmid transfection and siRNA gene interference were used to enhance or weaken the expression levels of YTHDC2 in human trabecular meshwork cells (HTMCs); Cell migration ability was detected by transwell chamber assay; Immunofluorescence staining assay was used to evaluate the expression of extracellular matrix (ECM) related proteins. RESULTS: Through the analysis of GSE27276 database, 5 m6A regulators with different expression in POAG were screened out. The results of random forest model showed that these 5 m6A regulators exhibited diagnostic potential and were characteristic genes of POAG. All POAG samples could be effectively divided into two groups based on the expression levels of these 5 hub m6A regulators. Immune cell infiltration analysis indicated that the levels of activated CD8+ T cells and regulatory T cells were different in the two subtypes. HTMC oxidative stress cell model and TGF-&#x3b2;2 stimulation cell model were further constructed to verify the expression of the aforementioned hub m6A regulators, and it was found that YTHDC2 mRNA showed the same expression trend in both models. The silencing of YTHDC2 enhanced the migration ability of HTMCs and increased the synthesis ability of ECM. However, when YTHDC2&#x394;YTH, which lacks the YTH domain, is overexpressed in HTMCs, there is no significant change in the ECM synthesis ability. CONCLUSIONS: The differentially expressed m6A regulators in TM tissues may serve as potential diagnostic biomarkers for POAG. And, in HTMCs, the expression level of YTHDC2 mRNA was changed under oxidative stress or TGF-&#x3b2;2 intervention, and then exerted its regulation on cell migration and ECM synthesis capability through m6A modification, which may be an important part of the disease process of POAG.

Humans

FOSB is a key factor in the genetic link between inflammatory bowel disease and acute myocardial infarction: multiple bioinformatics analyses and validation.

BACKGROUND: Inflammatory Bowel Disease (IBD), which includes Crohn's disease and ulcerative colitis, is associated with an increased risk of Acute Myocardial Infarction (AMI). The genetic mechanisms underlying this link are not well understood. METHODS: We downloaded IBD and AMI-related microarray datasets from the NCBI Gene Expression Omnibus (GEO) database. Differentially expressed genes (DEGs) were identified and analyzed using enrichment analysis and Weighted Gene Co-expression Network Analysis (WGCNA). Machine learning techniques, including LASSO, random forest, and Boruta, were employed to screen for hub genes. These genes were validated through qRT-PCR and Western blotting. Single-cell sequencing was used to confirm findings. Additionally, potential therapeutic targets were identified using the Connectivity Map (CMap) database. RESULTS: Five key hub genes-THBD, FOSB, ADGPR3, IL1R2, and PLAUR-were identified as significantly involved in both IBD and AMI pathogenesis. A diagnostic model for AMI constructed using these hub genes demonstrated high predictive accuracy. Single-cell sequencing analysis and several potential drugs targeting these hub genes were identified, offering new therapeutic avenues. CONCLUSION: This study highlights the crucial role of FOSB and other hub genes in the comorbidity of IBD and AMI. The findings provide novel insights for early diagnosis and potential therapeutic strategies, emphasizing the importance of further investigation into these genetic links.

Humans

Genetic targets related to aging for the treatment of coronary artery disease.

BACKGROUND: Coronary Artery Disease (CAD) is the most common cardiovascular disease worldwide, threatening human health, quality of life and longevity. Aging is a dominant risk factor for CAD. This study aims to investigate the potential mechanisms of aging-related genes and CAD, and to make molecular drug predictions that will contribute to the diagnosis and treatment. METHODS: We downloaded the gene expression profile of circulating leukocytes in CAD patients (GSE12288) from Gene Expression Omnibus database, obtained differentially expressed aging genes through "limma" package and GenaCards database, and tested their biological functions. Further screening of aging related characteristic genes (ARCGs) using least absolute shrinkage and selection operator and random forest, generating nomogram charts and ROC curves for evaluating diagnostic efficacy. Immune cells were estimated by ssGSEA, and then combine ARCGs with immune cells and clinical indicators based on Pearson correlation analysis. Unsupervised cluster analysis was used to construct molecular clusters based on ARCGs and to assess functional characteristics between clusters. The DSigDB database was employed to explore the potential targeted drugs of ARCGs, and the molecular docking was carried out through Autodock Vina. Finally, single-cell data (GSE159677) of arterial intima was used to further explore the expression of aging signature genes in different cell subpopulations. RESULTS: We identified 8 ARCGs associated with CAD, in which HIF1A and FGFR3 were up while NOX4, TCF7L2, HK3, CDK18, TFAP4, and ITPK1 were down in CAD patients. Based on this, CAD patients can be divided into two molecular clusters, among which cluster A mainly involves functional pathways such as ECM receptor interaction and focal adhesion; cluster B mainly involves functional pathways such as amimo sugar and nucleotide sugar metabolism and pyrimidine metabolism. In addition, the molecular docking results showed that retinoic acid and resveratrol had good binding affinity with targets genes. Further single-cell analysis results showed that NOX4, TCF7L2, ITPK1, and HIF1A were specifically expressed in different types of cells in atherosclerotic tissues. CONCLUSION: Our study identified several ARCGs that may be involved in the pathogenesis and progression of CAD. Further, retinoic acid and resveratrol were potential candidate molecule drugs for inhibiting these targets.

Humans