PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Random Forest”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Integrative Multi-Omics Analysis of Stem Growth Habit Divergence in Wild Soybean (Glycine soja).

Stem architecture is a major determinant of lodging resistance, biomass accumulation, and harvest efficiency in soybean. However, the molecular features associated with contrasting stem growth habits in wild soybean remain incompletely characterised. Here, we performed an integrated transcriptomic, metabolomic, and epigenomic analysis of stem growth-habit divergence in wild soybean, comparing the wild-type accession ZYD7068 with contrasting vining and erect mutant lines derived from carbon-ion beam mutagenesis. Pairwise transcriptomic comparisons identified between 20 311 and 28 705 differentially expressed genes per contrast, with a core set of 2672 genes consistently altered across the comparisons. Functional enrichment, gene set variation analysis, and gene set enrichment analysis converged on xylem and phloem pattern formation as a prominent molecular pathway associated with growth-habit divergence. Random forest analysis identified BBR-BPC and ARF transcription factor families as major molecular discriminators, while metabolomic profiling revealed distinct metabolic profiles involving amino-acid-derived and lipid-associated metabolites. Whole-genome bisulfite sequencing revealed context-specific DNA methylation differences, including substantial variation in CHG methylation among erect mutant lines. Integrated network and in silico perturbation analyses prioritised four candidate genes associated with vascular development for future functional validation. Together, these results provide a multi-layer molecular resource for investigating stem growth-habit divergence in G. soja and establish testable candidate pathways and genes for subsequent functional studies and soybean improvement.

glycine soja↗

Spatiotemporal mapping of tertiary lymphoid structure heterogeneity shapes immune niches and clinical outcomes in intrahepatic cholangiocarcinoma.

Intrahepatic cholangiocarcinoma (iCCA) is a highly lethal malignancy with limited therapeutic options. The spatial architecture and functional diversity of tertiary lymphoid structures (TLSs) in iCCA remain unclear. Here, we present a multimodal spatial atlas of TLSs and identified intratumoral TLSs (iTLSs) as independent prognostic markers. Bulk proteomic profiling of 214 discovery and 155 validation cases identified a four-tier TLS-based tumor microenvironment classification system and supported development of a TLS-predictive random forest classifier. Imaging mass cytometry revealed that iTLS+ tumors harbor structured immune architectures, where M1-like tissue-resident macrophages (RTMs), dendritic cells, and CXCL13+ CD4+ T cells colocalize to form antigen-presenting neighborhoods (apc-CNs) spatially coupled to TLS core regions (TLScore-CNs). Single-cell spatial transcriptomics further resolved 61 TLSs into 14 spatial niches and defined a pseudotemporal maturation continuum: aggregated, activated, and postactivated. Intraniche communication, primarily mediated by ifnCAFs, iCAFs, and CXCL12+ macrophages, evolved dynamically with maturation. Single-nucleus RNA sequencing combined with Tangram-based spatial mapping revealed CXCL12+ macrophages and iCAFs forming a peripheral band in aggregated TLSs, whereas ifnCAFs infiltrated TLS interiors during activation. These findings define TLS heterogeneity and provide insights for stroma-directed immunotherapy.

Cholangiocarcinoma↗

Microarray analysis in drug discovery: an uplifting view of depression.

Genomic profiling provides insights into drug evaluation for diseases without defined molecular mechanisms or cellular assays. Levy provides a brief background in the development of microarray analysis and discussion of the application of this technique to pharmacogenomics. Highlighted is the microarray analysis of primary human neurons treated with antidepressants, antipsychotics, or opioid receptor agonists, demonstrating that these classes of drugs can be properly categorized by using two different statistical analysis methods: classification tree and random forest. Not only is microarray analysis valuable for drug evaluation and leading candidate development, but the genes identified as markers for the various drug classifications point to new directions for research into the underlying pathways responsible for human diseases, such as depression and psychosis.

Animals↗

Biomarkers that discriminate multiple myeloma patients with or without skeletal involvement detected using SELDI-TOF mass spectrometry and statistical and machine learning tools.

Multiple Myeloma (MM) is a severely debilitating neoplastic disease of B cell origin, with the primary source of morbidity and mortality associated with unrestrained bone destruction. Surface enhanced laser desorption/ionization time-of-flight mass spectrometry (SELDI-TOF MS) was used to screen for potential biomarkers indicative of skeletal involvement in patients with MM. Serum samples from 48 MM patients, 24 with more than three bone lesions and 24 with no evidence of bone lesions were fractionated and analyzed in duplicate using copper ion loaded immobilized metal affinity SELDI chip arrays. The spectra obtained were compiled, normalized, and mass peaks with mass-to-charge ratios (m/z) between 2000 and 20,000 Da identified. Peak information from all fractions was combined together and analyzed using univariate statistics, as well as a linear, partial least squares discriminant analysis (PLS-DA), and a non-linear, random forest (RF), classification algorithm. The PLS-DA model resulted in prediction accuracy between 96-100%, while the RF model was able to achieve a specificity and sensitivity of 87.5% each. Both models as well as multiple comparison adjusted univariate analysis identified a set of four peaks that were the most discriminating between the two groups of patients and hold promise as potential biomarkers for future diagnostic and/or therapeutic purposes.

Adult↗

SNP-based analysis of genetic substructure in the German population.

OBJECTIVE: To evaluate the relevance and necessity to account for the effects of population substructure on association studies under a case-control design in central Europe, we analysed three samples drawn from different geographic areas of Germany. Two of the three samples, POPGEN (n = 720) and SHIP (n = 709), are from north and north-east Germany, respectively, and one sample, KORA (n = 730), is from southern Germany. METHODS: Population genetic differentiation was measured by classical F-statistics for different marker sets, either consisting of genome-wide selected coding SNPs located in functional genes, or consisting of selectively neutral SNPs from 'genomic deserts'. Quantitative estimates of the degree of stratification were performed comparing the genomic control approach [Devlin B, Roeder K: Biometrics 1999;55:997-1004], structured association [Pritchard JK, Stephens M, Donnelly P: Genetics 2000;155:945-959] and sophisticated methods like random forests [Breiman L: Machine Learning 2001;45:5-32]. RESULTS: F-statistics showed that there exists a low genetic differentiation between the samples along a north-south gradient within Germany (F(ST)(KORA/POPGEN): 1.7 . 10(-4); F(ST)(KORA/SHIP): 5.4 . 10(-4); F(ST)(POPGEN/SHIP): -1.3 . 10(-5)). CONCLUSION: Although the F(ST )-values are very small, indicating a minor degree of population structure, and are too low to be detectable from methods without using prior information of subpopulation membership, such as STRUCTURE [Pritchard JK, Stephens M, Donnelly P: Genetics 2000;155:945-959], they may be a possible source for confounding due to population stratification.

Case-Control Studies↗

Data-mining methods as useful tools for predicting individual drug response: application to CYP2D6 data.

OBJECTIVES: Selecting a maximally informative subset of polymorphisms to predict a clinical outcome, such as drug response, requires appropriate search methods due to the increased dimensionality associated with looking at multiple genotypes. In this study, we investigated the ability of several pattern recognition methods to identify the most informative markers in the CYP2D6 gene for the prediction of CYP2D6 metabolizer status. METHODS: Four data-mining tools were explored: decision trees, random forests, artificial neural networks, and the multifactor dimensionality reduction (MDR) method. Marker selection was performed separately in eight population samples of different ethnic origin to evaluate to what extent the most informative markers differ across ethnic groups. RESULTS: Our results show that the number of polymorphisms required to predict CYP2D6 metabolic phenotype with a high accuracy can be dramatically reduced owing to the strong haplotype block structure observed at CYP2D6. MDR and neural networks provided nearly identical results and performed the best. CONCLUSION: Data-mining methods, such as MDR and neural networks, appear as promising tools to improve the efficiency of genotyping tests in pharmacogenetics with the ultimate goal of pre-screening patients for individual therapy selection with minimum genotyping effort.

Cytochrome P-450 CYP2D6↗

Multiplex bead immunoassay analysis of aqueous humor reveals distinct cytokine profiles in uveitis.

PURPOSE: To extensively characterize the complex network of cytokines present in uveitis aqueous humor (AqH), and the relationships between cytokines and the cellular infiltrate. METHODS: AqH from noninflammatory control subjects and patients with idiopathic, Fuchs' heterochromic cyclitis (FHC), and herpes-viral or Behçet's uveitis were analyzed for IL-1beta, -2, -4, -5, -7, -8, -10, -12, -13, -15, TNFalpha, IFNgamma, CCL2 (MCP-1), CCL5 (RANTES), CCL11 (Eotaxin), TGFbeta2, and CXCL12 (SDF-1), using multiplex bead immunoassays. The cellular infiltrate was also determined for each sample. RESULTS: Idiopathic uveitis AqH, compared with noninflammatory controls, was characterized by high levels of IL-6, IL-8, CCL2 and IFNgamma, the levels of which correlated with each other. For IL-6 and IL-8 these levels were proportional to the number of neutrophils present. By contrast, the levels of both TGFbeta2 and CXCL12 decreased in idiopathic uveitis AqH with increasing inflammation. Cluster analysis showed a degree of segregation between noninflammatory and idiopathic uveitis AqH. Further examination using random forest analysis yielded a complete distinction between these two groups. The minimum cytokines required for this classification were IL-6, IL-8, CCL2, IL-13, TNFalpha, and IL-2. CONCLUSIONS: Application of multiplex bead immunoassays has allowed us to identify distinct patterns of cytokines that relate to both clinical disease and the cellular infiltrates present. Bioinformatics analysis allowed identification of cytokines that differentiate idiopathic uveitis from noninflammatory control AqH and are likely to be important for the pathogenesis of uveitis.

Adolescent↗

Joint analysis of two microarray gene-expression data sets to select lung adenocarcinoma marker genes.

BACKGROUND: Due to the high cost and low reproducibility of many microarray experiments, it is not surprising to find a limited number of patient samples in each study, and very few common identified marker genes among different studies involving patients with the same disease. Therefore, it is of great interest and challenge to merge data sets from multiple studies to increase the sample size, which may in turn increase the power of statistical inferences. In this study, we combined two lung cancer studies using microarray GeneChip, employed two gene shaving methods and a two-step survival test to identify genes with expression patterns that can distinguish diseased from normal samples, and to indicate patient survival, respectively. RESULTS: In addition to common data transformation and normalization procedures, we applied a distribution transformation method to integrate the two data sets. Gene shaving (GS) methods based on Random Forests (RF) and Fisher's Linear Discrimination (FLD) were then applied separately to the joint data set for cancer gene selection. The two methods discovered 13 and 10 marker genes (5 in common), respectively, with expression patterns differentiating diseased from normal samples. Among these marker genes, 8 and 7 were found to be cancer-related in other published reports. Furthermore, based on these marker genes, the classifiers we built from one data set predicted the other data set with more than 98% accuracy. Using the univariate Cox proportional hazard regression model, the expression patterns of 36 genes were found to be significantly correlated with patient survival (p < 0.05). Twenty-six of these 36 genes were reported as survival-related genes from the literature, including 7 known tumor-suppressor genes and 9 oncogenes. Additional principal component regression analysis further reduced the gene list from 36 to 16. CONCLUSION: This study provided a valuable method of integrating microarray data sets with different origins, and new methods of selecting a minimum number of marker genes to aid in cancer diagnosis. After careful data integration, the classification method developed from one data set can be applied to the other with high prediction accuracy.

Adenocarcinoma↗

Using the nucleotide substitution rate matrix to detect horizontal gene transfer.

BACKGROUND: Horizontal gene transfer (HGT) has allowed bacteria to evolve many new capabilities. Because transferred genes perform many medically important functions, such as conferring antibiotic resistance, improved detection of horizontally transferred genes from sequence data would be an important advance. Existing sequence-based methods for detecting HGT focus on changes in nucleotide composition or on differences between gene and genome phylogenies; these methods have high error rates. RESULTS: First, we introduce a new class of methods for detecting HGT based on the changes in nucleotide substitution rates that occur when a gene is transferred to a new organism. Our new methods discriminate simulated HGT events with an error rate up to 10 times lower than does GC content. Use of models that are not time-reversible is crucial for detecting HGT. Second, we show that using combinations of multiple predictors of HGT offers substantial improvements over using any single predictor, yielding as much as a factor of 18 improvement in performance (a maximum reduction in error rate from 38% to about 3%). Multiple predictors were combined by using the random forests machine learning algorithm to identify optimal classifiers that separate HGT from non-HGT trees. CONCLUSION: The new class of HGT-detection methods introduced here combines advantages of phylogenetic and compositional HGT-detection techniques. These new techniques offer order-of-magnitude improvements over compositional methods because they are better able to discriminate HGT from non-HGT trees under a wide range of simulated conditions. We also found that combining multiple measures of HGT is essential for detecting a wide range of HGT events. These novel indicators of horizontal transfer will be widely useful in detecting HGT events linked to the evolution of important bacterial traits, such as antibiotic resistance and pathogenicity.

Computational Biology↗

Comprehensive in silico genomics analysis of global trends and host-specific emergence of aminoglycoside resistance in Staphylococcus aureus: a One-Health perspective.

BACKGROUND: Aminoglycosides remain clinically valuable against Staphylococcus aureus. Aminoglycoside resistance in S. aureus represents a critical One Health concern and is primarily driven by aminoglycoside-modifying enzymes (AMEs), which are frequently plasmid-encoded. Although regional studies have provided valuable insights, the global epidemiology of aminoglycoside resistance determinants remains poorly characterized because comprehensive data integrating human, animal, and environmental reservoirs are still lacking. This study addresses this gap by analyzing over 110,000 S. aureus genomes (2000-2025) to map the global resistome, quantify temporal and host-specific trends, and assess the association between genetic determinants and phenotypic resistance. METHODS: We performed a retrospective One Health meta-analysis of 110,309 S. aureus genomes collected between 2000 and 2025 from 128 countries. Genomes were quality-filtered and aminoglycoside resistance determinants were identified using NCBI AMRFinderPlus (v4.0.23). Multilocus sequence typing and host-source harmonization (Human, Animal, Environment, Unknown) enabled clonal and reservoir stratification. Temporal trends in gene prevalence and resistance burden were modeled with robust regression. Geographic and host-associated structuring of key genes was assessed via &#x3c7;2 and enrichment tests. Machine-learning models (elastic-net, random forests, XGBoost) were benchmarked for minimum inhibitory concentration (MIC) prediction via nested cross-validation, with performance evaluated by mean absolute error, RMSE, and SHAP-based feature importance. All analyses were conducted in R and Python using publicly available, de-identified genomic data. RESULTS: Aminoglycoside resistance-associated genes were dominated by modifying enzyme determinants, with ant(6)-Ia, ant(9)-Ia, aph(3')-IIIa, sat4, aadD1, and aac(6')-Ie/aph(2'')-Ia occurring in 14-22% of isolates worldwide. Temporal analysis revealed significant declines in several major determinants, most notably ant(9)-Ia (-2.22 percentage points per year, p&#x2009;<&#x2009;0.001), whereas apmA exhibited a non-significant decreasing trend in animal isolates. Host structuring was marked: human clinical isolates concentrated common determinants, while animal and environmental isolates harbored rare alleles (apmA, spw, str, spd). Geographic mapping confirmed near-universal distribution of common genes but focal restriction of rare ones. Publicly available phenotypic data indicated strong activity of amikacin, whereas gentamicin showed a distinct resistant subpopulation that closely corresponded with AME gene carriage. Genotype-phenotype analyses demonstrated strong concordance, with gene-rich complements predicting resistant MIC strata and absence of determinants predicting susceptibility. Analysis across different gene classes revealed frequent co-occurrence of aminoglycoside resistance genes with determinants from other classes, such as mecA, blaZ, and MLS_B, embedding them within multidrug-resistant (MDR) genomic contexts. CONCLUSION: Over 25&#xa0;years, the prevalence of aminoglycoside resistance-associated genes in S. aureus has declined for several common determinants, while rare veterinary-linked alleles are emerging in animal isolates. Strong genotype-phenotype concordance supports genomic prediction for gentamicin and amikacin, where MIC data are available, although phenotypic confirmation remains essential. The frequent co-occurrence of aminoglycoside resistance genes with other antimicrobial resistance determinants indicates their integration within co-occurrence patterns of MDR genes, defined here as clusters of co-occurring resistance genes often carried on shared mobile genetic elements. These patterns highlight the need for integrated One Health surveillance combining clinical, veterinary, and environmental monitoring with plasmid-context resolution to anticipate emerging threats.

Aminoglycosides↗

Potential evaluation of SULT1A3 as an early diagnostic marker for nasopharyngeal carcinoma: a study based on serum proteomics screening and ELISA validation.

BACKGROUND: Nasopharyngeal carcinoma (NPC) represents a highly prevalent and aggressive malignancy endemic to Southeast Asia. Early and accurate diagnosis is critical to improving survival outcomes; however, the absence of robust, stage-specific biomarkers remains a key obstacle to clinical implementation of early screening strategies. METHODS: We performed untargeted serum proteomic profiling using mass spectrometry in 15 treatment-na&#xef;ve early-stage NPC patients and 15 VCA-IgA-positive healthy controls. Bioinformatics analyses were conducted to identify differentially expressed proteins (DEPs). Machine learning (random forest combined with recursive feature elimination) was employed to prioritize candidate biomarkers, which were subsequently verified using enzyme-linked immunosorbent assay (ELISA) in independent sample cohorts. RESULTS: In total, 1,428 serum proteins were identified, among which 1,410 were reliably quantified. We observed 31 upregulated and 189 downregulated proteins in NPC patients relative to controls. Spearman correlation analysis revealed significant associations: LTA4H (leukotriene A4 hydrolase) levels correlated with serum cell infiltration (r&#x2009;=&#x2009;0.383, p&#x2009;=&#x2009;0.032) and CD8&#x2009;+&#x2009;T-cell abundance (r&#x2009;=&#x2009;0.408, p&#x2009;=&#x2009;0.021); both SULT1A3 (sulfotransferase family 1&#xa0;A member 3) and FGL1 (fibrinogen-like protein 1) levels were positively associated with M1 macrophage infiltration (r&#x2009;=&#x2009;0.510, p&#x2009;=&#x2009;0.003 and r&#x2009;=&#x2009;0.430, p&#x2009;=&#x2009;0.015, respectively). In a preliminary validation cohort (n&#x2009;=&#x2009;80), ELISA yielded AUC values of 0.631 (95% CI: 0.515-0.736, p&#x2009;=&#x2009;0.04) for LTA4H, 0.787 (95% CI: 0.681-0.871, p&#x2009;<&#x2009;0.001) for SULT1A3, and 0.688 (95% CI: 0.575-0.787, p&#x2009;=&#x2009;0.002) for FGL1. In large-scale independent validation, SULT1A3 achieved an AUC of 0.826 (95% CI: 0.766-0.876; sensitivity&#x2009;=&#x2009;78.89%, specificity&#x2009;=&#x2009;75.47%) in cohort 1 (n&#x2009;=&#x2009;196) and 0.796 (95% CI: 0.723-0.857; sensitivity&#x2009;=&#x2009;76.67%, specificity&#x2009;=&#x2009;76.67%) in cohort 2 (n&#x2009;=&#x2009;150). CONCLUSIONS: Through an integrated workflow combining proteomic screening, machine learning prioritization, and multi-stage ELISA validation, we identified SULT1A3 as a candidate serum-based biomarker for early detection of NPC. Preliminary findings suggest that SULT1A3 may have potential utility in clinical screening, though further validation in independent, multi&#x2011;center cohorts is required.

Humans↗

Development and evaluation of a machine learning model for osteoporosis risk prediction in Korean women.

BACKGROUND: The aim of this study was to develop a machine learning (ML) model for classifying osteoporosis in Korean women based on a large-scale population cohort study. This study also aimed to assess ML model performance compared with traditional osteoporosis screening tools. Furthermore, this study aimed to examine the factors influencing the risk of osteoporosis through variable importance. METHODS: Data was collected from 4199 women aged 40-69 years in the baseline survey of the Ansan and Ansung cohort of the Korean Genome and Epidemiology Study. Osteoporosis was set as the dependent variable to develop ML classification models. Independent variables included 122 factors related to osteoporosis risk, such as socio-demographic characteristics, anthropometric parameters, lifestyle factors, reproductive factors, nutrient intakes, diet quality indices, medical history, medication history, family history, biochemical parameters, and genetic factors. The six classification models were developed using ML techniques, including decision tree, random forest, multilayer perceptron, support vector machine, light gradient boosting machine, and extreme gradient boosting (XGBoost). The six ML classification models were compared with two traditional osteoporosis screening tools, including the osteoporosis risk assessment instrument (ORAI) and the osteoporosis self-assessment tool (OST). The ML model performances were evaluated and compared using the confusion matrix and area under the curve (AUC) metrics. Variable importance was assessed using the XGBoost technique to investigate osteoporosis risk factors. RESULTS: The XGBoost model showed the highest performance out of the six ML classification models, with an accuracy of 0.705, precision of 0.664, recall of 0.830, and F1 score of 0.738. Moreover, the XGBoost model showed a higher performance on AUC than ORAI and OST. Variable importance scores were identified for 69 out of the 122 variables associated with osteoporosis risk factors. Age at menopause ranked first in variable importance. Variables of arthritis, physical activities, hypertension, education level, income level; alcohol intake, potassium intake, homeostatic model assessment for insulin resistance; energy intake, vitamin C intake, gout; and dietary inflammatory index ranked in the top 20 out of the 69 variables, using the XGBoost technique. CONCLUSIONS: This study found that an XGBoost model can be utilized to classify osteoporosis in Korean women. Age at menopause is a significant factor in osteoporosis risk, followed by arthritis, physical activities, hypertension, and education level.

Humans↗

Systematic review of machine learning approaches for predicting sickle cell crisis and mortality risk at the climate-health nexus.

BACKGROUND: Sickle cell anemia (SCA) is a severe genetic blood disorder characterized by recurrent vaso-occlusive crises and increased mortality, with the greatest burden occurring in low- and middle-income countries. Climatic and environmental conditions, including temperature variability, humidity, rainfall, air pollution, and seasonal changes, have been associated with disease exacerbation. However, the extent to which these factors have been incorporated into predictive models remains unclear. This study systematically reviews the application of machine learning (ML) models for predicting SCA crises and mortality in relation to climate and environmental factors. METHODOLOGY: The PRISMA guidelines were used, and 34 peer-reviewed studies published between 2005 and 2026 were analyzed to identify the climate variables, ML approaches employed, and predictive performance. The reviewed studies applied a range of ML techniques, including artificial neural networks, random forests, support vector machines, decision trees, logistic regression, and deep learning models. Temperature, humidity, rainfall, wind speed, air quality indicators, and seasonal patterns were the most frequently examined environmental variables. RESULTS: The findings indicate that most existing models rely predominantly on clinical and demographic data, with limited integration of climate information and inadequate representation of high-burden regions, especially Sub-Saharan Africa. Studies incorporating environmental variables reported improved predictive performance and highlighted the potential of climate-informed early warning systems for SCA management. CONCLUSION: The review recommends development of interdisciplinary, climate-aware ML frameworks, expansion of longitudinal environmental datasets, and increased research in underrepresented regions to support climate-resilient and patient-centered SCA care.

Humans↗

Major depletion of insulin sensitivity-associated taxa in the gut microbiome of persons living with HIV controlled by antiretroviral drugs.

BACKGROUND: Persons living with HIV (PWH) harbor an altered gut microbiome (higher abundance of Prevotella and lower abundance of Bacillota and Ruminococcus lineages) compared to non-infected individuals. Some of these alterations are linked to sexual preference and others to the HIV infection. The relationship between these lineages and metabolic alterations, often present in aging PWH, has been poorly investigated. METHODS: In this study, we compared fecal metagenomes of 25 antiretroviral-treatment (ART)-controlled PWH to three independent control groups of 25 non-infected matched individuals by means of univariate analyses and machine learning methods. Moreover, we used two external datasets to validate predictive models of PWH classification. Next, we searched for associations between clinical and biological metabolic parameters with taxonomic and functional microbiome profiles. Finally, we compare the gut microbiome in 7 PWH after a 17-week ART switch to raltegravir/maraviroc. RESULTS: Three major enterotypes (Prevotella, Bacteroides and Ruminococcaceae) were present in all groups. The first Prevotella enterotype was enriched in PWH, with several of characteristic lineages associated with poor metabolic profiles (low HDL and adiponectin, high insulin resistance (HOMA-IR)). Conversely butyrate-producing lineages were markedly depleted in PWH independently of sexual preference and were associated with a better metabolic profile (higher HDL and adiponectin and lower HOMA-IR). Accordingly with the worst metabolic status of PWH, butyrate production and amino-acid degradation modules were associated with high HDL and adiponectin and low HOMA-IR. Random Forest models trained to classify PWH vs. control on taxonomic abundances displayed high generalization performance on two external holdout datasets (ROC AUC of 80-82%). Finally, no significant alterations in microbiome composition were observed after switching to raltegravir/maraviroc. CONCLUSION: High resolution metagenomic analyses revealed major differences in the gut microbiome of ART-controlled PWH when compared with three independent matched cohorts of controls. The observed marked insulin resistance could result both from enrichment in Prevotella lineages, and from the depletion in species producing butyrate and involved into amino-acid degradation, which depletion is linked with the HIV infection.

Humans↗

Identification of key immune-related genes and potential therapeutic drugs in diabetic nephropathy based on machine learning algorithms.

BACKGROUND: Diabetic nephropathy (DN) is a major contributor to chronic kidney disease. This study aims to identify immune biomarkers and potential therapeutic drugs in DN. METHODS: We analyzed two DN microarray datasets (GSE96804 and GSE30528) for differentially expressed genes (DEGs) using the Limma package, overlapping them with immune-related genes from ImmPort and InnateDB. LASSO regression, SVM-RFE, and random forest analysis identified four hub genes (EGF, PLTP, RGS2, PTGDS) as proficient predictors of DN. The model achieved an AUC of 0.995 and was validated on GSE142025. Single-cell RNA data (GSE183276) revealed increased hub gene expression in epithelial cells. CIBERSORT analysis showed differences in immune cell proportions between DN patients and controls, with the hub genes correlating positively with neutrophil infiltration. Molecular docking identified potential drugs: cysteamine, eltrombopag, and DMSO. And qPCR and western blot assays were used to confirm the expressions of the four hub genes. RESULTS: Analysis found 95 and 88 distinctively expressed immune genes in the two DN datasets, with 14 consistently differentially expressed immune-related genes. After machine learning algorithms, EGF, PLTP, RGS2, PTGDS were identified as the immune-related hub genes associated with DN. In addition, the mRNA and protein levels of them were obviously elevated in HK-2 cells treated with glucose for 24&#xa0;h, as well as their mRNA expressions in kidney tissues of mice with DN. CONCLUSION: This study identified 4 hub immune-related genes (EGF, PLTP, RGS2, PTGDS), as well as their expression profiles and the correlation with immune cell infiltration in DN.

Diabetic Nephropathies↗

Prediction of metabolic syndrome using machine learning approaches based on genetic and nutritional factors: a 14-year prospective-based cohort study.

INTRODUCTION: Metabolic syndrome is a chronic disease associated with multiple comorbidities. Over the last few years, machine learning techniques have been used to predict metabolic syndrome. However, studies incorporating demographic, clinical, laboratory, dietary, and genetic factors to predict the incidence of metabolic syndrome in Koreans are limited. In the present study, we propose a genome-wide polygenic risk score for the prediction of metabolic syndrome, along with other factors, to improve the prediction accuracy of metabolic syndrome. METHODS: We developed 7 machine learning-based models and used Cox multivariable regression, deep neural network (DNN), support vector machine (SVM), stochastic gradient descent (SGD), random forest (RAF), Na&#xef;ve Bayes (NBA) classifier,&#xa0;and AdaBoost (ADB) to predict the incidence of metabolic syndrome at year 14 using the dataset from the Korean Genome and Epidemiology Study (KoGES) Ansan and Ansung. RESULTS: Of the 5440 patients, 2,120 were considered to have new-onset metabolic syndrome. The AUC values of model, which included sex, age, alcohol intake, energy intake, marital status, education status, income status, smoking status, dried laver intake, and genome-wide polygenic risk score (gPRS)&#xa0;Z-score based on 344,447 SNPs (p-value&#x2009;<&#x2009;1.0), were the highest for RAF (0.994 [95% CI 0.985, 1.000]) and ADB (0.994 [95% CI 0.986, 1.000]). CONCLUSIONS: Incorporating both gPRS and demographic, clinical, laboratory, and seaweed data led to enhanced metabolic syndrome risk prediction by capturing the distinct etiologies of metabolic syndrome development. The RAF- and ADB-based models predicted metabolic syndrome more accurately than the NBA-based model for the Korean population.

Humans↗

A machine learning model and identification of immune infiltration for chronic obstructive pulmonary disease based on disulfidptosis-related genes.

BACKGROUND: Chronic obstructive pulmonary disease (COPD) is a chronic and progressive lung disease. Disulfidptosis-related genes (DRGs) may be involved in the pathogenesis of COPD. From the perspective of predictive, preventive, and personalized medicine (PPPM), clarifying the role of disulfidptosis in the development of COPD could provide a opportunity for primary prediction, targeted prevention, and personalized treatment of the disease. METHODS: We analyzed the expression profiles of DRGs and immune cell infiltration in COPD patients by using the GSE38974 dataset. According to the DRGs, molecular clusters and related immune cell infiltration levels were explored in individuals with COPD. Next, co-expression modules and cluster-specific differentially expressed genes were identified by the Weighted Gene Co-expression Network Analysis (WGCNA). Comparing the performance of the random forest (RF), support vector machine (SVM), generalized linear model (GLM), and eXtreme Gradient Boosting (XGB), we constructed the ptimal machine learning model. RESULTS: DE-DRGs, differential immune cells and two clusters were identified. Notable difference in DRGs, immune cell populations, biological processes, and pathway behaviors were noted among the two clusters. Besides, significant differences in DRGs, immune cells, biological functions, and pathway activities were observed between the two clusters.A nomogram was created to aid in the practical application of clinical procedures. The SVM model achieved the best results in differentiating COPD patients across various clusters. Following that, we identified the top five genes as predictor genes via SVM model. These five genes related to the model were strongly linked to traits of the individuals with COPD. CONCLUSION: Our study demonstrated the relationship between disulfidptosis and COPD and established an optimal machine-learning model to evaluate the subtypes and traits of COPD. DRGs serve as a target for future predictive diagnostics, targeted prevention, and individualized therapy in COPD, facilitating the transition from reactive medical services to PPPM in the management of the disease.

Pulmonary Disease, Chronic Obstructive↗

Exploring diagnostic m6A regulators in primary open-angle glaucoma: insight from gene signature and possible mechanisms by which key genes function.

PURPOSE: The purpose of this study was to interrogate the potential role of N6-methyladenosine (m6A) regulators in the process of trabecular meshwork (TM) tissue damage in patients with primary open-angle glaucoma (POAG). METHODS: Firstly, the expression profile of m6A regulators in TM tissues of POAG patients was comprehensively analyzed by bioinformatics analysis; Plasmid transfection and siRNA gene interference were used to enhance or weaken the expression levels of YTHDC2 in human trabecular meshwork cells (HTMCs); Cell migration ability was detected by transwell chamber assay; Immunofluorescence staining assay was used to evaluate the expression of extracellular matrix (ECM) related proteins. RESULTS: Through the analysis of GSE27276 database, 5 m6A regulators with different expression in POAG were screened out. The results of random forest model showed that these 5 m6A regulators exhibited diagnostic potential and were characteristic genes of POAG. All POAG samples could be effectively divided into two groups based on the expression levels of these 5 hub m6A regulators. Immune cell infiltration analysis indicated that the levels of activated CD8+ T cells and regulatory T cells were different in the two subtypes. HTMC oxidative stress cell model and TGF-&#x3b2;2 stimulation cell model were further constructed to verify the expression of the aforementioned hub m6A regulators, and it was found that YTHDC2 mRNA showed the same expression trend in both models. The silencing of YTHDC2 enhanced the migration ability of HTMCs and increased the synthesis ability of ECM. However, when YTHDC2&#x394;YTH, which lacks the YTH domain, is overexpressed in HTMCs, there is no significant change in the ECM synthesis ability. CONCLUSIONS: The differentially expressed m6A regulators in TM tissues may serve as potential diagnostic biomarkers for POAG. And, in HTMCs, the expression level of YTHDC2 mRNA was changed under oxidative stress or TGF-&#x3b2;2 intervention, and then exerted its regulation on cell migration and ECM synthesis capability through m6A modification, which may be an important part of the disease process of POAG.

Humans↗