PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Random Forest”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Immunogenetic risk and protective factors for the idiopathic inflammatory myopathies: distinct HLA-A, -B, -Cw, -DRB1 and -DQA1 allelic profiles and motifs define clinicopathologic groups in caucasians.

The idiopathic inflammatory myopathies (IIM) are systemic connective tissue diseases in which autoimmune pathology is suspected to promote chronic muscle inflammation and weakness. We have performed low to high resolution genotyping to characterize the allelic profiles of HLA-A, -B, -Cw, -DRB1, and -DQA1 loci in a large population of North American Caucasian patients with IIM representing the major clinicopathologic groups (n = 571). We confirmed that alleles of the 8.1 ancestral haplotype were important risk markers for the development of IIM, and a random forests classification analysis suggested that within this haplotype, HLA-B*0801, DRB1*0301 and/ or closely linked genes are the principal HLA risk factors. In addition, we identified several novel HLA factors associated distinctly with 1 or more clinicopathologic groups of IIM. The DQA1*0201 allele and associated peptide-binding motif (KLPLFHRL) were exclusive protective factors for the CD8+ T cell-mediated IIM forms of polymyositis (PM) and inclusion body myositis (IBM) (pc < 0.005). In contrast, HLA-A*68 alleles were significant risk factors for dermatomyositis (DM) (pc = 0.0021), a distinct clinical group thought to involve a humorally mediated immunopathology. While the DQA1*0301 allele was detected as a possible risk factor for IIM, PM, and DM patients (p < 0.05), DQA1*03 alleles were protective factors for IBM (pc = 0.0002). Myositis associated with malignancies was the most distinctive group of IIM wherein HLA Class I alleles were the only identifiable susceptibility factors and a shared HLA-Cw peptide-binding motif (AGSHTLQWM) conferred significant risk (pc = 0.019). Together, these data suggest that HLA susceptibility markers distinguish different myositis phenotypes with divergent pathogenetic mechanisms. These variations in associated HLA polymorphisms may reflect responses to unique environmental triggers resulting in the tissue pathospecificity and distinct clinicopathologic syndromes of the IIM.

Adult↗

Immunogenetic risk and protective factors for the idiopathic inflammatory myopathies: distinct HLA-A, -B, -Cw, -DRB1, and -DQA1 allelic profiles distinguish European American patients with different myositis autoantibodies.

The idiopathic inflammatory myopathies (IIM) are systemic connective tissue diseases defined by chronic muscle inflammation and weakness associated with autoimmunity. We have performed low to high resolution molecular typing to assess the genetic variability of major histocompatibility complex loci (HLA-A, -B, -Cw, -DRB1, and -DQA1) in a large population of European American patients with IIM (n = 571) representing the major myositis autoantibody groups. We established that alleles of the 8.1 ancestral haplotype (8.1 AH) are important risk factors for the development of IIM in patients producing anti-synthetase/anti-Jo-1, -La, -PM/Scl, and -Ro autoantibodies. Moreover, a random forests classification analysis suggested that 8.1 AH-associated alleles B*0801 and DRB1*0301 are the principal HLA risk markers. In addition, we have identified several novel HLA susceptibility factors associated distinctively with particular myositis-specific (MSA) and myositis-associated autoantibody (MAA) groups of the IIM. IIM patients with anti-PL-7 (anti-threonyl-tRNA synthetase) autoantibodies have a unique HLA Class I risk allele, Cw*0304 (pcorr = 0.046), and lack the 8.1 AH markers associated with other anti-synthetase autoantibodies (for example, anti-Jo-1 and anti-PL-12). In addition, HLA-B*5001 and DQA1*0104 are novel potential risk factors among anti-signal recognition particle autoantibody-positive IIM patients (pcorr = 0.024 and p = 0.010, respectively). Among those patients with MAA, HLA DRB1*11 and DQA1*06 alleles were identified as risk factors for myositis patients with anti-Ku (pcorr = 0.041) and anti-La (pcorr = 0.023) autoantibodies, respectively. Amino acid sequence analysis of the HLA DRB1 third hypervariable region identified a consensus motif, 70D (hydrophilic)/71R (basic)/74A (hydrophobic), conferring protection among patients producing anti-synthetase/anti-Jo-1 and -PM/Scl autoantibodies. Together, these data demonstrate that HLA signatures, comprising both risk and protective alleles or motifs, distinguish IIM patients with different myositis autoantibodies and may have diagnostic and pathogenic implications. Variations in associated polymorphisms for these immune response genes may reflect divergent pathogenic mechanisms and/or responses to unique environmental triggers in different groups of subjects resulting in the heterogeneous syndromes of the IIM.

Alleles↗

Proteomic Immune Signatures of Severe HIV-Associated Tuberculosis in Sub-Saharan Africa: A Prospective, Multicenter Analysis From Uganda.

OBJECTIVES: Severe tuberculosis (TB) is a major cause of critical illness and death in people living with HIV (PLWH) worldwide. Despite this, the immunopathology of severe HIV-associated TB (HIV/TB) is poorly understood. We aimed to identify an immunopathologic signature of severe HIV/TB in sub-Saharan Africa. DESIGN AND SETTING: We analyzed proteomic data from two prospective observational cohorts of adults hospitalized with severe undifferentiated infection in Uganda: an urban discovery cohort (Entebbe, n = 241) and a rural validation cohort (Tororo, n = 253). PATIENTS: Adults (age &#x2265; 18 yr) hospitalized with severe febrile illness. INTERVENTIONS: None. MEASUREMENTS AND MAIN RESULTS: Across both cohorts, severe HIV/TB was common, affecting 18% of participants in the discovery cohort and 21% in the validation cohort. Overall mortality was significant (30-d mortality of 22% in the discovery cohort and 60-d mortality of 26% in the validation cohort). Participants were stratified into three HIV/TB phenotypes: HIV-negative without TB, PLWH without TB, and PLWH with microbiologically diagnosed TB. We applied ordinal random forest models in the discovery cohort as a supervised feature-selection approach to identify proteins associated with progressive HIV/TB phenotype. In both cohorts, PLWH with microbiologically diagnosed TB were at highest risk of critical illness and death (30-d mortality of 42% in the discovery cohort and 60-d mortality of 52% in the validation cohort). An eight-protein signature reliably distinguished this phenotype, reflecting mediators of macrophage/dendritic cell activation (lysosome-associated membrane glycoprotein 3), natural killer cell and T-cell stimulation and cytotoxicity (cluster of differentiation 70, class I-restricted T-cell-associated molecule), B-cell activation (immunoglobulin lambda constant 2), protease-mediated tissue injury (protease, serine 2 [trypsin-2]), dysregulated coagulation (serpin peptidase inhibitor, clade A [alpha-1 antitrypsin], member 5), extracellular matrix remodeling (epidermal growth factor-containing fibulin-like extracellular matrix protein 1), and growth hormone/insulin-like growth factor axis dysregulation (insulin-like growth factor binding protein 3). CONCLUSIONS: We identified an immunologic signature of severe HIV/TB defined by mediators of macrophage/dendritic cell and cytotoxic lymphocyte activation, extracellular matrix remodeling, and dysregulated coagulation. These findings offer new insight into HIV/TB pathobiology and highlight potential targets for host-directed therapies in this high-risk population.

Humans↗

Metabolism pathway-based subtyping in pancreatic adenocarcinoma: an integrated study by bulk RNA-sequence and machine learning algorithms.

BACKGROUND: Pancreatic adenocarcinoma (PAAD) is highly aggressive, and its tumor microenvironment has significant metabolic and immune microenvironment complexity and genomic instability. In this study, by integrating the metabolic pathway activity score and clinical data, we constructed a novel risk assessment model to reveal the unique biological behavior and clinical significance behind different PAAD subtypes. METHODS: In this study, the transcriptome and clinical data of TCGA and GSE57495 databases were integrated to explore the interaction between metabolic pathways. Based on unsupervised clustering analysis of pathway activity and survival prognosis, patients with PAAD were classified into metabolic subtypes with significant prognostic differences. Subsequently, we assessed the heterogeneity of these subtypes in terms of clinical outcomes, genomic characteristics, and immune microenvironment composition. Based on the differentially expressed genes (DEGs) among metabolic subtypes, a clinical prognostic risk model and nomogram were constructed, which were double-validated by GSE57495-independent cohort and GSE57495&#xa0;+&#xa0;TCGA-PAAD combined cohort. Finally, the correlations between risk scores (RSs) and signaling pathway activity and tumor immune microenvironment characteristics were evaluated. RESULTS: Based on metabolic pathway correlation and prognostic information, 240 patients in the TCGA-PAAD and GSE57495 datasets were divided into three subgroups. There were significant differences between subgroups in gene expression, pathway activity, clinical prognosis, and immune infiltration characteristics among the subtypes. Using machine learning algorithms, an RS model was constructed from DEGs among the subgroups, with the random forest method showing the best performance. A nomogram integrating the RS and clinical indicators demonstrated excellent predictive accuracy for 1-, 3-, and 5-year survival rates, confirming the RS as an independent prognostic factor. High- and low-risk groups exhibited significant differences in immune infiltration, pathway activity, and gene mutations. Drug sensitivity analysis showed that the high-risk group was more sensitive to AZD6244, ABT737, and other drugs. CONCLUSION: This study stratified patients with PAAD into three subgroups based on metabolic pathways and prognostic information, revealing significant differences in clinical outcomes, immune characteristics, and genetic mutations. The robust RS model developed from these findings demonstrated strong predictive power for patient survival and identified promising therapeutic strategies, providing valuable insights for advancing precision medicine in PAAD.

immune microenvironment↗

Predicting Weight Loss After Vertical Sleeve Gastrectomy Using a Whole-genome Sequencing-derived Polygenic Risk Score in the All of Us Cohort.

OBJECTIVE: To create a genome-wide polygenic risk score (PRS) to improve prediction of a 12-month percentage weight loss (WL) after vertical sleeve gastrectomy (VSG). BACKGROUND: Variability in post-VSG WL is not well explained by clinical factors. The All of Us program provides access to a 414,830 short-read whole-genome sequencing resource, enabling unbiased discovery of genetic predictors after VSG. METHODS: VSG counts, demographic, anthropomorphic and vital sign information were obtained from the linked electronic health record. The discovery cohort (DC) included participants from version 7 carried into version 8 while the validation cohort (VC) included those newly added to v8. We defined good responders and nonresponders as having WL&#xb1;1SD from the mean. Following quality filtering, we applied a 2-stage penalized-regression, followed by elastic-net logistic regression, to identify 1583 stable variants and derive &#x3b2;-weights. We then tested this PRS on the DC into a prediction model. RESULTS: We identified 395 participants in the DC and 336 participants in the VC, respectively. Of these, VSG, 44 were classified as good responders (&#x2265;37% WL) and 55 as nonresponders (&#x2264;19% WL). In the VC, 55 were classified as good responders and 48 as nonresponders. Adding the PRS to models to clinical predictors increased the area under the curve following logistic regression by 0.03; P <4.3 &#xd7; 10 -14 , random forest by 0.03; P <9.1 &#xd7; 10 -7 , decision tree by 0.05; P = 1.2 &#xd7; 10 -3 , and gradient boosting by 0.08; P <8.3 &#xd7; 10 -10 . CONCLUSIONS: Use of short-read whole-genome sequencing from All of Us (AoU) can be effectively used to generate PRS to enhance predictive WL accuracy. This work has implications for outcomes of both bariatric surgery and other surgical procedures.

Humans↗

Hot spots, indicator taxa, complementarity and optimal networks of taiga.

If hot spots for different taxa coincide, priority-setting surveys in a region could be carried out more cheaply by focusing on indicator taxa. Several previous studies show that hot spots of different taxa rarely coincide. However, in tropical areas indicator taxa may be used in selecting complementary networks to represent biodiversity as a whole. We studied beetles (Coleoptera), Heteroptera, polypores or bracket fungi (Polyporaceae) and vascular plants of old growth boreal taiga forests. Optimal networks for Heteroptera maximized the high overall species richness of beetles and vascular plants, but these networks were least favourable options for polypores. Polypores are an important group indicating the conservation value of old growth taiga forests. Random selection provided a better option. Thus, certain groups may function as good indicators for maximizing the overall species richness of some taxonomic groups, but all taxa should be examined separately.

Animals↗

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results across studies. Here, we performed multiple modeling experiments integrating clinical and demographic data from electronic health records (EHR) and genetic data to understand which decision points may affect performance. Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from two large independent health systems and polygenic risk scores (PRS) were generated across all patients with genetic data in the corresponding biobanks. Crohn's disease was used as the model phenotype based on its substantial genetic component, established EHR-based definition, and sufficient prevalence for model training and testing. We investigated the impact of PRS integration method, as well as choices regarding training sample, model complexity, and performance metrics. Overall, our results show that including PRS resulted in higher performance by some metrics but the gain in performance was only robust when combined with demographic data alone. Improvements were inconsistent or negligible after including additional clinical information. The impact of genetic information on performance also varied by PRS integration method, with a small improvement in some cases from combining PRS with the output of a clinical model (late-fusion) compared to its inclusion an additional feature (early-fusion). The effects of other modeling decisions varied between institutions though performance increased with more compute-intensive models such as random forest. This work highlights the importance of considering methodological decision points in interpreting the impact on prediction performance when including PRS information in clinical models.

Preprint↗

Demographics, Overlap, and Latency of Severe Cutaneous Adverse Reactions in an FDA Database.

IMPORTANCE: Severe cutaneous adverse reactions (SCARs), including Stevens-Johnson syndrome/toxic epidermal necrolysis (SJS-TEN), drug reaction with eosinophilia and systemic symptoms (DRESS), acute generalized exanthematous pustulosis (AGEP), and generalized bullous fixed drug eruption (GBFDE), are rare but life-threatening drug hypersensitivity syndromes. Due to their low incidence and diagnostic complexity, large-scale characterization of SCAR is challenging. OBJECTIVE: To characterize the demographics, causative agents, trends, latency, and phenotypic overlap of SCAR using a large-scale, sanitized pharmacovigilance dataset from FAERS (FDA Adverse Event Reporting System). DESIGN: Cross-sectional study of spontaneous adverse event reports. Cases were drawn from the U.S. Food and Drug Administration Adverse Event Reporting System (FDA FAERS) from January 2004 to December 2023 and subjected to sanitization and deduplication. Disproportionality analysis was used to characterize causative agents. Machine learning (random forest classifiers) was used to analyze predictors of drug latency and mortality. SETTING: Global pharmacovigilance reports submitted to FAERS. PARTICIPANTS: A total of 56,683 deduplicated SCAR reports were identified, representing 0.33% of reports during the study period. EXPOSURES: Suspected causative drugs, including both small molecules and biologics. MAIN OUTCOMES AND MEASURES: Main outcomes included the frequency and distribution of SCAR syndromes, reporting trends over time, latency from drug start to reaction onset, drug-specific disproportionality (PRR, ROR, IC), and co-reporting between SCAR types and related conditions. RESULTS: A total of 56,683 unique SCAR reports were identified, including SJS-TEN (28,871), DRESS (22,444), AGEP (6,183), and GBFDE (150). We identified 237 drugs with significant disproportionality for SCAR overall. Co-reporting between SCARs was significantly enriched (p < 1e-200), suggesting overlapping phenotypes. Latency varied by drug and syndrome (median: GBFDE 3 days, AGEP 4 days, SJS-TEN 12 days, DRESS 20 days). CONCLUSIONS AND RELEVANCE: SCAR syndromes display distinct but overlapping phenotypes, with variable latency and diverse causative agents. These findings, based on the largest SCAR dataset to date, highlight the need for improved classification frameworks and molecular validation. Large-scale pharmacovigilance, integrated with genomic and histopathologic data, will be critical to improving diagnosis, mechanistic understanding, and clinical management of SCAR.

Acute Generalized Exanthematous Pustulosis↗

Integrating explainable AI with multiomics systems biology and EHR data mining for personalized drug repurposing in Alzheimer's disease.

Alzheimer's disease (AD) is characterized by region- and patient-specific molecular heterogeneity, which hinders therapeutic design. In this study, we introduce PRISM-ML (PRecision-medicine using Interpretable Systems and Multiomics with Machine Learning), an open-source integrated analysis pipeline that combines interpretable machine learning with systems biology and electronic health record (EHR) data mining to elucidate the molecular diversity of AD and predict promising drug repurposing opportunities. First, we integrated and harmonized transcriptomic (bulk RNA-seq) and genomic (genome-wide association study) data from 2105 brain samples, each with matched data from the same individual (1363 AD patients, 742 controls; nine tissues), sourced from three independent studies. Random forest classifiers with SHapley Additive exPlanations (SHAP) identified patient-specific biomarkers; unsupervised clustering resolved 36 molecularly distinct "subtissues" (clusters of samples); and gene-gene co-expression networks prioritized 262 high-centrality bottleneck genes as putative regulators of dysregulated pathways. Next, knowledge graph-based drug repurposing predicted six FDA-approved drugs that simultaneously target multiple bottleneck genes and multiple AD-relevant pathways. Notably, in a large U.S. de-identified insurance-claims database (n = 364733), exposure to promethazine, one of the candidate drugs, was associated with a 57-62 % lower incidence of AD versus an active antihistamine comparator (adjusted hazard ratio 0.38; inverse-probability weighted 0.43; both p < 0.001), providing real-world support for its repurposing potential. In summary, PRISM-ML, as an explainable multi-omics analysis pipeline, is readily transferable to other complex diseases, advancing precision medicine.

Computational Biology↗

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 684 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

Journal Article↗

The mouse gut microbiota responds to predator odor and predicts host behavior.

Chronic stressors can alter the mammalian gut microbiota in ways that mediate host stress responses, but the impacts of acute stressors on these interactions are less well understood. Here, we show that brief exposure of wild-derived mice to predator odor altered gut-microbiota composition, which in turn predicted host behavior. We investigated the individual and combined effects of 15-minute exposures to synthetic fox fecal odor and 30 days of chronic social isolation, an established chronic stressor. Using ethological assays, visceral adipose tissue transcriptomics, and genome-resolved metagenomics, we found that predator-odor exposure significantly affected mouse behavior, gene expression, and gut microbiota. Predator odor-responsive bacteria were associated with the expression of genes involved in anti-microbial defense, and host behavioral responses were predicted by random forest models trained on gut-microbiota profiles. These findings indicate interactions between the gut microbiota and wild-mouse responses to the threat of predation, an ecologically relevant acute stressor.

Journal Article↗

Gene Specific Pathogenicity Predictor for Chromatin-Remodeling BAF Complex-Associated Neurodevelopmental Disorders.

Advancements in whole genome sequencing have increased the number of variants of uncertain significance (VUS) identified in patient genomes. This has created a diagnostic bottleneck for genetic counselors tasked with sifting through these variants and determining those most likely to be causative for a patient's clinical presentation. Machine learning (ML) tools can aid in identifying pathogenic variants from VUS, but there is a need for gene-specific algorithms that predict pathogenic variants with high accuracy. To address this need, we present a workflow for developing gene-specific, ensemble-learning ML tools, that leverage outputs from other algorithms, locations of variants within the gene, and evolutionary conservation data to make a prediction of pathogenicity. Variants in SMARCA2 and SMARCA4 that are associated with rare neurodevelopmental diseases were used to screen 15 ML algorithms. A random forest learner was tuned to yield a final accuracy of 0.93 on holdout data. Generalizing this predictor to other BAF complex proteins resulted in a sharp decline in performance. We trained a final predictor for all genes in the study to create a predictor that identifies pathogenic variants in these BAF subunits with an accuracy of 0.91 on holdout data. This predictor specific to BAF complex proteins performs with higher accuracy and AUROC than any other predictor. The decline in performance when generalized to other proteins emphasizes the need for the gene-specific calibration of predictors. Our workflow for the development of such models provides a quick, computationally inexpensive route for improving the ML tools available to genetic counselors.

Journal Article↗

Robust and accurate cancer classification with gene expression profiling.

Robust and accurate cancer classification is critical in cancer treatment. Gene expression profiling is expected to enable us to diagnose tumors precisely and systematically. However, the classification task in this context is very challenging because of the curse of dimensionality and the small sample size problem. In this paper, we propose a novel method to solve these two problems. Our method is able to map gene expression data into a very low dimensional space and thus meets the recommended samples to features per class ratio. As a result, it can be used to classify new samples robustly with low and trustable (estimated) error rates. The method is based on linear discriminant analysis (LDA). However, the conventional LDA requires that the within-class scatter matrix S(w) be nonsingular. Unfortunately, Sw is always singular in the case of cancer classification due to the small sample size problem. To overcome this problem, we develop a generalized linear discriminant analysis (GLDA) that is a general, direct, and complete solution to optimize Fisher's criterion. GLDA is mathematically well-founded and coincides with the conventional LDA when S(w) is nonsingular. Different from the conventional LDA, GLDA does not assume the nonsingularity of S(w), and thus naturally solves the small sample size problem. To accommodate the high dimensionality of scatter matrices, a fast algorithm of GLDA is also developed. Our extensive experiments on seven public cancer datasets show that the method performs well. Especially on some difficult instances that have very small samples to genes per class ratios, our method achieves much higher accuracies than widely used classification methods such as support vector machines, random forests, etc.

Algorithms↗

Automatic tracking, feature extraction and classification of C elegans phenotypes.

This paper presents a method for automatic tracking of the head, tail, and entire body movement of the nematode Caenorhabditis elegans (C. elegans) using computer vision and digital image analysis techniques. The characteristics of the worm's movement, posture and texture information were extracted from a 5-min image sequence. A Random Forests classifier was then used to identify the worm type, and the features that best describe the data. A total of 1597 individual worm video sequences, representing wild type and 15 different mutant types, were analyzed. The average correct classification ratio, measured by out-of-bag (OOB) error rate, was 90.9%. The features that have most discrimination ability were also studied. The algorithm developed will be an essential part of a completely automated C. elegans tracking and identification system.

Algorithms↗

Recognizing plankton images from the shadow image particle profiling evaluation recorder.

We present a system to recognize underwater plankton images from the shadow image particle profiling evaluation recorder (SIPPER). The challenge of the SIPPER image set is that many images do not have clear contours. To address that, shape features that do not heavily depend on contour information were developed. A soft margin support vector machine (SVM) was used as the classifier. We developed a way to assign probability after multiclass SVM classification. Our approach achieved approximately 90% accuracy on a collection of plankton images. On another larger image set containing manually unidentifiable particles, it also provided 75.6% overall accuracy. The proposed approach was statistically significantly more accurate on the two data sets than a C4.5 decision tree and a cascade correlation neural network. The single SVM significantly outperformed ensembles of decision trees created by bagging and random forests on the smaller data set and was slightly better on the other data set. The 15-feature subset produced by our feature selection approach provided slightly better accuracy than using all 29 features. Our probability model gave us a reasonable rejection curve on the larger data set.

Algorithms↗

A Deep Model Framework for Morphological Trait Imputation Across Taxonomic Groups.

Incomplete morphological trait data pose major hurdles for trait-based analyses, particularly when missing values, multicollinearity, and sparse sampling constrain inference. These issues limit our ability to quantify trait variation and explore broad patterns of functional differentiation across taxa. Here, we introduce FS-DeepRBFNet, which overcomes these pitfalls through integrating correlation-based feature selection with a dual-layer adaptive radial basis function (RBF) network. This end-to-end approach effectively reduces noise and captures both linear allometric trends and nonlinear morphological relationships. We tested the framework on a large species-level morphological trait dataset of Chinese birds and further validated its cross-taxon transferability using the Amphibian Database (Caudata). FS-DeepRBFNet consistently outperformed conventional methods such as KNN, Random Forest, and XGBoost, demonstrating superior predictive accuracy across multiple traits. Beyond improvements, the model revealed biologically interpretable trait associations and stable cross-taxon generalization. These results demonstrate that FS-DeepRBFNet provides a robust and biologically grounded solution for morphological trait prediction, enabling reliable imputation for comparative phylogenetics, functional ecology, and biodiversity forecasting in data-limited situations.

cross&#x2010;taxon transferability↗

NR3C1 Modulates Wnt Signalling to Influence the Invasiveness and Immune Features of Nonfunctioning Invasive Pituitary Adenomas.

Pituitary adenomas (PAs) are common intracranial tumours, and invasiveness in nonfunctioning invasive pituitary adenomas (NIPAs) predicts poor prognosis. The molecular mechanisms driving this phenotype remain unclear. This study explored the role of nuclear receptor subfamily 3 group C member 1 (NR3C1) in NIPA invasiveness and its regulation of Wnt signalling. mRNA expression profiles of 32 PA samples were generated by RNA-seq, and proteomic data from 19 samples were obtained by mass spectrometry. Immune-related differentially expressed genes (DEGs) were retrieved from GeneCards. Weighted gene coexpression network analysis identified modules and hub genes linked to invasiveness, while machine learning methods (support vector machine, LASSO, random forest) prioritised key genes. Gene set enrichment analysis (GSEA) assessed pathways associated with candidate gene expression. NR3C1 expression and function were validated by immunohistochemistry, Western blotting and invasion assays. Integration of transcriptomic, proteomic and immune-related datasets yielded 11 overlapping genes, with NR3C1 emerging as the top candidate. NR3C1 was significantly upregulated in NIPAs and demonstrated good discriminatory power by ROC analysis. GSEA associated high NR3C1 expression with Wnt pathway activation. Functional experiments confirmed that NR3C1 overexpression enhances the invasive capacity of PA cells. NR3C1 promotes the invasive phenotype of NIPAs by activating Wnt signalling. These findings suggest NR3C1 as a potential biomarker and therapeutic target for invasive pituitary adenomas.

Humans↗

Integrated Bulk and Single-Cell RNA-Seq Analysis Reveals Transcriptional Activation of PTGS2 by FOS in Progression From T2DM to T2DM-Associated NAFLD.

Type 2 diabetes mellitus (T2DM) and nonalcoholic fatty liver disease (NAFLD) frequently coexist, exacerbating disease burden. However, the molecular mechanisms underlying the progression from T2DM to T2DM-associated NAFLD remain unclear. This study investigated the regulatory function of FOS-mediated PTGS2 activation in this transition. We integrated bulk RNA-seq data from GEO, single-cell transcriptomic data and transcriptomes from patients with T2DM-associated NAFLD. Differentially expressed genes were identified using the limma package, and T2DM-related gene modules were defined by weighted gene co-expression network analysis. LASSO regression and random forest identified 14 candidate genes, with PTGS2 and FOS prioritised. Single-cell analysis showed increased FOS and PTGS2 expression in monocytes, CD8+ T cells and Kupffer cells. Transcription factor prediction and dual-luciferase assays confirmed that FOS directly binds the PTGS2 promoter and drives its transcription. In&#xa0;vitro, FOS silencing decreased PTGS2 expression, cytokine secretion and apoptosis under high-glucose and free fatty acid conditions, whereas PTGS2 overexpression exacerbated inflammation and apoptosis independently of FOS expression. These findings demonstrate that FOS transcriptionally activates PTGS2, contributing to hepatic inflammation and apoptosis during the progression from T2DM to NAFLD. PTGS2 may serve as a promising biomarker and therapeutic target for T2DM-associated NAFLD.

Single-Cell Gene Expression Analysis↗