PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Random Forest”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Artificial intelligence (AI) uses in stereotactic radiosurgery (SRS): diagnosis with brain metastasis (BM) - A systematic review.

BACKGROUND: Brain metastases (BM) are the most common intracranial tumors in adults, and stereotactic radiosurgery (SRS) has become a mainstay of management. However, several diagnostic challenges persist in the SRS pathway, particularly the differentiation of radiation necrosis (RN) from true tumor progression, which conventional MRI and even advanced imaging techniques often cannot reliably resolve. Recent advances in artificial intelligence (AI) offer the potential to address these diagnostic limitations. This systematic review synthesizes current literature on AI applications for MRI-based diagnostic decision support in BM patients undergoing SRS, with a focus on radiomics and deep learning tools for distinguishing RN from progression, classifying molecular and histologic subtypes, and predicting treatment response. METHODS: A systematic review was performed in accordance with PRISMA guidelines. PubMed, Web of Science, and Scopus were searched using a targeted query combining terms related to AI, brain metastasis, diagnosis or imaging, and SRS. After screening 483 records and applying strict inclusion and exclusion criteria, 18 studies published between 2015 and 2025 were included. Data were extracted on study design, cohort characteristics, imaging modality, AI methodology, validation strategy, and reported diagnostic performance. RESULTS: Among the 18 included studies, AI models demonstrated strong performance across diagnostic tasks in the BM-SRS pathway. The differentiation of RN from true tumor progression was the most extensively studied application, addressed by 14 of 18 studies, with reported AUCs ranging from 0.71 to 0.94. Support vector machines, random-forest ensembles, convolutional neural networks, and transformer-based multimodal architectures were widely used. The literature evolved from single-sequence radiomic classifiers in 2018 to multimodal deep learning frameworks fusing imaging with clinical and genomic data in 2025. Contrast-enhanced T1-weighted MRI was the dominant imaging input, and texture-based radiomic features (GLCM, GLSZM, GLDM, and wavelet-derived features) were the most consistently predictive. The highest-performing models reached AUCs of 0.85-0.91 through multimodal integration of imaging with clinical and genomic features, and consistently outperformed expert neuroradiologist read on matched cases. Remaining studies addressed longitudinal segmentation-based detection of local failure and adverse radiation effects, BRAF mutation status in melanoma BM, early Gamma Knife treatment response, and primary tumor histology classification, with more variable performance. CONCLUSION: AI models, particularly those integrating MRI-derived radiomic features with clinical and genomic data, show high accuracy in supporting diagnostic decisions for BM patients treated with SRS. The post-SRS differentiation of radiation necrosis from true tumor progression has reached the greatest level of maturity and is closest to clinical translation, with potential to reduce unnecessary biopsies, personalize surveillance intervals, and rationalize treatment-pathway decisions. Other diagnostic applications, including molecular subtyping and primary tumor histology classification, remain exploratory and require further multicenter validation. Integration of AI tools into multidisciplinary tumor-board workflows, combined with prospective validation and standardized reporting, will be essential to realize the full clinical benefits of AI in SRS for brain metastases.

Humans↗

Rapid diagnosis of common, undetected, and uncultivable bloodstream infections from positive blood cultures using Oxford Nanopore sequencing: a metagenomic pipeline analysis.

BACKGROUND: Metagenomic sequencing can potentially transform clinical microbiology by enabling rapid pathogen identification and antimicrobial resistance (AMR) prediction in critically ill patients with bloodstream infections. However, the clinical use of metagenomic sequencing has been constrained by its speed, accuracy, and technical feasibility. Our aim was to develop and evaluate a direct-from-positive blood culture workflow using Oxford Nanopore sequencing that overcomes these limitations and delivers rapid, accurate results. METHODS: In this metagenomic pipeline analysis, 211 positive (130 aerobic and 81 anaerobic) and 62 negative (30 aerobic and 32 anaerobic) randomly selected blood cultures were processed from Oxford University Hospitals for comparing species identification, AMR detection, and time-to-result against standard culture-based diagnostics performed by the hospital's routine microbiology laboratory. Species prediction was performed using Kraken2 with a comprehensive standard database, applying heuristic and random forest classification models. Additionally, we benchmarked AMR classification tools and databases, including ResFinder, CARD, and NCBI AMRFinderPlus. FINDINGS: Across all samples, our method achieved 97% sensitivity and 94% specificity for species identification compared with that of routine culture and matrix-assisted laser desorption ionisation time-of-flight-based diagnostics; both sensitivity and specificity increased to 100% after adjudication of plausible additional infections. We detected 19 additional infections (13 polymicrobial, five previously unidentifiable, and one in a culture-negative sample) and delivered species identification results within 3 h 20 min (IQR 3 h 7 min-3 h 27 min), approximately 10 h earlier than routine diagnostic methods. For the ten most common clinically relevant pathogens, our method yielded AMR results 20 h earlier than current antimicrobial susceptibility testing, with an overall sensitivity of 88% and specificity of 93%. Performance varied by species. For Staphylococcus aureus, the AMR prediction sensitivity was 100% and specificity was 99%, and for Escherichia coli, the prediction sensitivity was 91% and specificity was 94%. INTERPRETATION: These findings show that metagenomic sequencing has the potential to rapidly and comprehensively detect pathogens and AMR in bloodstream infections. Integration into clinical practice could help to close diagnostic gaps, reduce empirical antibiotic use, and enable rapid targeted treatment. Nonetheless, improvements in AMR prediction for some species and drugs, along with further multisite validation, are required before clinical implementation. FUNDING: National Institute for Health Research (NIHR) Oxford Biomedical Research Centre.

Humans↗

Direct and spillover hospitalisation patterns during climate hazards across regions of different health-system resilience levels in China: a nationwide retrospective analysis.

BACKGROUND: Health-system resilience serves as a key contributor in mitigating adverse health impacts during climate hazards. However, quantitative insights into resilience-associated health-care utilisation patterns and targeted adaptation policies remain scarce. We aimed to capture the spatiotemporal health impacts in disaster-exposed counties and their neighbouring counties in China during storms, floods, tropical cyclones, and blizzards or winter storms; understand the association between health-system resilience metrics and hazard-attributable hospitalisations; and develop evidence-based adaptation policies towards climate extremes. METHODS: In this retrospective, observational analysis of county-level aggregated hospitalisation data, we used a propensity score matching-difference-in-differences framework to assess the spatiotemporal changes of nine types of disease-specific hospitalisations in both disaster-exposed and neighbouring regions during storms, floods, tropical cyclones, and blizzards in China. We quantified the relative importance and health gains of health-system metrics during such hazards through random forest approach with interpretable partial dependence plots to derive evidence-based adaptation recommendations. FINDINGS: We included hospitalisation data from Jan 1, 2016 to Dec 31, 2023. In this period, 3241 county-hazard event combinations and 41 747 482 hospitalisations were recorded across 955 Chinese counties. The disaster-exposed regions experienced an initial decline in hospitalisation rates, followed by admission surges after disasters. For example, infectious disease admissions decreased by 11·92% (95% CI -10·53 to -13·31) during the flood-active period but increased by 7·68% (6·46-8·91) after 1-2 weeks of floods. Neighbouring zones were also affected through spillover effects, with infectious disease admissions increasing by 3·18% (1·76-4·61) after 1-2 weeks of the floods. Cardiovascular disease, injuries, infectious, respiratory, and mental disorders were more sensitive across all regions. Particularly for disaster-exposed counties, cardiovascular hospitalisations increased by 14·31% (7·34-21·29) during the tropical cyclone-active period. Notably, compared with low-resilience counties, high-resilience counties were associated with 19·48-30·03% smaller hazard-related relative changes in hospitalisation rates during the hazard-active period and 27·07-31·08% smaller hazard-related relative changes in hospitalisation rates in post-hazard periods. For instance, during the storm-active period, the increase in respiratory hospitalisations was 7·21% (0·67-13·75) in high-resilience counties versus 12·13% (5·20-19·05) in low-resilience counties. Health workforce (relative importance 14·58% during the hazard-active period and 13·80% during the post-hazard period) and service delivery (14·10% during the hazard-active period and 14·17% during the post-hazard period) were identified as key contributors of health-system resilience. Empirical synergistic effects were observed when combining interventions during the post-hazard period, with the combined effect of service delivery (individual contribution 8%) and workforce (individual contribution 4%) exceeding the sum of their individual contributions (16% reduction in cumulative excess admissions) by 33%. INTERPRETATION: Climate hazards are associated with substantial changes in hospitalisation rates in both disaster-exposed and neighbouring regions. Health-system resilience is essential in addressing disaster-health challenges. Targeted adaptation interventions should be context-appropriate and threshold-aware, thereby maximising the public health benefits relative to resilience-oriented investments in health systems. FUNDING: Gates Foundation and the National Natural Science Foundation of China.

Journal Article↗

Exploration of methods to identify polymorphisms associated with variation in DNA repair capacity phenotypes.

Elucidating the relationship between polymorphic sequences and risk of common disease is a challenge. For example, although it is clear that variation in DNA repair genes is associated with familial cancer, aging and neurological disease, progress toward identifying polymorphisms associated with elevated risk of sporadic disease has been slow. This is partly due to the complexity of the genetic variation, the existence of large numbers of mostly low frequency variants and the contribution of many genes to variation in susceptibility. There has been limited development of methods to find associations between genotypes having many polymorphisms and pathway function or health outcome. We have explored several statistical methods for identifying polymorphisms associated with variation in DNA repair phenotypes. The model system used was 80 cell lines that had been resequenced to identify variation; 191 single nucleotide substitution polymorphisms (SNPs) are included, of which 172 are in 31 base excision repair pathway genes, 19 in 5 anti-oxidation genes, and DNA repair phenotypes based on single strand breaks measured by the alkaline Comet assay. Univariate analyses were of limited value in identifying SNPs associated with phenotype variation. Of the multivariable model selection methods tested: the easiest that provided reduced error of prediction of phenotype was simple counting of the variant alleles predicted to encode proteins with reduced activity, which led to a genotype including 52 SNPs; the best and most parsimonious model was achieved using a two-step analysis without regard to potential functional relevance: first SNPs were ranked by importance determined by random forests regression (RFR), followed by cross-validation in a second round of RFR modeling that included ever more SNPs in declining order of importance. With this approach six SNPs were found to minimize prediction error. The results should encourage research into utilization of multivariate analytical methods for epidemiological studies of the association of genetic variation in complex genotypes with risk of common diseases.

Cell Line↗

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 688 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

DNA methylation↗

Meta-analysis of growth and inactivation kinetics of Legionella.

Quantitative risk assessments intended to inform evidence-based water management plans and public health targets for Legionella in engineered water systems are constrained by fragmented and heterogeneous growth and inactivation kinetics. We conducted a meta-analysis of 25 growth and 39 thermal- and chemical-inactivation studies, fitting microbial persistence models to harmonize parameters. Nonlinear models outperformed first-order formulations, indicating that lag phases and resistant or protected subpopulations are central to Legionella persistence. Random forest analysis identified environmental and methodological drivers of variability based on 226 growth rates and reduction times for thermal (209) and chemical (135) inactivation. Growth was primarily governed by temperature, nutrient availability, and compatible Legionella-host pairings; thermal inactivation by quantification method, temperature, and turbidity; and chemical inactivation by inoculum size, disinfectant type, concentration, and host-associations. Accordingly, temperature-dependent growth parameters and exposure metrics for heat, free-chlorine, and monochloramine, expressed as TT (Temperature×time) and CT (Concentration×time), were derived as condition-specific inputs for predictive models. Growth optima around 37-40 °C, together with lag-time estimates, indicate that hot-water temperature setbacks and energy-saving practices may favor Legionella proliferation under repeated or prolonged lukewarm exposure. Culture- and viability-based TT differences highlight the need to consider viable‑but-non-culturable persistence in monitoring programs. CT comparisons suggest monochloramine may be advantageous because of its lower apparent sensitivity to host-associated protection. Although limited by restricted experimental conditions, the findings show that predictive models should account for microbial ecology, water matrix effects, and quantification endpoints. Future kinetic studies should prioritize realistic multi-host systems, strain pre-adaptation, complementary viability measurements, and standardized protocols and reporting to ensure reproducibility and enable robust system-level predictive modeling.

Legionella↗

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)─a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC ≥ 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median ρ ∼ 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals↗

Attribution of PM2.5-Induced Transcriptomic Perturbation to Toxic Components.

Ambient fine particulate matter (PM2.5) is a chemically complex mixture whose health impacts are not fully captured by particle mass. Here, we developed an interpretable chemotranscriptomic framework to attribute PM2.5-induced molecular perturbations to toxicity-relevant components. PM2.5 collected from urban roadside and coastal environments was separated into whole, extractable, and unextractable fractions, characterized by LC/GC × GC-HRMS-based nontarget analysis and inductively coupled plasma mass spectrometry (ICP-MS), and evaluated using cytotoxicity testing and transcriptomic profiling in human bronchial epithelial cells. Urban PM2.5 exhibited greater cytotoxic potency per unit mass than coastal PM2.5, with extractable fractions accounting for most cytotoxic and pathway-level responses. Transcriptomics revealed distinct site-specific modes of action: urban PM2.5 preferentially induced oxidative stress, xenobiotic metabolism, and cell cycle suppression, consistent with acute, nonapoptotic injury, whereas coastal PM2.5 elicited weaker cytotoxicity but stronger interferon-mediated immune and apoptosis-related signaling. Integrating chemical abundance with pathway activity using random forest regression, SHAP interpretation, and mechanistic corroboration reduced 5,033 detected features to 444 pathway-linked candidate drivers. Fewer than 5% of features explained ∼95% of cumulative model contribution. Standard-confirmed contributors included plasticizer-related compounds, aromatic and heteroaromatic combustion products, and copper for urban PM2.5 and secondary/aged organics and nickel for coastal PM2.5. These findings support mechanism-informed prioritization of hazardous PM2.5 components beyond mass-based assessment.

Particulate Matter↗

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, 〈MAE〉 = 0.11 eV, and 〈RMSE〉 = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by ∼10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature↗

Boosting: an ensemble learning tool for compound classification and QSAR modeling.

A classification and regression tool, J. H. Friedman's Stochastic Gradient Boosting (SGB), is applied to predicting a compound's quantitative or categorical biological activity based on a quantitative description of the compound's molecular structure. Stochastic Gradient Boosting is a procedure for building a sequence of models, for instance regression trees (as in this paper), whose outputs are combined to form a predicted quantity, either an estimate of the biological activity, or a class label to which a molecule belongs. In particular, the SGB procedure builds a model in a stage-wise manner by fitting each tree to the gradient of a loss function: e.g., squared error for regression and binomial log-likelihood for classification. The values of the gradient are computed for each sample in the training set, but only a random sample of these gradients is used at each stage. (Friedman showed that the well-known boosting algorithm, AdaBoost of Freund and Schapire, could be considered as a particular case of SGB.) The SGB method is used to analyze 10 cheminformatics data sets, most of which are publicly available. The results show that SGB's performance is comparable to that of Random Forest, another ensemble learning method, and are generally competitive with or superior to those of other QSAR methods. The use of SGB's variable importance with partial dependence plots for model interpretation is also illustrated.

ATP Binding Cassette Transporter, Subfamily B, Mem↗

Structure-based classification of chemical reactions without assignment of reaction centers.

The automatic classification of chemical reactions is of high importance for the analysis of reaction databases, reaction retrieval, reaction prediction, or synthesis planning. In this work, the classification of photochemical reactions was investigated with no explicit assignment of the reacting centers. Classifications were explored with Random Forests or Kohonen neural networks in three different situations, using different levels of information: (a) pairs of reactants were classified according to the type of reaction they produce, (b) products were classified according to the type of reaction from which they can be synthesized, and (c) reactions were classified from the difference between the descriptors of the product and the descriptors of the reactants. In all cases molecular maps of atom-level properties (MOLMAPs) were used as descriptors. They are generated by a self-organizing map and encode physicochemical properties of the bonds available in a molecule. Correct classification could be achieved for approximately 90% of the 78 reactions in an independent test set.

Journal Article↗

Development and evaluation of an in silico model for hERG binding.

It has been recognized that drug-induced QT prolongation is related to blockage of the human ether-a-go-go-related gene (hERG) ion channel. Therefore, it is prudent to evaluate the hERG binding of active compounds in early stages of drug discovery. In silico approaches provide an economic and quick method to screen for potential hERG liability. A diverse set of 90 compounds with hERG IC(50) inhibition data was collected from literature references. Fragment-based QSAR descriptors and three different statistical methods, support vector regression, partial least squares, and random forests, were employed to construct QSAR models for hERG binding affinity. Important fragment descriptors relevant to hERG binding affinity were identified through an efficient feature selection method based on sparse linear support vector regression. The support vector regression predictive model built upon selected fragment descriptors outperforms the other two statistical methods in this study, resulting in an r(2) of 0.912 and 0.848 for the training and testing data sets, respectively. The support vector regression model was applied to predict hERG binding affinities of 20 in-house compounds belonging to three different series. The model predicted the relative binding affinity well for two out of three compound series. The hierarchical clustering and dendrogram results show that the compound series with the best prediction has much higher structural similarity and more neighbors of training compounds than the other two compound series, demonstrating the predictive scope of the model. The combination of a QSAR model and postprocessing analysis, such as clustering and visualization, provides a way to assess the confidence level of QSAR prediction results on the basis of similarity to the training set.

Cell Line↗

Assessing different classification methods for virtual screening.

How well do different classification methods perform in selecting the ligands of a protein target out of large compound collections not used to train the model? Support vector machines, random forest, artificial neural networks, k-nearest-neighbor classification with genetic-algorithm-optimized feature selection, trend vectors, naïve Bayesian classification, and decision tree were used to divide databases into molecules predicted to be active and those predicted to be inactive. Training and predicted activities were treated as binary. The database was generated for the ligands of five different biological targets which have been the object of intense drug discovery efforts: HIV-reverse transcriptase, COX2, dihydrofolate reductase, estrogen receptor, and thrombin. We report significant differences in the performance of the methods independent of the biological target and compound class. Different methods can have different applications; some provide particularly high enrichment, others are strong in retrieving the maximum number of actives. We also show that these methods do surprisingly well in predicting recently published ligands of a target on the basis of initial leads and that a combination of the results of different methods in certain cases can improve results compared to the most consistent method.

Algorithms↗

Physicochemical stereodescriptors of atomic chiral centers.

Physicochemical atomic stereodescriptors (PAS) were implemented that represent the chirality of an atomic chiral center on the basis of empirical physicochemical properties of the ligands. The ligands are ranked according to a specific property, and the chiral center takes an S/R-like descriptor relative to that property. The procedure is performed for a series of properties, yielding a chirality profile. Application of the PAS descriptors to the prediction of enantioselectivity in chemical reactions, from the molecular structures, is illustrated here. The relationship between the molecular structures, represented by the PAS descriptors, and the enantioselectivity was learned by neural networks, decision trees, or random forests. In a first application, a data set was employed with chiral amino alcohols that enantioselectively catalyze the addition of diethylzinc to benzaldehyde. Prediction of the major enantiomer obtained in the reaction, from the molecular structure of the catalyst, was achieved with accuracy up to 90%. The second application investigated the enantiopreference of Pseudomonas cepacia lipase (PCL) toward primary alcohols. The learned models could make correct predictions about the preferred enantiomer, from the molecular structure of the substrate, in up to 93% of the cases. These included substrates with and without O-atoms bonded to the chiral center. The properties automatically selected to build the models can give indications on the relevant factors guiding the observed chemical behavior.

Alcohols↗

PostDOCK: a structural, empirical approach to scoring protein ligand complexes.

In this work we introduce a postprocessing filter (PostDOCK) that distinguishes true binding ligand-protein complexes from docking artifacts (that are created by DOCK 4.0.1). PostDOCK is a pattern recognition system that relies on (1) a database of complexes, (2) biochemical descriptors of those complexes, and (3) machine learning tools. We use the protein databank (PDB) as the structural database of complexes and create diverse training and validation sets from it based on the "families of structurally similar proteins" (FSSP) hierarchy. For the biochemical descriptors, we consider terms from the DOCK score, empirical scoring, and buried solvent accessible surface area. For the machine-learners, we use a random forest classifier and logistic regression. Our results were obtained on a test set of 44 structurally diverse protein targets. Our highest performing descriptor combinations obtained approximately 19-fold enrichment (39 of 44 binding complexes were correctly identified, while only allowing 2 of 44 decoy complexes), and our best overall accuracy was 92%.

Ligands↗

Th2 bias and T-cell exhaustion characterize the immunopathology of non-tuberculous mycobacterial pulmonary disease.

Non-tuberculous mycobacterial pulmonary disease (NTM-PD) is an escalating global health concern with poorly defined immunological mechanisms, necessitating comprehensive profiling to guide therapeutic advances. We analyzed peripheral blood from 28 treatment-naïve NTM-PD patients (19 Mycobacterium avium complex, 9 Mycobacterium abscessus) and 27 matched controls using 42-marker mass cytometry (CyTOF) and Luminex multiplex assays. A random forest model identified predictive markers, while an in vitro murine macrophage model evaluated chemokine production. NTM-PD patients displayed significant immune shifts, including increased classical monocytes (CD14+ CD16-), reduced NKT-like cells (CD3+ CD56+), and elevated T-cell exhaustion markers (PD-1, TOX). This coincided with a Th1/Th2 balance shift characterized by heightened IL-13. Elevated IFN-γ-inducible chemokines CXCL9 and CXCL10 coexisted with this Th2-biased signature, indicating a complex, dysregulated inflammatory state. A model integrating immune-cell frequencies and cytokine profiles achieved robust diagnostic accuracy (AUC = 0.922) with prognostic potential. In vitro, NTM-infected macrophages produced substantial CXCL9 and CXCL10 levels relative to the LPS maximal activation benchmark, identifying them as a major cellular source. These findings propose an immunological framework wherein T-cell exhaustion and a Th2-biased microenvironment strongly correlate with NTM-PD pathogenesis. CXCL9, CXCL10, and IL-13 emerge as candidate therapeutic targets, while our predictive model offers a foundational approach for risk stratification.

Humans↗

Mapping key mitochondrial genes in Alzheimer's disease through human tissue and iPSC derived neurons.

Alzheimer's disease (AD) is a progressive neurodegenerative condition that has become a global health challenge due to an aging world population and no available effective treatment. Mitochondrial dysfunction plays a crucial role in the development of AD due to its critical role in neuronal survival and function. However, the specific mitochondrial genes and pathways involved in AD pathogenesis remain poorly defined. In this study, we incorporated seven AD human postmortem and three AD iPSC-derived neurons (iNs) gene expression datasets to identify mitochondria-related Differentially Expressed Genes (mitoDEGs) between AD and control. The Gene Ontology (GO) analysis is conducted to investigate the AD biological mechanisms, and a random forest model is developed to assess how well the key mitoDEGs differentiate AD and control groups. Through our analysis, we identified fourteen key mitochondria related genes that show significant dysregulation in both postmortem brain tissues and iNs derived from AD patients. These genes have strong connections to oxidative stress, indicating mitochondrial dysfunction plays a crucial role in Alzheimer's disease pathology. Our study identified the key genes and pathways as promising targets for future research and therapeutic interventions, highlighting the importance of mitigating oxidative stress and restoring mitochondrial function in AD.

Humans↗

Predicting interpretability of metabolome models based on behavior, putative identity, and biological relevance of explanatory signals.

Powerful algorithms are required to deal with the dimensionality of metabolomics data. Although many achieve high classification accuracy, the models they generate have limited value unless it can be demonstrated that they are reproducible and statistically relevant to the biological problem under investigation. Random forest (RF) generates models, without any requirement for dimensionality reduction or feature selection, in which individual variables are ranked for significance and displayed in an explicit manner. In metabolome fingerprinting by mass spectrometry, each metabolite can be represented by signals at several m/z. Exploiting a prior understanding of expected biochemical differences between sample classes, we aimed to develop meaningful metrics relevant to the significance both of the overall RF model and individual, potentially explanatory, signals. Pair-wise comparison of related plant genotypes with strong phenotypic differences demonstrated that robust models are not only reproducible but also logically structured, highlighting correlated m/z derived from just a small number of explanatory metabolites reflecting the biological differences between sample classes. RF models were also generated by using groupings of samples known to be increasingly phenotypically similar. Although classification accuracy was often reasonable, we demonstrated reproducibly in both Arabidopsis and potato a performance threshold based on margin statistics beyond which such models showed little structure indicative of either generalizability or further biological interpretability. In a multiclass problem using 25 Arabidopsis genotypes, despite the complicating effects of ecotype background and secondary metabolome perturbations common to several mutations, the ranking of metabolome signals by RF provided scope for deeper interpretability.

Arabidopsis↗