PubMed HealthSearch

SEARCH · PubMed Health

Results for “Random Forest”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

In silico screening of anti-atherosclerotic compounds from Morus alba leaves by machine learning and network pharmacology.

OBJECTIVE: This study integrates machine learning with network pharmacology, molecular docking, and molecular dynamics simulations to screen bioactive compounds from Mulberry leaves and elucidate their potential mechanisms against atherosclerosis (AS). METHODS: A training dataset of anti-AS active compounds was compiled and encoded as Morgan fingerprints. Three machine learning classifiers, specifically Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XG-Boost), were constructed and evaluated using multiple performance metrics. Potential active components from Mulberry leaves and AS-related targets were retrieved, followed by protein-protein interaction network construction and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis. Molecular docking was then performed to evaluate binding affinities between core targets and candidate compounds, and the most stable complex was subjected to molecular dynamics simulations using GROMACS (2025). RESULTS: The RF model achieved superior performance (accuracy= 0.8354, F1 = 0.8408, AUC = 0.9119) with 100% external validation accuracy. Thirteen anti-AS candidates were prioritized from mulberry leaves, four of which have been previously documented. Network pharmacology revealed AKT1 and IL6 as core targets, enriched in pathways such as endocrine resistance. Molecular docking and dynamics simulations confirmed strong binding between oxysanguinarine and AKT1, with the complex exhibiting high stability. CONCLUSION: The RF model provides a reliable computational tool for prioritizing anti-AS compounds from Mulberry leaves. The integrated analysis reveals that Mulberry leaves exert anti-atherosclerotic effects through multi-target (e.g., AKT1, IL6) and multi-pathway (e.g., PI3K-Akt) mechanisms, offering a framework for further experimental validation.

Morus

Uncovering hub genes and key pathways responsive to drought stress in rice via meta-analysis of transcriptomic data.

Drought stress presents a formidable threat to global rice cultivation, triggering complex molecular responses that impact plant growth and productivity. To decipher the underlying gene expression dynamics, we performed a comprehensive meta-analysis of transcriptomic datasets derived from drought-tolerant rice genotypes. Via microarray data from three independent studies, we identified a set of consistently expressed differentially expressed genes (DEGs) under drought conditions. Integration of functional annotation tools, including GO and KEGG pathway enrichment, revealed key biological processes and signaling cascades involved in stress mitigation, such as ABA signaling, protein folding, and photosynthesis suppression. Protein-protein interaction (PPI) network construction, followed by hub gene identification via maximal clique centrality (MCC), highlighted pivotal regulators including LEA proteins, dehydrins, HSP70, and several transcription factors. Machine learning approaches further prioritize potential biomarkers, with Random Forest models achieving high classification accuracy and pinpointing key predictive genes. Chromosomal localization analysis provided spatial insights into the distribution of these hub genes, whose expression patterns were further compared against qRT-PCR data from previously published studies. This integrative approach identifies candidate genomic markers and mechanistic insights that may support future breeding strategies for drought-tolerant rice, pending experimental validation.

Cytoscape

Liquid Biopsy-Multiomics Link Adhesion Pathway Dysregulation to Kidney Injury Severity.

INTRODUCTION: Severe acute kidney injury (AKI) is strongly associated with the risk of developing chronic kidney disease; however, little is known about the cell type-specific mechanisms driving kidney injury severity. METHODS: In this multicenter observational study, we used clinically obtained liquid biopsy proteomics and machine learning (ML) to predict severe outcomes in patients with COVID-associated and non-COVID AKI. Further, we orthogonally combined 169 urine proteomics with 437 plasma proteomics samples and 40 urine sediment single-cell transcriptomics samples to identify complementary dysregulated mechanisms. RESULTS: Using a 10-fold cross-validated random forest algorithm, we identified a set of urinary proteins that demonstrate predictive power for both discovery and validation set with AUC of 87% and 76%, respectively. These predictive proteomics features obtained demonstrate that cell adhesion and autophagy-associated pathways are uniquely impacted in severe AKI. Differentially abundant proteins (DAPSs) associated with these pathways are highly expressed in cells of the juxtamedullary nephron, endothelial cells (ECs), and podocytes, indicating that these kidney cell types could be potential targets. Single-cell transcriptomic analysis in the in vitro model of kidney organoids infected with SARS-CoV-2 reveal dysregulation of extracellular matrix (ECM) organization in multiple nephron segments, recapitulating the clinically observed fibrotic response across multiomics datasets. Ligand-receptor interaction analysis of the podocyte and tubule organoid clusters shows significant reduction and loss of interaction between integrins and basement membrane receptors in the infected kidney organoids. CONCLUSION: Collectively, these data suggest that ECM degradation and adhesion-associated mechanisms could be the main driver of severe kidney injury.

AKI

Metagenomic insights into ecological risk of antibiotic resistome and mobilome in riverine plastisphere under impact of urbanization.

Microplastics (MPs) are of increasing concern due to their role as reservoirs for antibiotic resistance genes (ARGs) and pathogens. To date, few studies have explored the influence of anthropogenic activities on ARGs and mobile genetic elements (MGEs) within various riverine MPs, in comparison to their natural counterparts. Here an in-situ incubation was conducted along heavily anthropogenically-impacted Houxi River to characterize the geographical pattern of antibiotic resistome, mobilome and pathogens inhabiting MPs- and leaf-biofilms. The metagenomics result showed a clear urbanization-driven profile in the distribution of ARGs, MGEs and pathogens, with their abundances sharply increasing 4.77 to 19.90 times from sparsely to densely populated regions. The significant correlation between human fecal marker crAssphage and ARG (R2 = 0.67, P=0.003) indicated the influence of anthropogenic activity on ARG proliferation in plastisphere and natural leaf surfaces. And mantel tests and random forest analysis revealed the impact of 17 socio-environmental factors, e.g., population density, antibiotic concentrations, and pore volume of materials, on the dissemination of ARGs. Partial least squares-path modeling further unveiled that intensifying human activities not only directly boosted ARGs abundance but also exerted a comparable indirect impact on ARGs propagation. Furthermore, the polyvinylchloride plastisphere created a pathogen-friendly habitat, harboring higher abundances of ARGs and MGEs, while polylactic acid are not likely to serve as vectors for pathogens in river, with a lower resistome risk score than that in leaf-biofilms. This study highlights the diverse ecological risks associated with the dissemination of ARGs and pathogens in varied MPs, offering insights for the policymaking of usage and control of plastics within urbanization.

Urbanization

Temporal shifts in gyrA mutation types and sublineage replacement in ST11 Salmonella enterica serovar Enteritidis over a decade (2014-2023): A genomic epidemiological study in Guangxi, China.

The overuse or abuse of antibiotics drives the global health threat of antimicrobial resistance. Although bans on certain veterinary antibiotics, such as colistin, have proven effective, the impact of fluoroquinolone stewardship on the evolution of the foodborne pathogen Salmonella enterica serovar Enteritidis (S. Enteritidis) remains unclear. Here, we conducted a decade-long (2014-2023) retrospective longitudinal genomic epidemiological analysis of 441 ST11 S. Enteritidis isolates from Guangxi, China, alongside a global reference dataset of 4297 genomes. Our aim was to elucidate the effect of real-world antibiotic stewardship on the shift of gyrA point mutations and lineage distribution. Surveillance identified three global epidemic clade sublineages (GEC-L2, L3, L4), with the multidrug-resistant GEC-L4 (i.e., GC-c or MMC2), characterized by the gyrA mutation with amino acid substitution D87Y, being domestically dominant (70.07%, 309/441). Following China's 2016 ban on the veterinary use of critical fluoroquinolones, the proportion of the highly resistant GEC-L4 sublineage decreased continuously (from 86.84% in 2017 to 56.00% in 2023), while the less resistant GEC-L3 sublineage (i.e., GC-b or MMC1), mainly characterized by gyrA D87G, increased simultaneously (from 13.16% to 44.00%). This phenomenon might be attributed to the fact that the GEC-L4 sublineage exhibited a higher fitness cost compared with the GEC-L3 sublineage, as confirmed by the competition assay. A Random Forest Model validated that the gyrA mutation with amino acid substitution D87Y was the paramount feature for these sublineages' identification. In contrast, global data showed a continuous increase in gyrA mutations (from 8.63% in 2006 to 68.85% in 2024), primarily D87Y (from 1.44% to 31.15%) and D87N (from 4.32% to 22.95%), correlating with rising average fluoroquinolone consumption. This study provides direct genomic evidence that national-level antibiotic stewardship can drive the replacement of highly resistant sublineages with moderately resistant ones. These findings offer crucial scientific evidence for evaluating the impact of antibiotic management policies and inform strategies for the rational use of antimicrobials.

China

Evaluating the impact of modeling choices on the performance of integrated genetic and clinical models.

PURPOSE: The value of genetic information for improving the performance of clinical risk prediction models has yielded variable conclusions. Many methodological decisions have the potential to contribute to differential results. We performed multiple modeling experiments integrating clinical and demographic data from electronic health records with genetic data to understand which decisions may affect performance. METHODS: Clinical data in the form of structured diagnostic codes, medications, procedural codes, and demographics were extracted from 2 large independent health systems, and polygenic risk scores (PRS) were generated across all patients of European ancestry with genetic data in the corresponding biobanks. Crohn's disease was studied based on its substantial genetic component, established electronic health records-based definition, and sufficient prevalence for training and testing. We investigated the impact of choices regarding the PRS integration method, training sample, model complexity, and performance metrics. RESULTS: Overall, our results showed that including PRS resulted in higher performance, but this gain was only robust in situations with limited clinical information. We found consistent performance increases from more compute-intensive models, such as random forest, but the impact of other decisions varied by site. CONCLUSION: This work highlights the importance of considering methodological decision points in interpreting the impact of PRS on prediction performance in clinical models.

Humans

Development and validation of a machine learning prognostic model based on an epigenomic signature in patients with pancreatic ductal adenocarcinoma.

BACKGROUND: In Pancreatic Ductal Adenocarcinoma (PDAC), current prognostic scores are unable to fully capture the biological heterogeneity of the disease. While some approaches investigating the role of multi-omics in PDAC are emerging, the analysis of methylation data is under exploited. MATERIALS AND METHODS: We analyzed CpG sites from two publicly available datasets, the TCGA-PAAD used as discovery set and the CPTAC-PDA as external test set. Single mutations and co-mutation of KRAS and TP53 genes were identified as targets, and differentially methylated CpG sites (DMC) were detected accordingly. We trained and validated Random Forest (RF) models to predict each target. Area Under the Receiver Operating Characteristic curve (AUROC) and Area Under the Precision-Recall curve (AUPRC) were used as performance metrics. Then, we performed consensus clustering from the DMCs to identify novel patients' profiles. Finally, we trained and validated a combination of eXtreme Gradient Boosting (XGB) and tree models to select an epigenomic prognostic determinant. RESULTS: From 598 DMCs extracted, an RF model predicted KRAS and TP53 co-mutation on the external test set with AUROC of 0.77 and AUPRC of 0.87. The consensus clustering allowed us to identify 4 clusters (C1, C2, C3, and C4) of patients. The C4 cluster captured a subgroup of patients with favorable Overall Survival (OS) with respect to others. The XGB model perfectly predicted C4 vs other clusters on the discovery set. In both cohorts, patients were stratified into two risk groups according to methylation levels of cg16854533, individuated as the most important CpG site. CONCLUSION: We analyzed methylation data to develop a classifier for the TP53 and KRAS mutational status. Four prognostic clusters were pointed out and a prognostic model using a CpG site was validated in an independent cohort. Our results evidence that the proposed use of methylation data facilitates risk stratification for PDAC.

Humans

Machine learning-assisted plasma PEA proteomics enables differential diagnosis of melancholic depression and bipolar disorder.

Differentiating bipolar disorder (BD) from major depressive disorder (MDD) remains a critical unmet need in psychiatry due to overlapping clinical presentations and the absence of reliable biological markers. In this study, we assessed the capacity of multivariate machine learning models to accurately differentiate BD from MDD with melancholic features using plasma proteomic profiles obtained via Proximity Extension Assay (PEA) technology. A total of 67 participants were included (23 BD, 20 MDD, and 24 HC), and plasma protein expression was assessed using the Olink Target 96 Neurology panel. Differential proteomic analysis revealed distinct disorder-specific expression patterns, identifying 21 differentially expressed proteins in BD versus MDD, 18 in BD versus healthy controls, and 7 in MDD versus healthy controls. Using a stepwise feature reduction strategy, machine learning models were trained on three feature sets comprising all proteins, the top 20 most informative proteins, and the top 5 most beneficial proteins, and evaluated across BD-MDD, BD-HC, and MDD-HC classification tasks using five algorithms. For BD-MDD discrimination, the Random Forest model achieved the highest performance when trained on the top 5 protein set (LXN, HAGH, MATN3, PLXNB1, and CTSC), yielding an AUC of 0.905, with similarly strong performance observed using the top 20 protein set. Feature importance analysis highlighted proteins involved in neurodevelopmental processes, immune regulation, and extracellular matrix organization. Overall, these findings demonstrate that integrating plasma proteomics with machine learning enables robust differentiation between BD and MDD with melancholic features, supporting the development of scalable and biologically informed diagnostic tools for precision psychiatry.

Bipolar disorder

Integrated Multi-Omics Analyses Reveal Lipid Metabolic Signature in Osteoarthritis.

Osteoarthritis (OA) is the most common degenerative joint disease and the second leading cause of disability worldwide. Single-omics analyses are far from elucidating the complex mechanisms of lipid metabolic dysfunction in OA. This study identified a shared lipid metabolic signature of OA by integrating metabolomics, single-cell and bulk RNA-seq, as well as metagenomics. Compared to the normal counterparts, cartilagesin OA patients exhibited significant depletion of homeostatic chondrocytes (HomCs) (P&#xa0;=&#xa0;0.03) and showed lipid metabolic disorders in linoleic acid metabolism and glycerophospholipid metabolism which was consistent with our findings obtained from plasma metabolomics. Through high-dimensional weighted gene co-expression network analysis (hdWGCNA), weidentified PLA2G2A as a hub gene associated with lipid metabolic disorders in HomCs. And an OA-associated subtype of HomCs, namely HomC1 (marked by PLA2G2A, MT-CO1, MT-CO2, and MT-CO3) was identified, which also exhibited abnormal activation of lipid metabolic pathways. This suggests the involvement of HomC1 in OA progression through the shared lipid metabolism aberrancies, which were further validated via bulk RNA-Seq analysis. Metagenomic profiling identified specific gut microbial species significantly associated with the key lipid metabolism disorders, including Bacteroides uniformis (P&#xa0;<&#xa0;0.001, R&#xa0;=&#xa0;-0.52), Klebsiella pneumonia (P&#xa0;=&#xa0;0.003, R&#xa0;=&#xa0;0.42), Intestinibacter_bartlettii (P&#xa0;=&#xa0;0.009, R&#xa0;=&#xa0;0.38), and Streptococcus anginosus (P&#xa0;=&#xa0;0.009, R&#xa0;=&#xa0;0.38). By integrating the multi-omics features, a random forest diagnostic model with outstanding performance was developed (AUC&#xa0;=&#xa0;0.97). In summary, this study deciphered the crucial role of a integrated lipid metabolic signature in OA pathogenesis, and established a regulatory axis of gut microbiota-metabolites-cell-gene, providing new insights into the gut-joint axis and precision therapy for OA.

Humans

Artificial intelligence (AI) uses in stereotactic radiosurgery (SRS): diagnosis with brain metastasis (BM) - A systematic review.

BACKGROUND: Brain metastases (BM) are the most common intracranial tumors in adults, and stereotactic radiosurgery (SRS) has become a mainstay of management. However, several diagnostic challenges persist in the SRS pathway, particularly the differentiation of radiation necrosis (RN) from true tumor progression, which conventional MRI and even advanced imaging techniques often cannot reliably resolve. Recent advances in artificial intelligence (AI) offer the potential to address these diagnostic limitations. This systematic review synthesizes current literature on AI applications for MRI-based diagnostic decision support in BM patients undergoing SRS, with a focus on radiomics and deep learning tools for distinguishing RN from progression, classifying molecular and histologic subtypes, and predicting treatment response. METHODS: A systematic review was performed in accordance with PRISMA guidelines. PubMed, Web of Science, and Scopus were searched using a targeted query combining terms related to AI, brain metastasis, diagnosis or imaging, and SRS. After screening 483 records and applying strict inclusion and exclusion criteria, 18 studies published between 2015 and 2025 were included. Data were extracted on study design, cohort characteristics, imaging modality, AI methodology, validation strategy, and reported diagnostic performance. RESULTS: Among the 18 included studies, AI models demonstrated strong performance across diagnostic tasks in the BM-SRS pathway. The differentiation of RN from true tumor progression was the most extensively studied application, addressed by 14 of 18 studies, with reported AUCs ranging from 0.71 to 0.94. Support vector machines, random-forest ensembles, convolutional neural networks, and transformer-based multimodal architectures were widely used. The literature evolved from single-sequence radiomic classifiers in 2018 to multimodal deep learning frameworks fusing imaging with clinical and genomic data in 2025. Contrast-enhanced T1-weighted MRI was the dominant imaging input, and texture-based radiomic features (GLCM, GLSZM, GLDM, and wavelet-derived features) were the most consistently predictive. The highest-performing models reached AUCs of 0.85-0.91 through multimodal integration of imaging with clinical and genomic features, and consistently outperformed expert neuroradiologist read on matched cases. Remaining studies addressed longitudinal segmentation-based detection of local failure and adverse radiation effects, BRAF mutation status in melanoma BM, early Gamma Knife treatment response, and primary tumor histology classification, with more variable performance. CONCLUSION: AI models, particularly those integrating MRI-derived radiomic features with clinical and genomic data, show high accuracy in supporting diagnostic decisions for BM patients treated with SRS. The post-SRS differentiation of radiation necrosis from true tumor progression has reached the greatest level of maturity and is closest to clinical translation, with potential to reduce unnecessary biopsies, personalize surveillance intervals, and rationalize treatment-pathway decisions. Other diagnostic applications, including molecular subtyping and primary tumor histology classification, remain exploratory and require further multicenter validation. Integration of AI tools into multidisciplinary tumor-board workflows, combined with prospective validation and standardized reporting, will be essential to realize the full clinical benefits of AI in SRS for brain metastases.

Humans

Rapid diagnosis of common, undetected, and uncultivable bloodstream infections from positive blood cultures using Oxford Nanopore sequencing: a metagenomic pipeline analysis.

BACKGROUND: Metagenomic sequencing can potentially transform clinical microbiology by enabling rapid pathogen identification and antimicrobial resistance (AMR) prediction in critically ill patients with bloodstream infections. However, the clinical use of metagenomic sequencing has been constrained by its speed, accuracy, and technical feasibility. Our aim was to develop and evaluate a direct-from-positive blood culture workflow using Oxford Nanopore sequencing that overcomes these limitations and delivers rapid, accurate results. METHODS: In this metagenomic pipeline analysis, 211 positive (130 aerobic and 81 anaerobic) and 62 negative (30 aerobic and 32 anaerobic) randomly selected blood cultures were processed from Oxford University Hospitals for comparing species identification, AMR detection, and time-to-result against standard culture-based diagnostics performed by the hospital's routine microbiology laboratory. Species prediction was performed using Kraken2 with a comprehensive standard database, applying heuristic and random forest classification models. Additionally, we benchmarked AMR classification tools and databases, including ResFinder, CARD, and NCBI AMRFinderPlus. FINDINGS: Across all samples, our method achieved 97% sensitivity and 94% specificity for species identification compared with that of routine culture and matrix-assisted laser desorption ionisation time-of-flight-based diagnostics; both sensitivity and specificity increased to 100% after adjudication of plausible additional infections. We detected 19 additional infections (13 polymicrobial, five previously unidentifiable, and one in a culture-negative sample) and delivered species identification results within 3 h 20 min (IQR 3 h 7 min-3 h 27 min), approximately 10 h earlier than routine diagnostic methods. For the ten most common clinically relevant pathogens, our method yielded AMR results 20 h earlier than current antimicrobial susceptibility testing, with an overall sensitivity of 88% and specificity of 93%. Performance varied by species. For Staphylococcus aureus, the AMR prediction sensitivity was 100% and specificity was 99%, and for Escherichia coli, the prediction sensitivity was 91% and specificity was 94%. INTERPRETATION: These findings show that metagenomic sequencing has the potential to rapidly and comprehensively detect pathogens and AMR in bloodstream infections. Integration into clinical practice could help to close diagnostic gaps, reduce empirical antibiotic use, and enable rapid targeted treatment. Nonetheless, improvements in AMR prediction for some species and drugs, along with further multisite validation, are required before clinical implementation. FUNDING: National Institute for Health Research (NIHR) Oxford Biomedical Research Centre.

Humans

Direct and spillover hospitalisation patterns during climate hazards across regions of different health-system resilience levels in China: a nationwide retrospective analysis.

BACKGROUND: Health-system resilience serves as a key contributor in mitigating adverse health impacts during climate hazards. However, quantitative insights into resilience-associated health-care utilisation patterns and targeted adaptation policies remain scarce. We aimed to capture the spatiotemporal health impacts in disaster-exposed counties and their neighbouring counties in China during storms, floods, tropical cyclones, and blizzards or winter storms; understand the association between health-system resilience metrics and hazard-attributable hospitalisations; and develop evidence-based adaptation policies towards climate extremes. METHODS: In this retrospective, observational analysis of county-level aggregated hospitalisation data, we used a propensity score matching-difference-in-differences framework to assess the spatiotemporal changes of nine types of disease-specific hospitalisations in both disaster-exposed and neighbouring regions during storms, floods, tropical cyclones, and blizzards in China. We quantified the relative importance and health gains of health-system metrics during such hazards through random forest approach with interpretable partial dependence plots to derive evidence-based adaptation recommendations. FINDINGS: We included hospitalisation data from Jan 1, 2016 to Dec 31, 2023. In this period, 3241 county-hazard event combinations and 41&#x2009;747&#x2009;482 hospitalisations were recorded across 955 Chinese counties. The disaster-exposed regions experienced an initial decline in hospitalisation rates, followed by admission surges after disasters. For example, infectious disease admissions decreased by 11&#xb7;92% (95% CI -10&#xb7;53 to -13&#xb7;31) during the flood-active period but increased by 7&#xb7;68% (6&#xb7;46-8&#xb7;91) after 1-2 weeks of floods. Neighbouring zones were also affected through spillover effects, with infectious disease admissions increasing by 3&#xb7;18% (1&#xb7;76-4&#xb7;61) after 1-2 weeks of the floods. Cardiovascular disease, injuries, infectious, respiratory, and mental disorders were more sensitive across all regions. Particularly for disaster-exposed counties, cardiovascular hospitalisations increased by 14&#xb7;31% (7&#xb7;34-21&#xb7;29) during the tropical cyclone-active period. Notably, compared with low-resilience counties, high-resilience counties were associated with 19&#xb7;48-30&#xb7;03% smaller hazard-related relative changes in hospitalisation rates during the hazard-active period and 27&#xb7;07-31&#xb7;08% smaller hazard-related relative changes in hospitalisation rates in post-hazard periods. For instance, during the storm-active period, the increase in respiratory hospitalisations was 7&#xb7;21% (0&#xb7;67-13&#xb7;75) in high-resilience counties versus 12&#xb7;13% (5&#xb7;20-19&#xb7;05) in low-resilience counties. Health workforce (relative importance 14&#xb7;58% during the hazard-active period and 13&#xb7;80% during the post-hazard period) and service delivery (14&#xb7;10% during the hazard-active period and 14&#xb7;17% during the post-hazard period) were identified as key contributors of health-system resilience. Empirical synergistic effects were observed when combining interventions during the post-hazard period, with the combined effect of service delivery (individual contribution 8%) and workforce (individual contribution 4%) exceeding the sum of their individual contributions (16% reduction in cumulative excess admissions) by 33%. INTERPRETATION: Climate hazards are associated with substantial changes in hospitalisation rates in both disaster-exposed and neighbouring regions. Health-system resilience is essential in addressing disaster-health challenges. Targeted adaptation interventions should be context-appropriate and threshold-aware, thereby maximising the public health benefits relative to resilience-oriented investments in health systems. FUNDING: Gates Foundation and the National Natural Science Foundation of China.

Journal Article

Machine Learning-Based Identification of Survival-Associated CpG Biomarkers in Pancreatic Ductal Adenocarcinoma.

Pancreatic ductal adenocarcinoma (PDAC) is an exceptionally aggressive cancer with a 5-year survival rate of less than 10%, driven by late-stage diagnosis, limited treatment options, and a lack of reliable biomarkers for early detection and prognosis. In this study, we integrated DNA methylation data from TCGA and ICGC cohorts, categorizing samples based on survival time, and identified 688 differentially methylated CpG sites, along with 224 CpG biomarkers significantly associated with patient survival through statistical and machine learning-based analyses. We developed a random forest model to predict patient survival, achieving 85.2% accuracy for short-survival patients and 70.0% for long-survival patients in the validation set. External dataset validation further confirmed the model's robustness and accuracy. De novo motif analysis of genomic regions surrounding the 224 CpG biomarkers identified TWIST1 and FOXA2 as key transcriptional regulators enriched in survival-associated CpG sites, linking their activity to patient survival outcomes. Collectively, our findings highlight valuable epigenetic biomarkers and provide a predictive model to assess PDAC risk levels post-surgery, offering the potential for improved patient stratification and personalized therapeutic strategies.

DNA methylation

Meta-analysis of growth and inactivation kinetics of Legionella.

Quantitative risk assessments intended to inform evidence-based water management plans and public health targets for Legionella in engineered water systems are constrained by fragmented and heterogeneous growth and inactivation kinetics. We conducted a meta-analysis of 25 growth and 39 thermal- and chemical-inactivation studies, fitting microbial persistence models to harmonize parameters. Nonlinear models outperformed first-order formulations, indicating that lag phases and resistant or protected subpopulations are central to Legionella persistence. Random forest analysis identified environmental and methodological drivers of variability based on 226 growth rates and reduction times for thermal (209) and chemical (135) inactivation. Growth was primarily governed by temperature, nutrient availability, and compatible Legionella-host pairings; thermal inactivation by quantification method, temperature, and turbidity; and chemical inactivation by inoculum size, disinfectant type, concentration, and host-associations. Accordingly, temperature-dependent growth parameters and exposure metrics for heat, free-chlorine, and monochloramine, expressed as TT (Temperature&#xd7;time) and CT (Concentration&#xd7;time), were derived as condition-specific inputs for predictive models. Growth optima around 37-40 &#xb0;C, together with lag-time estimates, indicate that hot-water temperature setbacks and energy-saving practices may favor Legionella proliferation under repeated or prolonged lukewarm exposure. Culture- and viability-based TT differences highlight the need to consider viable&#x2011;but-non-culturable persistence in monitoring programs. CT comparisons suggest monochloramine may be advantageous because of its lower apparent sensitivity to host-associated protection. Although limited by restricted experimental conditions, the findings show that predictive models should account for microbial ecology, water matrix effects, and quantification endpoints. Future kinetic studies should prioritize realistic multi-host systems, strain pre-adaptation, complementary viability measurements, and standardized protocols and reporting to ensure reproducibility and enable robust system-level predictive modeling.

Legionella

Spectral-Proteomic Integration Analysis (SPIA) Deciphers Molecular Trajectories of Breast Cancer and Enables Multitarget Therapeutic Assessment.

Raman spectroscopy and mass spectrometry-based proteomics offer deeply complementary yet largely disconnected views of cancer biology: the former provides a label-free, real-time biochemical phenotype, while the latter delivers a quantitative inventory of specific protein effectors. Bridging this gap remains a fundamental challenge in analytical biomedicine. Here, we introduce Spectral-Proteomic Integration Analysis (SPIA)&#x2500;a novel, data-driven integrative framework that systematically links Raman spectroscopic phenotypes with quantitative proteomic profiles through machine learning and statistical correlation. Using a DMBA-induced rat breast cancer model with and without Toremifene (TOR) intervention, SPIA dynamically maps tumor microenvironment remodeling, capturing progressive collagen deposition and lipid metabolic reprogramming. An SVM classifier trained on Raman spectra achieves exceptional diagnostic accuracy (AUC &#x2265; 99.0%) and successfully predicts TOR therapeutic response. Proteomic analysis identifies 1,350 differentially expressed proteins, with convergent machine learning feature selection (LASSO, Random Forest, XGBoost) pinpointing core regulators including Luc7l2, Nucb1, Cbx3, and Csnk2a1. Crucially, Spearman correlation analysis between key Raman bands and core DEPs reveals strong, statistically robust associations (median &#x3c1; &#x223c; 0.75 in the 1533-1669 cm-1 region), empirically validating SPIA's core integrative logic. Leveraging this multimodal map, we elucidate a multitarget mechanism for TOR involving concurrent suppression of collagen deposition and correction of aberrant lipid metabolism. SPIA establishes a powerful, generalizable paradigm for integrating phenotypic and molecular data, with broad implications for biomarker discovery, drug mechanism elucidation, and precision oncology.

Animals

Attribution of PM2.5-Induced Transcriptomic Perturbation to Toxic Components.

Ambient fine particulate matter (PM2.5) is a chemically complex mixture whose health impacts are not fully captured by particle mass. Here, we developed an interpretable chemotranscriptomic framework to attribute PM2.5-induced molecular perturbations to toxicity-relevant components. PM2.5 collected from urban roadside and coastal environments was separated into whole, extractable, and unextractable fractions, characterized by LC/GC &#xd7; GC-HRMS-based nontarget analysis and inductively coupled plasma mass spectrometry (ICP-MS), and evaluated using cytotoxicity testing and transcriptomic profiling in human bronchial epithelial cells. Urban PM2.5 exhibited greater cytotoxic potency per unit mass than coastal PM2.5, with extractable fractions accounting for most cytotoxic and pathway-level responses. Transcriptomics revealed distinct site-specific modes of action: urban PM2.5 preferentially induced oxidative stress, xenobiotic metabolism, and cell cycle suppression, consistent with acute, nonapoptotic injury, whereas coastal PM2.5 elicited weaker cytotoxicity but stronger interferon-mediated immune and apoptosis-related signaling. Integrating chemical abundance with pathway activity using random forest regression, SHAP interpretation, and mechanistic corroboration reduced 5,033 detected features to 444 pathway-linked candidate drivers. Fewer than 5% of features explained &#x223c;95% of cumulative model contribution. Standard-confirmed contributors included plasticizer-related compounds, aromatic and heteroaromatic combustion products, and copper for urban PM2.5 and secondary/aged organics and nickel for coastal PM2.5. These findings support mechanism-informed prioritization of hazardous PM2.5 components beyond mass-based assessment.

Particulate Matter

Transfer Learning across Material Properties Using Center-Environment Features: From Energetics to Mechanical Properties in Multicomponent Mo Alloys.

Transfer learning (TL) provides a viable approach to mitigate data scarcity in materials informatics. While conventional TL focuses on predicting identical properties across different systems, this work demonstrates a cross-property extension of TL from energy to mechanical properties via end-to-end model weight pre-training and fine-tuning: knowledge learned from predicting substitution energies is transferred to predict distinctly different mechanical properties, substantially improving computational efficiency given the typically higher cost of acquiring target-domain data. To accelerate computational alloy design, machine learning models using center-environment (CE) features were first developed to predict substitution energies of alloying elements in molybdenum (Mo)-based alloys. The Random Forest models achieved the optimal performance and transferability-R2 = 0.97, &#x3008;MAE&#x3009; = 0.11 eV, and &#x3008;RMSE&#x3009; = 0.16 eV-against the density functional theory (DFT) benchmark. The model dependency of feature selection and importance analysis was discussed. The transferability of the energy models was validated on unknown systems with new elements. Subsequently, the energy models were fine-tuned using limited mechanical property data to construct energy-to-property (E2P) TL models capable of predicting elastic properties, including bulk modulus, Young's modulus, shear modulus, and elastic constants, achieving an improved accuracy over the non-transferred ML by &#x223c;10-30%, with its transferability verified by additional DFT calculations. This cross-property E2P transfer learning framework opens a new avenue for accelerating computational materials discovery and may be extended to other multiproperty predictions governed by similar physical principles.

center-environment feature

Th2 bias and T-cell exhaustion characterize the immunopathology of non-tuberculous mycobacterial pulmonary disease.

Non-tuberculous mycobacterial pulmonary disease (NTM-PD) is an escalating global health concern with poorly defined immunological mechanisms, necessitating comprehensive profiling to guide therapeutic advances. We analyzed peripheral blood from 28 treatment-na&#xef;ve NTM-PD patients (19 Mycobacterium avium complex, 9 Mycobacterium abscessus) and 27 matched controls using 42-marker mass cytometry (CyTOF) and Luminex multiplex assays. A random forest model identified predictive markers, while an in vitro murine macrophage model evaluated chemokine production. NTM-PD patients displayed significant immune shifts, including increased classical monocytes (CD14+ CD16-), reduced NKT-like cells (CD3+ CD56+), and elevated T-cell exhaustion markers (PD-1, TOX). This coincided with a Th1/Th2 balance shift characterized by heightened IL-13. Elevated IFN-&#x3b3;-inducible chemokines CXCL9 and CXCL10 coexisted with this Th2-biased signature, indicating a complex, dysregulated inflammatory state. A model integrating immune-cell frequencies and cytokine profiles achieved robust diagnostic accuracy (AUC&#x2009;=&#x2009;0.922) with prognostic potential. In vitro, NTM-infected macrophages produced substantial CXCL9 and CXCL10 levels relative to the LPS maximal activation benchmark, identifying them as a major cellular source. These findings propose an immunological framework wherein T-cell exhaustion and a Th2-biased microenvironment strongly correlate with NTM-PD pathogenesis. CXCL9, CXCL10, and IL-13 emerge as candidate therapeutic targets, while our predictive model offers a foundational approach for risk stratification.

Humans