PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Augmented kurtosis-based projection pursuit: a novel, advanced machine learning approach for multi-omics data analysis and integration.

Due to the heterogeneity of multi-omics data, exacting their maximum information potential remains a challenge. Whereas some solutions have been offered, most cannot overcome the large linear dynamic range associated with such data, while others require large biological effect sizes to produce meaningful models. Here, we (i) perform a comprehensive benchmarking of multi-omics data analysis tools, and (ii) introduce kurtosis-based projection pursuit analysis, augmented with classification and regression trees (kPPA-CART) as a robust, easy-to-implement alternative. Using ground truth data, we demonstrate that kPPA-CART exhibits superiority in inferring biological significance from low-intensity (low-count) features and studies with small biological effect sizes. Applying it to experimental breast cancer data from The Cancer Genome Atlas, we identify novel genes that cluster the samples into subtypes that mimic the canonical PAM50 classes with notable improvements. Validating with external metastatic breast cancer data from the AURORA US consortium, kPPA-CART identifies genes that are associated with poor event-free survival and additional clustering associated with increased tumor mutational burden. Finally, we provide an R package and an online implementation of kPPA-CART.

Humans

Spectral Transforms as a Tool to Optimize Digital Phenotyping in Biological Images.

Modern livestock breeding has mastered genotyping. Genome-wide association studies, genomic selection, and SNP arrays enable genetic merit prediction at lower cost. However, phenotyping remains the bottleneck, as manual measurement is slow, expensive, subjective, and unable to capture spatial or temporal trait organization. Digital phenotyping via artificial intelligence could resolve this, but deep learning requires thousands of labelled examples, impractical when phenotyping cost itself limits datasets to hundreds of individuals. This creates a paradox: AI could accelerate phenotyping but requires large numbers of samples to train the models. Here, we demonstrate that integrating computer vision with machine learning offers sample-efficient digital phenotyping using eggshell colour as a model system. Rather than learning features from scratch (deep learning), we engineer physically motivated features via Wavelet transforms that decompose images into multi-scale spatial components. Wavelet features captured 14.2 percentage points more variance (R2&#x2009;=&#x2009;0.976 vs. 0.834, p&#x2009;<&#x2009;0.001) than standard colorimetry, with 50% better sample efficiency (achieving at n&#x2009;=&#x2009;60 what colorimetry required n&#x2009;=&#x2009;120). Variance decomposition revealed 77% of discriminative capacity derives from spatial patterns (bands, spots, gradients) invisible to scalar averages. Additionally, we identified "cryptic phenotypes" (3.3%) where spatial patterns contradicted average colour, cases where colorimeters failed but Wavelets succeeded. The underlying principle-that spatial decomposition can recover organizational information lost by scalar averaging-may be applicable to other traits with spatial or temporal structure, such as marbling, dermatitis, or pigmentation rhythms, although whether comparable performance gains would be observed remains to be tested empirically. Hence, for breeding programs implementing genomic selection, computer vision-based digital phenotyping captures complex trait variation without massive training datasets, addressing the bottleneck that increasingly limits genetic progress as genotyping becomes trivial.

Wavelet transform

Machine learning-assisted plasma PEA proteomics enables differential diagnosis of melancholic depression and bipolar disorder.

Differentiating bipolar disorder (BD) from major depressive disorder (MDD) remains a critical unmet need in psychiatry due to overlapping clinical presentations and the absence of reliable biological markers. In this study, we assessed the capacity of multivariate machine learning models to accurately differentiate BD from MDD with melancholic features using plasma proteomic profiles obtained via Proximity Extension Assay (PEA) technology. A total of 67 participants were included (23 BD, 20 MDD, and 24 HC), and plasma protein expression was assessed using the Olink Target 96 Neurology panel. Differential proteomic analysis revealed distinct disorder-specific expression patterns, identifying 21 differentially expressed proteins in BD versus MDD, 18 in BD versus healthy controls, and 7 in MDD versus healthy controls. Using a stepwise feature reduction strategy, machine learning models were trained on three feature sets comprising all proteins, the top 20 most informative proteins, and the top 5 most beneficial proteins, and evaluated across BD-MDD, BD-HC, and MDD-HC classification tasks using five algorithms. For BD-MDD discrimination, the Random Forest model achieved the highest performance when trained on the top 5 protein set (LXN, HAGH, MATN3, PLXNB1, and CTSC), yielding an AUC of 0.905, with similarly strong performance observed using the top 20 protein set. Feature importance analysis highlighted proteins involved in neurodevelopmental processes, immune regulation, and extracellular matrix organization. Overall, these findings demonstrate that integrating plasma proteomics with machine learning enables robust differentiation between BD and MDD with melancholic features, supporting the development of scalable and biologically informed diagnostic tools for precision psychiatry.

Bipolar disorder

Machine Learning and Metabolomics to Characterize Warburg-Like Metabolic Subtypes in Human Retinal Endothelial Cells Exposed to Risk Factors Associated With Proliferative Diabetic Retinopathy.

PURPOSE: High glucose (HG), hypoxia (Hyp), and their combination are major risk factors for proliferative diabetic retinopathy (PDR). Although these conditions induce features of the Warburg-like metabolic reprogramming in human retinal endothelial cells (HRECs), it remains unclear whether they produce distinct metabolic and angiogenic subtypes. This study aimed to characterize the Warburg-like-associated metabolic heterogeneity induced by these PDR-related risk factors and evaluate the ability of supervised machine-learning models to distinguish these subtypes. METHODS: HRECs were cultured under normoglycemic, HG, Hyp (2% O2), and combined HG-Hyp conditions. Untargeted LC-MS/MS metabolomics quantified metabolites spanning carbohydrates, amino acids, nucleotides, and lipids. Principal component analysis (PCA) assessed overall metabolic variation, and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis identified metabolic pathways associated with angiogenesis. In vitro angiogenesis assays measured endothelial tube formation and branching. Nine supervised classifiers (decision tree, logistic regression, na&#xef;ve Bayes, random forest, K-Nearest Neighbors, neural network, gradient boosting, AdaBoost, and Support Vector Machine) were trained on the highest-ranked metabolites selected by the Information Gain Ratio feature-ranking approach. Model performance was evaluated using 10-fold cross-validation, leave-one-out cross-validation (LOOCV), permutation testing, and a classifier stability analysis under biologically meaningful distributional shift using an independent chemically induced hypoxia model (CoCl2). RESULTS: PCA revealed partial separation of metabolic profiles across conditions, indicating different Warburg-like metabolic subtypes. The combined HG-Hyp condition exhibited enhanced angiogenic potential relative to either HG or Hyp alone. KEGG pathway enrichment analysis identified fatty acid biosynthesis and elongation among the most significantly enriched pathways in HRECs under combined HG-Hyp conditions, alongside amino sugar and nucleotide sugar metabolism, glycerophospholipid metabolism, the pentose phosphate pathway, and glycolysis/gluconeogenesis. Supervised machine-learning classifiers distinguished these metabolic subtypes, with AdaBoost and gradient Boosting showing the most balanced, reproducible performance across 10-fold cross-validation, LOOCV, and permutation testing, and remaining the most reliable classifiers under domain-shift testing (area under the curve = 0.88, P = 0.0061). CONCLUSIONS: In this exploratory analysis, HG, Hyp, and their combination drive metabolically and functionally distinct subtypes of Warburg-like metabolic reprogramming in HRECs, with HG-Hyp in combination producing a highly angiogenic phenotype. Boosting-based ensemble classifiers provide a promising framework for detecting these subtypes even under domain-shift conditions, warranting validation in larger independent datasets. TRANSLATIONAL RELEVANCE: Integrating metabolomics with machine-learning classification offers a strategy to identify Warburg-like metabolic subtypes in retinal endothelial cells, providing insights into angiogenic mechanisms and guiding the development of targeted diagnostics or therapeutics for PDR.

Humans

Unraveling 'F' factor: towards a genetic-clinical framework for the musculoskeletal-heart crosstalk in metabolic aging.

BACKGROUND: The rising co-occurrence of cardiometabolic diseases and musculoskeletal degeneration poses a critical challenge to healthy aging, yet the shared biological mechanisms underlying this multimorbidity remain poorly defined. This study aimed to establish an integrative clinical-genetic framework to elucidate the common frailty factor, the 'F' factor, that captures the systemic vulnerability linking cardiometabolic multimorbidity (CMM) and musculoskeletal aging. METHODS: Utilizing the prospective China Health and Retirement Longitudinal Study (CHARLS) cohort, we developed and validated novel Frailty-Integrated Indices for CMM risk prediction, evaluated with machine learning models interpreted via SHapley Additive exPlanations (SHAP). Independently, we applied genomic structural equation modeling (Genomic-SEM) to integrate genome-wide association data from six traits-coronary artery disease, type 2 diabetes, hypertension, bone mineral density, frailty, and telomere length-to model a shared latent genetic factor ('F' factor). This was followed by multivariate GWAS, fine-mapping, transcriptome-wide association study (TWAS), gene-based analysis, and functional annotation to prioritize causal genes, pathways, and cell types. RESULTS: Clinically, several Frailty-Integrated Indices significantly improved CMM risk prediction, with the optimal model achieving an AUC of 0.727. Genetically, we modeled a significant shared latent genetic factor ('F' factor), pinpointing novel risk loci and implicating key genes such as APOE and SLC22A3. These genes were enriched in pathways including cellular senescence and cholesterol metabolism and showed specific expression patterns in developmental brain stages and across multi-organ endothelial cells. CONCLUSION: Our findings provide converging evidence for Musculoskeletal&#x2011;Heart crosstalk of metabolic aging and inferred the 'F' factor as a genetic correlate of a transdiagnostic state, which links genetic predisposition to metabolic dysregulation, and systemic functional decline. This work provides a multi-level biological characterization of multimorbidity liability, informing early-risk detection and preventive strategies for complex aging-related comorbidities.

Humans

Sex Hormone Receptors, HBV Integrations and Their Prognostic Predictive Value Among Hepatocellular Carcinoma Patients.

Hepatocellular carcinoma (HCC) related to hepatitis B virus (HBV) infection predominantly affects males, yet few studies have investigated the association between sex hormones and HBV integrations, and their involvement in HCC prognosis. We assessed estrogen receptor alpha (ER&#x3b1;) and androgen receptor (AR) expression via immunohistochemistry on tissue microarrays constructed from 426 HBV-related HCC samples. HBV integration features were determined using HBV-captured sequencing data. Logistic regression models were utilized to evaluate the association between sex hormone receptor expression level and HBV integration features. Cox regression models, combined with machine learning (ML) methods, were implemented to investigate the prognostic value of sex hormone receptors and HBV integrations concerning overall survival. We found high AR expression level was significantly associated with higher HBV integration levels (adjusted odds ratio [aOR]&#x2009;=&#x2009;1.84, 95% confidence interval [CI]: 1.09-3.11, P for trend&#x2009;=&#x2009;0.012), TERT integration (aOR&#x2009;=&#x2009;2.34, 95% CI: 1.16-4.74, P for trend&#x2009;=&#x2009;0.047), intergenic integration (aOR&#x2009;=&#x2009;2.25, 95% CI: 1.20-4.24, P for trend&#x2009;=&#x2009;0.021), and promoter integration (aOR&#x2009;=&#x2009;1.81, 95% CI: 1.00-3.31, P for trend&#x2009;=&#x2009;0.034). The inclusion of sex hormone receptors and HBV integrations in the predictive models led to improvements across all performance metrics in the Cox regression analyses (AUC improvement: 0.014 [Training], 0.026 [Validation]) and the ML (AUC improvement: 0.022 [Training]), although a slight deterioration in performance was noted in the ML validation set. The results suggested a relationship between AR expression level and HBV integration events, as well as the potential utility of HBV integration biomarkers and sex hormone receptor profiles in assessing post-surgical prognosis among HCC patients.

Humans

Livestock Multi-Omics Integration: A Systematic Framework From Statistical Association to Causal Interpretation.

Livestock multi-omics integration is key to unraveling complex trait regulation, yet systematic, livestock-specific strategies remain scarce. This review traces the progression from single-omics accumulation to multi-dimensional integration, highlighting how large-scale genomic, epigenomic, and transcriptomic projects lay the foundation for functional dissection. We identify core impediments: extreme species diversity, marked data heterogeneity, limited sample sizes, and a pervasive reduction of multi-omics data to simplistic differential screens, resulting in low translational efficiency. We critically appraise four common pitfalls-overinterpreting correlation as causation, relegating proteomics to corroborating transcriptomics, incomplete microbiome-host integration lacking environmental context, and systematic neglect of metabolic fluxomics-and show how exposomics and fluxomics add necessary causal and dynamic dimensions. To address these, we propose a livestock-adapted three-tier analytical framework: (1) statistical association of cross-omics covariation patterns; (2) machine learning-driven feature mining and integrative modeling; and (3) causal interpretation encompassing Mendelian randomization, prior-knowledge-guided network inference, and physical causal evidence via fluxomics and metabolic control analysis. We further discuss how multimodal sequencing (single-cell, spatial, temporal) and generative AI can fundamentally mitigate heterogeneity and strengthen causal evidence. Finally, we outline future priorities in database standardization, livestock-specific benchmarking, and translational pipelines, charting a path from correlation-centric reporting to mechanistic causality and precision breeding.

Animals

Machine learning and multi-omics clustering to map cellular rewiring and immune evasion in ccRCC.

Immune checkpoint blockade (ICB) efficacy in clear cell renal cell carcinoma (ccRCC) is limited by tumor microenvironment (TME) heterogeneity. Because traditional bulk-derived models lack spatial resolution, we developed an integrated framework connecting macroscopic survival risks to microscopic TME structures. We applied ten algorithms to establish multi-omics subtypes and evaluated 101 machine-learning combinations across three independent cohorts to generate a Consensus Machine Learning-driven Signature (CMLS). The signature's spatial and cellular origins were decoded using spatial transcriptomics (ST) and a 140,000-cell scRNA-seq atlas. Expression of key genes was experimentally validated via RT-qPCR in 17 paired ccRCC clinical tissues. We identified two molecular subtypes with distinct clinical and epigenetic profiles. SuperPC optimization yielded a 24-gene CMLS serving as an independent prognostic factor. scRNA-seq and ST deconvolution revealed these signals predominantly originate from cancer-associated fibroblasts (CAFs) and malignant epithelial cells, which collaborate to drive spatial immune exclusion. RT-qPCR confirmed significant overexpression of five core CMLS genes in ccRCC versus adjacent normal tissues. Low CMLS scores correlated with enhanced ICB responsiveness, whereas high-CMLS tumors demonstrated specific vulnerability to dasatinib and dabrafenib. The CMLS translates spatial immune-exclusion dynamics into a quantifiable metric, outperforming tumor mutational burden in predicting ICB benefits, providing a robust tool for patient stratification in ccRCC.

Humans

SSB deficiency-induced R-loop accumulation triggers podocyte inflammation in DKD.

INTRODUCTION: Diabetic kidney disease (DKD) is fundamentally a podocytopathy in which sterile inflammation plays a central pathogenic role, yet the upstream triggers that initiate inflammatory cascades in podocytes remain elusive. R-loops are critical regulators of genomic stability, and their pathological accumulation triggers DNA damage and innate immune activation. Whether R-loop dysregulation contributes to podocyte-driven inflammation in DKD is unknown. METHODS: We integrated single-cell transcriptomic profiling, dual machine learning algorithms, and functional experiments to dissect the R-loop regulatory network in the diabetic kidney. RESULTS: Integrated analysis of human diabetic kidney single-cell RNA-seq data revealed a globally compromised R-loop regulatory network selectively within podocytes. Intersection of podocyte-specific transcriptomic shifts with validated R-loop regulators identified 93 candidate genes, from which dual machine learning algorithms pinpointed SSB (Sj&#xf6;gren syndrome antigen B) as the principal podocyte-selective R-loop resolver and a superior diagnostic biomarker (AUC = 0.983). SSB expression was selectively downregulated in diabetic podocytes and showed the strongest positive correlation with the R-loop resolution module. Mechanistically, SSB loss impaired RNA splicing and stability pathways, leading to aberrant R-loop accumulation that activated the cGAS-dependent inflammatory signaling in podocytes. In two murine DKD models and high glucose-challenged podocytes, SSB was markedly reduced. Remarkably, SSB knockdown in podocytes alone sufficed to trigger R-loop accumulation and pro-inflammatory cytokine expression, whereas both RNase H1-mediated R-loop removal and cGAS co-depletion blunted this response. DISCUSSION: These findings suggest that an SSB-governed R-loop -cGAS -inflammatory signaling axis may link genomic instability to podocyte inflammation and contribute to DKD progression, nominating R-loop homeostasis as a previously unrecognized potential therapeutic target.

Podocytes

Blood-based DNA methylation markers for autism spectrum disorder identification using machine learning.

BACKGROUND: Autism spectrum disorder (ASD) is a complex neurodevelopmental disorder lacking objective biomarkers for early diagnosis. DNA methylation is a promising epigenetic marker, and machine learning offers a data-driven classification approach. However, few studies have examined whole-blood, genome-wide DNA methylation profiles for ASD diagnosis in school-aged children. METHODS: We analyzed genome-wide DNA methylation data from GEO dataset GSE113967, including 52 children with ASD and 48 typically developing (TD) controls. Differentially methylated positions (DMPs) were identified, and feature selection was performed using support vector machine-recursive feature elimination with cross-validation (SVM-RFECV). Classification models were developed using random forest (RF), extreme gradient boosting (XGBoost), and decision tree (DT) classifiers. A nomogram visualized feature contributions. RESULTS: A total of 138 DMPs differentiated ASD from TD children. Eleven CpG sites selected by SVM-RFECV formed the basis for model construction. RF and XGBoost achieved the highest accuracy (75%), with DT reaching 70%. Functional annotation indicated enrichment in cell adhesion and immune-related pathways. CONCLUSIONS: This exploratory study demonstrates the feasibility of integrating peripheral blood DNA methylation data with machine learning to distinguish children with ASD. While limited by sample size and moderate accuracy, this study provides methodological insights into the feasibility of integrating epigenetic and computational approaches for ASD-related biomarker exploration.

Humans

Unveiling the Molecular Secrets of Seaweeds: A Comprehensive Review of Bioinformatics Applications in Algal Research.

Recent advances in high-throughput sequencing, bioinformatics, and multi-omics technologies have transformed seaweed research by overcoming long-standing challenges associated with complex genomes, diverse life cycles, and limited genomic resources. This review provides a comprehensive overview of bioinformatics approaches used to investigate seaweed genomics, transcriptomics, proteomics, metabolomics, microbiomes, and functional genomics, with emphasis on the computational tools and databases that support these analyses. Applications of bioinformatics in phylogenetics, drug discovery, microbiome characterization, and the development of biofuels, nutraceuticals, pharmaceuticals, and sustainable agriculture are also discussed. Particular attention is given to emerging strategies involving multi-omics integration, genome editing, artificial intelligence, machine learning, and synthetic biology that are reshaping seaweed research. The review further examines current challenges, including incomplete genomic resources, data standardization, and the need for experimental validation of computational predictions. Collectively, these advances highlight the growing role of bioinformatics in enabling systems-level understanding of seaweed biology and accelerating their translation into sustainable biotechnological and marine bioeconomy applications.

macroalgal genomics

Potential evaluation of SULT1A3 as an early diagnostic marker for nasopharyngeal carcinoma: a study based on serum proteomics screening and ELISA validation.

BACKGROUND: Nasopharyngeal carcinoma (NPC) represents a highly prevalent and aggressive malignancy endemic to Southeast Asia. Early and accurate diagnosis is critical to improving survival outcomes; however, the absence of robust, stage-specific biomarkers remains a key obstacle to clinical implementation of early screening strategies. METHODS: We performed untargeted serum proteomic profiling using mass spectrometry in 15 treatment-na&#xef;ve early-stage NPC patients and 15 VCA-IgA-positive healthy controls. Bioinformatics analyses were conducted to identify differentially expressed proteins (DEPs). Machine learning (random forest combined with recursive feature elimination) was employed to prioritize candidate biomarkers, which were subsequently verified using enzyme-linked immunosorbent assay (ELISA) in independent sample cohorts. RESULTS: In total, 1,428 serum proteins were identified, among which 1,410 were reliably quantified. We observed 31 upregulated and 189 downregulated proteins in NPC patients relative to controls. Spearman correlation analysis revealed significant associations: LTA4H (leukotriene A4 hydrolase) levels correlated with serum cell infiltration (r&#x2009;=&#x2009;0.383, p&#x2009;=&#x2009;0.032) and CD8&#x2009;+&#x2009;T-cell abundance (r&#x2009;=&#x2009;0.408, p&#x2009;=&#x2009;0.021); both SULT1A3 (sulfotransferase family 1&#xa0;A member 3) and FGL1 (fibrinogen-like protein 1) levels were positively associated with M1 macrophage infiltration (r&#x2009;=&#x2009;0.510, p&#x2009;=&#x2009;0.003 and r&#x2009;=&#x2009;0.430, p&#x2009;=&#x2009;0.015, respectively). In a preliminary validation cohort (n&#x2009;=&#x2009;80), ELISA yielded AUC values of 0.631 (95% CI: 0.515-0.736, p&#x2009;=&#x2009;0.04) for LTA4H, 0.787 (95% CI: 0.681-0.871, p&#x2009;<&#x2009;0.001) for SULT1A3, and 0.688 (95% CI: 0.575-0.787, p&#x2009;=&#x2009;0.002) for FGL1. In large-scale independent validation, SULT1A3 achieved an AUC of 0.826 (95% CI: 0.766-0.876; sensitivity&#x2009;=&#x2009;78.89%, specificity&#x2009;=&#x2009;75.47%) in cohort 1 (n&#x2009;=&#x2009;196) and 0.796 (95% CI: 0.723-0.857; sensitivity&#x2009;=&#x2009;76.67%, specificity&#x2009;=&#x2009;76.67%) in cohort 2 (n&#x2009;=&#x2009;150). CONCLUSIONS: Through an integrated workflow combining proteomic screening, machine learning prioritization, and multi-stage ELISA validation, we identified SULT1A3 as a candidate serum-based biomarker for early detection of NPC. Preliminary findings suggest that SULT1A3 may have potential utility in clinical screening, though further validation in independent, multi&#x2011;center cohorts is required.

Humans

Large-scale discovery platform enables identification of peptides targeting drug-resistant candidiasis.

Natural products have an unparalleled track record as sources of clinical drugs. Among them, nonribosomal peptides (NRPs) stand as one of the most therapeutically significant classes, encompassing numerous approved anti-infective and anticancer agents. Yet, discovering bioactive NRPs remains profoundly challenging due to their complex biosynthesis and chemical architecture. Here, we present NPDiscover, a pathogen-oriented, scalable bioinformatics platform that integrates genome mining, metabolomics, and machine learning to identify NRPs active against drug-resistant pathogens. Applying NPDiscover to Actinobacteria datasets, we discovered edaphochelin A, a previously unreported NRP that kills multi-drug-resistant Candida auris and Candida glabrata by disrupting respiratory chain proteins. Structural elucidation via nuclear magnetic resonance and mass spectrometry, alongside in vitro and in vivo validation, confirmed its efficacy, safety, and a mode of action distinct from existing antifungals-establishing edaphochelin A as a compelling drug candidate and NPDiscover as a powerful engine for scalable natural product discovery.

CP: biotechnology

AutoPVPrimer: A comprehensive AI-Enhanced pipeline for efficient plant virus primer design and assessment.

Plant viruses pose a significant threat to global agriculture and require efficient tools for their timely detection. We present AutoPVPrimer, an innovative pipeline that integrates artificial intelligence (AI) and machine learning to accelerate the development of plant virus primers. The pipeline uses Biopython to automatically retrieve different genomic sequences from the NCBI database to increase the robustness of the subsequent primer design. The design_primers_with_tuning module uses a random forest classifier that optimizes parameters and provides flexibility for different experimental conditions. Quality control measures, including the evaluation of poly-X content and melting temperature, increase primer reliability. Unique to AutoPVPrimer is the visualize_primer_dimer module, which supports the visual evaluation of primer dimers-a feature missing in other tools. Primer specificity is validated via primer BLAST, which contributes to the overall efficiency of the pipeline. AutoPVPrimer has been successfully applied to the tomato mosaic virus, proving its adaptability and efficiency. The modular design allows customization by the user and extends the applicability to different plant viruses and experimental scenarios. The pipeline represents a significant advance in primer design and provides researchers with an effective tool to accelerate molecular biology experiments. Future developments aim to extend compatibility and incorporate user feedback to consolidate AutoPVPrimer as an innovative contribution to the bioinformatics toolbox and a promising resource for the advancement of plant virology research.

DNA Primers

The Computational Revolution in Natural Product Research: A Data-Driven Roadmap for Next-Generation Drug Development.

Natural products (NPs) have historically provided the foundational scaffolds for drug development, yet traditional bioprospecting faces critical limitations: high rediscovery rates, laborious isolation workflows, and substantial attrition during clinical translation. The emergence of big data technologies is fundamentally transforming this landscape, enabling a shift from serendipity-based discovery toward systematic, data-driven approaches. This review examines how the integration of artificial intelligence (AI), machine learning (ML), and multi-omics datasets is accelerating natural product research across three key domains: (1) genome mining for biosynthetic gene cluster identification using platforms such as antiSMASH, (2) cheminformatics-driven prediction of structure-activity relationships and ADMET properties, and (3) metabolomics-guided dereplication to prioritize novel bioactive scaffolds. We evaluate the convergence of genomics, metabolomics, and computational chemistry in enabling in silico lead optimization and the discovery of cryptic metabolites from previously inaccessible microbial taxa. While challenges in data standardization and scalability persist, the synergy between big data and NP research is accelerating clinical translation. Despite persistent challenges in data standardization, scalability, and equitable benefit-sharing, the convergence of big data and NP research is poised to redefine drug development. These advances position computational NP research as a cornerstone of next-generation drug development.

big data analytics

Divergent microbial preludes to necrotising enterocolitis defined by gut phages and bacterial resistomes.

BACKGROUND: Translating microbiome correlations into robust predictive features for complex gut disorders remains elusive, partly due to oversimplified models of pathogenesis and neglect of the virome, a key player in microbial ecosystems. Necrotising enterocolitis (NEC), a devastating disease of preterm infants with no reliable clinical predictors, exemplifies this challenge. OBJECTIVE: To determine the predictive potential of the gut prophageome and polymicrobial aetiologies for NEC. DESIGN: We applied integrated metagenomic and metatranscriptomic analyses and machine learning to 1825 longitudinal stool samples from 43 preterm infants who later developed NEC and 86 gestational age-matched and birthweight-matched controls across three US hospitals. We characterised gut prophageome acquisitions and their association with clinical exposures, including antibiotics, diet and pharmacotherapies. To predict NEC risk, we integrated pre-onset prophageome, antibacterial resistome and bacteriome profiles with neonatal pathology, stratifying the cohort by disease onset timing (early: &#x2264;40 days; late: >40&#x2009;days) for separate analysis. RESULTS: NEC cases exhibited distinct viral diversity trajectories before disease onset. Early-onset NEC was best predicted by phage-bacterial interaction signatures (75% accuracy, 81% sensitivity). Metatranscriptomics revealed increased phage DNA abundance with low gene expression, suggesting a lysogenic lifestyle that may stabilise pathobionts. These phages encode metabolic genes potentially enhancing pathobiont resilience. Late-onset NEC was best predicted by antibacterial resistome profiles (83% accuracy). CONCLUSION: The gut prophageome serves as both a source of pre-symptomatic predictive signals and an active modulator of NEC pathogenesis, with distinct microbial mechanisms driving early-onset and late-onset disease. These polymicrobial etiologies inform strategies for early detection, risk stratification and the development of microbiome-targeted preventive and therapeutic interventions.

BIOMARKERS

Dual-Matrix Platform for Highly Specific Multi-Omics Profiling of Renal Cell Carcinoma.

Multiomics interrogation provides complementary information beyond single-omics approaches for improved disease characterization. To enable such multilayer profiling, we expanded the rapid functionalized mesoporous nanoparticle-coupled laser desorption/ionization mass spectrometry (fMNPLDI-MS) platform by designing two structurally homologous but functionally tailored fMNPs. This design enables efficient acquisition of both serum metabolic and peptide fingerprints from a total of only 2.05 &#x3bc;L of serum, with an LDI MS analysis time of approximately 90 s per sample, while addressing the limitation of single-matrix systems in simultaneously optimizing analytical performance for different biomolecular species. Through statistical analysis and machine learning-based feature selection, an integrated multiomics biomarker panel was established, comprising 5 peptides and 4 metabolites. Notably, this integrated panel outperformed both single-omics panels across all evaluation metrics in the validation set, improving the area under curve from 0.985 to 1.000 and increasing the classification accuracy from 0.947 (metabolites) and 0.930 (peptides) to 0.965, while showing consistent improvements in F1-score, precision, and recall. Collectively, these results demonstrate the robust performance of the dual-matrix design and multiomics integration for renal cell carcinoma classification, with potential relevance for broader applications in complex disease profiling.

Carcinoma, Renal Cell

Opportunities for machine learning to predict cross-neutralization in FMDV serotype O.

Accurately estimating cross-neutralization between serotype O foot-and-mouth disease viruses (FMDVs) is critical for guiding vaccine selection and disease management. In this study, we developed a machine learning approach to estimate r1 values-an established measure of antigenic similarity-using VP1 sequence data and published virus neutralization titer (VNT) results. Our dataset comprised 108 serum-virus pairs representing 73 distinct FMDV strains. We applied Boruta feature selection and random forest classifiers, optimizing model performance through tenfold cross-validation and sub-sampling to address class imbalance. Predictors included pairwise amino acid distances, site-specific polymorphisms, and differences in potential N-glycosylation sites. Using a 0.3 r1 threshold to define cross-neutralization, the final model achieved high accuracy (0.96), sensitivity (0.93), and specificity (0.96) in training, and performed robustly on independent test sets - accuracy was 0.75 (95% CI 0.60 and 0.90), F1 score 0.86% and PPV 0.77. Importantly, key VP1 residues-positions 48, 100, 135, 150, and 151-emerged as strong predictors of antigenic relationships. Our results demonstrate the utility of integrating routinely generated genomic data with machine learning to inform vaccine candidate selection and anticipate immune interactions among circulating FMDV strains. This approach offers a practical tool for accelerating vaccine decision-making and can be adapted to other FMDV serotypes. The latest version of the r1 predictive model is available for access via a Shiny dashboard (https://dmakau.shinyapps.io/PredImmune-FMD/).

Foot-and-Mouth Disease Virus