PubMed HealthSearch

SEARCH · PubMed Health

Results for “Machine learning integration”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

ToxiVerse: chemical bioprofiling, toxicity data sharing and customizable predictive modeling.

MOTIVATION: Chemical toxicity assessment is critical for drug development and environmental safety. Computational models have emerged as a promising alternative to animal testing and now play a significant role in efficiently evaluating new chemicals. To address the urgent need for user-friendly machine learning tools in computational toxicology, we developed ToxiVerse, a public web-based platform. RESULTS: ToxiVerse provides automatic chemical bioprofiling, curated toxicity datasets, and a predictive modeling interface designed for researchers who lack programming expertise. The platform comprises three integrated modules: (i) Bioprofiler, which provides chemical descriptors by combining chemical-bioactivity data from PubChem assays with a machine learning-based data gap-filling procedure; (ii) Database, which hosts ∼50 000 curated chemicals covering diverse toxicity endpoints; and (iii) Cheminformatics, which enables dataset upload, chemical curation, and automatic generation of quantitative structure-activity relationship models for toxicity prediction. AVAILABILITY: The tool is accessible at www.toxiverse.com, and source code is available at https://github.com/zhu-research-group/toxiverse.

Quantitative Structure-Activity Relationship

Target and biomarker exploration portal for drug discovery.

MOTIVATION: The discovery of novel drug targets and precision biomarkers remains a major challenge in drug development, with traditional differential expression analysis often overlooking key regulatory proteins. Here, we present a novel, web-based bioinformatics tool, the Target and Biomarker Exploration Portal (TBEP), designed to accelerate the drug discovery process by integrating large-scale biomedical data with network analysis techniques. RESULTS: TBEP harnesses machine-learning approaches to mine and combine multimodal datasets, including human genetics, functional genomics, and protein-protein interaction networks, to decode causal disease mechanisms and uncover novel therapeutic targets and precision biomarkers for specific phenotypes. A unique feature of the tool is its ability to process large-scale data in real-time, facilitated by an efficient cloud-based architecture. Additionally, the tool incorporates an integrated large language model (LLM), which assists researchers in exploring and interpreting complex biological relationships within the generated networks and multi-omics data using natural language (English). By offering an intuitive, interactive interface, the LLM enhances the exploration of biological insights, making it easier for scientists to derive actionable conclusions. This powerful integration of network analysis, multi-omics data, and LLM provides a robust framework for accelerating the identification of novel drug targets. AVAILABILITY AND IMPLEMENTATION: The tool is publicly available at https://tbep.missouri.edu. The source code, documentation and installation instructions are available at GitHub repository: https://github.com/mizzoudbl/tbep.

Drug Discovery

Long-read based detection of large copy number variants with potential functional significance using the ContextSV structural variant caller.

Long-read sequencing enables improved detection of structural variants (SVs) in the human genome due to its substantially increased read lengths. However, currently widely used long-read SV callers primarily rely on alignment-based evidence, limiting their ability to detect large and complex SVs and potentially missing disease-relevant events. To address these limitations, we developed ContextSV, a framework that integrates alignment evidence with copy number predictions derived from sequencing coverage and single-nucleotide variant allele frequencies to improve SV detection, particularly for large copy number variants (CNVs). We additionally developed ContextScore, a machine learning-based classification model to assign SV confidence scores based on genomic context features and integrated it within ContextSV. Through benchmarking analyses on both simulated and real datasets, we demonstrate that ContextSV improves detection of large CNVs and inversions that may be missed by existing long-read SV callers. We further illustrate its utility by identifying and experimentally validating multiple large SVs in the KOLF2.1J reference stem cell line that were not detected by other methods. Collectively, our results demonstrate that ContextSV serves as a valuable complement to existing long-read SV detection approaches by improving sensitivity for large and clinically relevant SVs.

Humans

Data-driven approaches in green microbiology: strategies for plant growth-promoting bacteria.

Plant growth-promoting bacteria (PGPB) are gaining attention as scalable biological solutions to enhance crop productivity and resilience. However, accurately identifying and characterizing PGPB remains challenging, particularly under variable environmental conditions where microbial functions are context-dependent and shaped by complex plant-microbe interactions. Advances in high-throughput sequencing have shifted the field from culture-dependent approaches to genome-informed strategies, enabling large-scale taxonomic and functional profiling. Although trait-based databases support the prediction of plant-beneficial genes, they capture only a fraction of the underlying biological complexity and often require labor-intensive analyses. Machine learning (ML) and deep learning (DL) have emerged as powerful tools to integrate genomic, physiological, and ecological data, enabling the prioritization of candidate strains with plant growth-promoting potential. To evaluate advances in the field, we conducted a systematic review of studies integrating ML and DL with PGPB characterization, assessing algorithm selection, performance, and target plant systems. Across 248 observations, only 6.0% of studies directly addressed PGPB screening, whereas the majority (77.4%) focused on plant disease detection, revealing a substantial gap in the application of AI to beneficial microorganisms for plant growth. Convolutional neural networks (CNNs) were the most frequently applied algorithms, largely driven by image-based phenotyping tasks. Overall, the field is constrained by limited datasets, high computational demands, and challenges in modeling multispecies and host-associated interactions. We highlight the need for integrative and interpretable ML and DL frameworks that bridge genomic data and functional validation. Such approaches represent a promising path toward scalable, data-driven discovery and deployment of bioinoculants in sustainable agriculture.

Agriculture

The molecular landscape of chordoma: Current frontiers from multi-omics to artificial intelligence.

Chordoma is a rare and aggressive malignant bone tumor of the axial skeleton that has historically challenged clinicians due to its complex anatomical locations and a high recurrence rate of up to 85%. This review synthesizes the most recent advances in chordoma research and offers an overview of how multi-omics, advanced immunology, and artificial intelligence are reshaping the treatment paradigm. Central to its pathogenesis is the T-box transcription factor Brachyury, which this review highlights as both the pathognomonic diagnostic marker and the primary therapeutic vulnerability. Cutting-edge innovations targeting this driver include covalent small-molecule binders, targeted protein degradation, and peptide-centric CAR-T cells designed to attack the intracellular oncoprotein. The tumor immune microenvironment is functionally dynamic, and new dimensions in cellular therapy, such as dual-specific CAR constructs and NK-cell platforms, are being engineered to neutralize immunosuppressive factors. Beyond biological insights, the review emphasizes the role of computational biology, specifically how deep-learning and machine-learning models achieve expert-level precision in tumor segmentation and personalized survival forecasting. By integrating genomic, transcriptomic, epigenomic, and proteomic data, multiomics approaches can fully elucidate chordoma subtypes and underlying resistance mechanisms, ultimately paving the way for more precise and personalized therapeutic strategies.

Humans

Esketamine multi-omic biomarker evaluation in major depressive disorder (EMBER-MDD): concept, objectives and methodologies of a non-clinical investigator-initiated study.

Treatment resistance (TR) in major depressive disorder (MDD) affects a substantial minority of patients and is hard to recognize early, delaying intensified care. The Esketamine multi-omic biomarker evaluation in MDD (EMBER-MDD) is a non-interventional, investigator-initiated, in-vitro study within the EU Psych-STRATA programme, analyzing biospecimens collected in the randomized INTENSIFY study and the mirror OBS-TR cohort after participants complete treatment. EMBER-MDD aims to discover individual-omic and integrated multi-omic (hypothesis-free) biomarkers and signatures associated with TR risk, and molecular correlates of clinical response to esketamine nasal spray versus treatment as usual (TAU). Biomaterials will derive from approximately 420 adults with MDD (estimated n = 210 esketamine; n = 210 TAU) and include whole blood, RNA-stabilized whole blood, plasma and serum, sampled at baseline and, when feasible, during and after treatment (up to ~ 5,040 aliquots stored at - 80 °C). Genomics will use baseline DNA genotyping on Illumina Infinium GSA v3.0+MD arrays; epigenomics will profile genome-wide DNA methylation across time points using MethylationEPIC v2.0; transcriptomics will employ mRNA-seq (NovaSeq X/ X Plus); and proteomics/ metabolomics will be generated using high-throughput Olink and/ or Biocrates platforms. Each layer will undergo state-of-the-art preprocessing and analyses (e.g., GWAS/ PRS, EWAS, differential expression, WGCNA, pathway and network analyses), followed by integrative strategies including QTL mapping (meQTL/ eQTL/ pQTL/ mQTL) and intermediate-fusion machine learning with nested cross-validation, explainable AI (SHAP/ LIME) and treatment-effect modelling. All outputs are research-only and will not support individual efficacy, tolerability, or clinical decision-making. The study will deliver robust biosignatures and mechanistic hypotheses to guide future validation and inform stratified, molecularly guided intervention strategies in subsequent prospective trials. Trial registration number: 2023-506617-21-00 and 2025-178-f-S.

Humans

Inferring metabolic objectives and trade-offs in single cells during embryogenesis.

While proliferating cells optimize their metabolism to produce biomass, the metabolic objectives of cells that perform non-proliferative tasks are unclear. The opposing requirements for optimizing each objective result in a trade-off that forces single cells to prioritize their metabolic needs and optimally allocate limited resources. Here, we present single-cell optimization objective and trade-off inference (SCOOTI), which infers metabolic objectives and trade-offs in biological systems by integrating bulk and single-cell omics data, using metabolic modeling and machine learning. We validated SCOOTI by identifying essential genes from CRISPR-Cas9 screens in embryonic stem cells, and by inferring the metabolic objectives of quiescent cells, during different cell-cycle phases. Applying this to embryonic cell states, we observed a decrease in metabolic entropy upon development. We further uncovered a trade-off between glutathione and biosynthetic precursors in one-cell zygote, two-cell embryo, and blastocyst cells, potentially representing a trade-off between pluripotency and proliferation. A record of this paper's transparent peer review process is included in the supplemental information.

Single-Cell Analysis

From wild to domestic: Single-cell transcriptomic perspectives on hippocampal regulation and evolution.

How domestication shapes brain evolution remains an open question. In this study, we integrated single-nucleus RNA sequencing (snRNA-seq), population genomics, and machine learning to investigate the hippocampal evolution under domestication. Across-species comparisons revealed that hippocampal cell type profiles are largely conserved across vertebrate species, while supporting the presence of adult hippocampal neurogenesis in birds. We further found that domestication and selective breeding likely influence the cellular composition and molecular regulation of the hippocampus. Our findings provide cellular evidence supporting the hypothesis that domestication affects adult hippocampal neurogenesis. Additionally, we showed that genes associated with neural progenitor cells (NPC) states and cell-marker programs are enriched for signatures of selection. Many of these genes function as regulators of neurogenesis and pathways mediating stress and fear reduction. Specifically, we identified selection at the FKBP5 promoter that may influence its expression in the NPC lineage, potentially contributing to stress-response regulation during domestication. Collectively, these results suggest that domestication is associated with hippocampal remodeling as part of an adaptive response to human-managed environments. This study provides a cellular and genetic perspective on how domestication reshapes the brain and offers a basis for further investigation into the mechanisms of neural evolution within the context of microevolution.

Animals

GRUMB: a genome-resolved metagenomic framework for monitoring urban microbiomes and diagnosing pathogen risk.

SUMMARY: Urban infrastructure hosts dynamic microbial communities that complicate biosurveillance and AMR monitoring. Existing tools rarely combine genome-resolved reconstruction with ecological modeling and batch-aware analytics tailored to infrastructure-scale studies. We present GRUMB (Genome-Resolved Urban Microbiome Biosurveillance), an open-source, SLURM-compatible pipeline that reconstructs high-quality metagenome-assembled genomes (MAGs) from shotgun sequencing reads and integrates taxonomic/functional annotation (CARD, VFDB), batch-aware normalization, ecological diagnostics and machine learning classification of environment types with uncertainty and risk scoring. GRUMB accepts either SRA project accessions or paired-end FASTQ files with metadata, and produces assemblies, MAGs, taxonomic and functional profiles, ecological outputs and risk-informed classification. Its modular design enables reproducible, infrastructure-scale biosurveillance across diverse environments. AVAILABILITY AND IMPLEMENTATION: GRUMB is freely available under the MIT License at: https://github.com/SuleimanAminu/genome-resolved-urban-microbiome-biosurveillance; Zenodo DOI: https://doi.org/10.5281/zenodo.15505402. Requirements: Linux (Ubuntu 20.04+), Python 3.11, R 4.2+, SLURM. Issues and feature requests are tracked on GitHub.

Microbiota

Alzheimer's subtypes A supervised, unsupervised, multimodal, multilayered embedded recursive (SUMMER) AI study.

Since Alzheimer's disease (AD) is a heterogeneous disease, different subtypes may have distinct biological, genetic, and clinical characteristics, requiring tailored interventions. While several proposed subtypes of AD exist, there is still no clear consensus on a definitive classification. By leveraging complementary AI approaches, including supervised and unsupervised learning, within a recursive pipeline (SUMMER) that integrates multimodal datasets encompassing MRI measurements, phenotypes, and genetic data, our goal was to generate robust scientific evidence for identifying AD subtypes. Data was downloaded from the Alzheimer's Disease Neuroimaging Initiative (ADNI) database and included neuroimaging data (MRI), genetics (SNPs), clinical diagnosis, and demographics. 1133 European American participants' images, aged 55-95, were included in this study. The analysis was multi-fold, where the first step involved applying an unsupervised application to a subset of the MRI sample (AD + cognitively normal (CN) aged matched groups, 100 men aged 68-85 years, and 76 women aged 68-85 years). The MRI brain gray matter was segmented into 44 regions of interest (ROIs) according to a standard atlas, and 618 features were extracted, including ROI voxel intensity measurements such as minimum, maximum, and histogram variables. Results identified a cluster of subtype AD men and a cluster of subtype AD women that were distinct from the rest of their respective samples. In the next step, the integrity of the identified subtype AD clusters was investigated using the XGBoost supervised machine learning application with genetic features (SNPs, N=36,724) and labels: the identified subtype AD cluster vs. the rest of the sample, stratified by sex. A significant AD subtype men model (accuracy=0.85, F1=0.72, AUC=0.83) and a significant women AD subtype model (accuracy=0.81, F1=0.81, AUC=0.81) were built, confirming the homogeneity of the isolated AD subtype clusters. Discriminative biomarkers were extracted from the significant models, including selected ROIs and SNPs. Finally, the subtype models were tested on an unseen subset of ADNI data. The genetic-based models identified clusters of AD subtype participants consisting of 34% of the men AD group and 47% of the women AD group. Phenotypic analysis indicates that lower body weight was associated with the women's AD subtype. Complex diseases like AD demand a sophisticated, multimodal approach for precise diagnosis. Effectively identifying disease subtypes enhances the potential for personalized treatment, ultimately improving patient outcomes.

Journal Article

Dynamic evolution of chaperone-mediated autophagy is associated with tumor microenvironment remodeling and prognostic stratification in lung adenocarcinoma: insights from single-cell transcriptomics, ensemble machine learning, and experimental validation.

BACKGROUND: Lung adenocarcinoma (LUAD) shows prognostic heterogeneity, and tumor-node-metastasis (TNM) staging is limited for individualized management. Chaperone-mediated autophagy (CMA) maintains proteostasis, but its role during adenocarcinoma in situ (AIS)-minimally invasive adenocarcinoma (MIA)-invasive adenocarcinoma (IAC) progression remains unclear. METHODS: Single-cell RNA sequencing (scRNA-seq) data from GSE189357 and bulk transcriptomes from The Cancer Genome Atlas (TCGA)-LUAD and Gene Expression Omnibus (GEO) cohorts were integrated. CMA activity, cell-cell communication, weighted gene co-expression network analysis (WGCNA), tumor-normal differential expression, machine-learning survival modeling, tumor microenvironment (TME) features, drug sensitivity, and EPC1 function were analyzed. RESULTS: CMA-high tumor epithelial cells increased from AIS (58.1%) to MIA (65.7%) but declined in IAC (44.4%; p < 0.001). CMA-low cells preferentially received fibroblast-derived extracellular matrix cues. A CMA-negatively correlated module identified 69 core genes. Random survival forest (RSF) performed best among 117 machine-learning combinations (mean concordance index > 0.873). High-risk patients had worse survival across cohorts, and the risk score was independently associated with overall survival (hazard ratio = 16.013, 95% confidence interval: 9.579-26.768, p < 0.001). High-risk tumors showed proliferative activation and M0 macrophage enrichment, whereas low-risk tumors showed stronger immune-related signaling. EPC1 overexpression suppressed malignant phenotypes in A549 cells. CONCLUSION: CMA dynamics are associated with stromal and immune remodeling during LUAD progression. A CMA-based model provides robust prognostic stratification and may offer a basis for future TME-guided studies.

Chaperone-mediated autophagy

Integrative TWAS and multi-omics analyses prioritize HSPE1 as a candidate risk gene for bipolar disorder with immune cell-specific regulatory evidence.

BACKGROUND: Bipolar disorder (BD) is a severe psychiatric disorder associated with substantial disability. Although genome-wide association studies have identified multiple BD-associated loci, the underlying genes and mechanisms remain incompletely understood. METHODS: We integrated a European-ancestry BD genome-wide association dataset with cross-tissue and tissue-specific transcriptome-wide association studies (TWAS) and complementary gene-based analysis. Candidate genes were further evaluated using differential expression analysis, consensus clustering, immune infiltration analysis, machine learning, summary-data-based Mendelian randomization, Mendelian randomization using single-cell expression quantitative trait locus data, single-nucleus transcriptomics, phenome-wide association analysis, and virtual screening. RESULTS: The integrative analyses prioritized 37 candidate genes. Peripheral-blood differential-expression analysis identified 14 genes that remained significant after FDR correction, and their expression profiles separated BD samples into two expression-defined clusters. Machine-learning analysis selected UNC50, LMAN2L, LYG2, HSPE1, and KANSL3 for an exploratory classification nomogram. SMR associated genetically predicted higher HSPE1 expression with increased BD risk in two blood eQTL datasets. Cell-type-specific analyses indicated HSPE1-related associations in T-cell and natural killer cell subsets, while single-nucleus analysis descriptively showed higher HSPE1 expression in medial thalamic T cells from BD samples. PheWAS identified no genome-wide significant associations for HSPE1, whereas virtual screening identified candidate compounds with favorable predicted docking scores against the HSPE1 structure. CONCLUSION: This integrative multi-omics study identified HSPE1 as a candidate BD risk gene with immune-cell-related regulatory evidence, providing insight into BD pathogenesis and supporting functional validation.

Humans

Proteome-wide structural and interaction analysis using cross-linking mass spectrometry and its applications.

Deciphering the mechanisms of protein-protein interactions (PPIs) and protein structural changes within the native cellular environment is crucial for advancing drug discovery. In vivo chemical cross-linking coupled with mass spectrometry (XL-MS) captures weak, transient, and higher-order interactions that are often dysregulated under altered physiological conditions and remain challenging to detect using conventional methods. Applications of in vivo XL-MS range from targeted mapping of PPIs to large-scale identification of interactome networks within the cells. The integration of quantitative approaches further facilitates comparison across different physiological conditions. The recent incorporation of machine learning (ML) tools into XL-MS workflows is transforming the depth and efficiency of this technology. AI-driven algorithms now enable more accurate identification of cross-linked peptides and the mapping of interaction topologies. Furthermore, the synergistic coupling of in vivo XL-MS data with AI-assisted structural modeling platforms such as AlphaFold allows dynamic and high-throughput prediction of protein networks. This review discusses the broader applications of in vivo XL-MS in complex biological samples, ranging from organelles and cells to whole tissues, and highlights how AI integration is expanding structural biology toward a systems-level understanding of proteome architecture.

Mass Spectrometry

Immune-Like Malignant Epithelial Programs Shape Tumor-Immune Interactions and Inform Prognostic Stratification in Lung Adenocarcinoma.

Lung adenocarcinoma (LUAD) is characterized by marked cellular heterogeneity, yet how malignant epithelial states contribute to immune regulation and clinical outcomes remains incompletely defined. We integrated single-cell RNA-sequencing data to map the cellular landscape of LUAD and identify malignant epithelial cells based on inferred copy-number alterations. Epithelial states were further examined through trajectory inference, transcription factor analysis, and cell-cell communication profiling. Single-cell-derived genes were subsequently integrated with TCGA and independent GEO cohorts to construct and validate a machine learning-based prognostic signature. Malignant epithelial cells displayed distinct functional programs, including an immune-like state associated with genomic instability, immune-related transcriptional activity, tumor-immune communication, and patient outcomes. The resulting immune-like malignant epithelial cell signature (IMEC-Sig) consistently stratified survival across multiple cohorts. Low IMEC-Sig scores were accompanied by greater immune infiltration, higher immune checkpoint expression, and increased immunophenoscore, whereas high scores were linked to a comparatively immunosuppressive phenotype. Pan-cancer analyses further identified KRT8 as a gene associated with unfavorable prognosis, and functional experiments showed that KRT8 silencing suppressed proliferation, migration, invasion, and colony formation in LUAD cells. Together, these findings connect malignant epithelial heterogeneity with the immune context and clinical outcomes, support IMEC-Sig as a biologically informed prognostic tool, and nominate KRT8 as a potential therapeutic target in LUAD.

Humans

Machine learning vs. traditional methods for predicting postoperative cardiac complications after non-cardiac surgery: a systematic review and Bayesian network meta-analysis.

INTRODUCTION: Accurate prediction of peri-operative cardiac complications is critical to optimise pre-operative decision-making. Traditional risk prediction scores, such as the Revised Cardiac Risk Index, show only modest discrimination. Machine learning can model complex, non-linear relationships but their predictive performance compared with traditional scores remains unclear. METHODS: We performed a systematic review and Bayesian network meta-analysis. The primary outcome was postoperative adverse cardiac events following non-cardiac surgery. Prediction models were assessed relative to the Revised Cardiac Risk Index. As many studies evaluated multiple versions of each model type, the highest performing ('best version') and lowest performing ('worst version') results were analysed. Models were ranked using the surface under the cumulative ranking curve (SUCRA). RESULTS: Thirteen studies evaluating 54 models and 927,113 patients were included. Machine learning approaches generally outperformed traditional risk scores. Automated machine learning ranked highest (SUCRA 96.6) showed the greatest improvement in the best version analysis (mean difference (MD) 0.28 (95%CrI 0.16-0.40)) and remained superior in the sensitivity analysis (MD 0.30 (95%CrI 0.14-0.45)). Gradient boosting models showed superior performance over the Revised Cardiac Risk Index across analysis (best version: MD 0.20 (95%CrI 0.14-0.26), worst version: MD 0.18 (95%CrI 0.12-0.25), SUCRA 82.4). The Gupta Perioperative Risk for Myocardial Infarction or Cardiac Arrest score outperformed the Revised Cardiac Risk Index in the best version analysis (MD 0.16 (95%CrI 0.01-0.32)). Between-study heterogeneity was low. None of the included studies externally validated their machine learning models and only six were judged to be at low risk of bias. DISCUSSION: Most machine learning models showed better discrimination than traditional risk scores, with automated machine learning and gradient boosting models ranking highest. However, study quality, calibration reporting and absence of external validation limit immediate clinical adoption. Prospective, multicentre evaluation is required before integration of these models into peri-operative practice.

Humans

Measuring Cell Dimensions in Fission Yeast Using Machine Learning.

In fission yeast (Schizosaccharomyces pombe), cell length is a crucial indicator of cell cycle progression. Microscopy screens that examine the effect of agents or genotypes suspected of altering genomic or metabolic stability and thus cell size are crucial for studying disruptions to cell cycle dynamics. This method is based on using an automated cell segmentation algorithm to measure S. pombe cells imaged by brightfield (BF) microscopy methods. PhotoPhenosizer (PP) is a machine learning-based tool designed for automated cell measuring and dimensional analysis of morphology frequency distributions. Integration of this method into large-scale pipelines for tracking cell dimension change streamlines morphological measurements, which facilitates the examination of cellular responses to genomic and metabolic stresses. In this protocol, we use PP to observe the effect of genomic instability on cell size dynamics over a 12-day chronological lifespan assay. Our results show that relative to wild-type cells, a replication stress mutant shows larger cells during chronological aging in excess glucose media. Our results are consistent with activation of checkpoints that regulate cell morphology in response to DNA damage. This method's application highlights the relevance of its incorporation in experimental routines that require large-scale image processing and its adoption by users with routine needs in S. pombe molecular research projects.

Schizosaccharomyces

Q RadFusion: Hybrid Quantum Classical Radiogenomic Framework for Breast Cancer Diagnosis.

BACKGROUND AND PURPOSE: Breast cancer remains the most common cancer in women worldwide, with early and accurate diagnosis critical for patient survival. Radiogenomics integrates imaging phenotypes with genomic profiles, offering a pathway to precision diagnostics. However, existing classical machine learning models often struggle with the high dimensionality and heterogeneity of multimodal data, leading to issues in calibration and reproducibility. This study presents Q RadFusion, a hybrid quantum-classical framework designed to enhance breast cancer diagnosis by fusing mammography and genomics data. METHODS: Q RadFusion was implemented on two publicly available datasets: CBIS-DDSM (2,600 curated mammography cases, TCIA) and TCGA-BRCA (1,000 genomic profiles, GDC). Imaging preprocessing included bias-field correction, segmentation, and harmonization, while genomic data underwent normalization and imputation. Feature selection was performed using the Quantum Approximate Optimization Algorithm (QAOA), and features were mapped into a quantum Hilbert space using Variational Quantum Circuits (VQC). For multimodal fusion, ResNet encoded mammography features, and a Transformer encoded genomic features. Patient-level and site-held-out splits were used for evaluation. RESULTS: Q RadFusion achieved an AUC of 0.96 and accuracy of 94%, outperforming baselines including CNN-LSTM, ResNet + XGBoost, and multimodal Transformers. Ablation studies confirmed the contribution of quantum components, with optimal performance observed at circuit depth, qubits, and QAOA layers. The model also demonstrated improved calibration and ~ 80% fewer parameters compared to deep fusion networks. CONCLUSION: Q RadFusion demonstrates that hybrid quantum-classical radiogenomic integration can deliver accurate, reproducible, and clinically meaningful diagnostic support for breast cancer, with strong potential for future clinical translation.

Breast Cancer

Engineering Bacillus Subtilis for Efficient Biosynthesis of Riboflavin: Current Knowledge and Future Perspectives.

Riboflavin is an essential water-soluble vitamin that serves as a precursor for the biosynthesis of the flavin cofactors FMN and FAD, which play pivotal roles in numerous redox and energy metabolism reactions. With the growing global demand for sustainable vitamin production, microbial fermentation has become an attractive alternative to chemical synthesis due to its environmental and economic advantages. Among microbial hosts, Bacillus subtilis has emerged as a leading cell factory for riboflavin production owing to its GRAS status, well-characterized genetics, and efficient protein secretion system. This review provides a comprehensive overview of recent advances in metabolic engineering strategies to enhance riboflavin biosynthesis in B. subtilis. Key topics include strengthening biosynthetic and precursor pathways, relieving feedback inhibition, balancing metabolic flux and cell growth, employing adaptive laboratory evolution, and utilizing omics-guided optimization and 13C metabolic flux analysis. Moreover, the integration of synthetic biology tools such as riboswitch engineering, regulatory element design, and high-throughput screening has significantly accelerated strain improvement. Despite remarkable progress, challenges remain in achieving precise regulatory control, optimizing multi-gene expression, and enhancing genome integration efficiency. Future research combining multi-omics data, synthetic regulatory design, and machine learning-driven predictive modeling is expected to further advance the development of intelligent B. subtilis cell factories. However, the practical implementation of these systems remains constrained by the metabolic burden of overproduction and the lack of universal regulatory models that can predict strain performance across varying industrial scales.

Bacillus subtilis