PubMed HealthSearch

SEARCH · PubMed Health

Results for “Deep-learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

SIMS: A deep-learning label transfer tool for single-cell RNA sequencing analysis.

Cell atlases serve as vital references for automating cell labeling in new samples, yet existing classification algorithms struggle with accuracy. Here we introduce SIMS (scalable, interpretable machine learning for single cell), a low-code data-efficient pipeline for single-cell RNA classification. We benchmark SIMS against datasets from different tissues and species. We demonstrate SIMS's efficacy in classifying cells in the brain, achieving high accuracy even with small training sets (<3,500 cells) and across different samples. SIMS accurately predicts neuronal subtypes in the developing brain, shedding light on genetic changes during neuronal differentiation and postmitotic fate refinement. Finally, we apply SIMS to single-cell RNA datasets of cortical organoids to predict cell identities and uncover genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Single-Cell Analysis

Ophthalmic imaging as a measure of cardiovascular and neurological health: a multi-omic analysis of deep-learning derived phenotypes.

The eye is a recognised source of biomarkers for cardiovascular and neurodegenerative disease risk. Here, we characterise the breadth of these associations and identify biological axes that may mediate them. Using UK Biobank data, we developed a multi-omic analysis pipeline integrating physiological, radiomic, metabolomic, and genomic information. We trained adversarial autoencoders (Ret-AAE) to represent optical coherence tomography (OCT) images and colour fundus photographs as 256-dimensional embeddings. Ret-AAE derived embeddings were associated with a range of cardiovascular and neurodegenerative diseases, including ischaemic heart disease, cerebrovascular disease, Parkinson's disease, and dementia. Examining associations across diverse omics datasets, we provide evidence linking ophthalmic imaging features to neurological and cardiovascular anatomy and function, lipid metabolism, and gene sets associated with neurodegenerative pathology. Collectively, our findings demonstrate that ophthalmic features reflect complex, multisystem biological processes, and reinforce the role of the eye as a composite indicator of systemic health.

Journal Article

Deep-Learning Model for Tumor-Type Prediction Using Targeted Clinical Genomic Sequencing Data.

UNLABELLED: Tumor type guides clinical treatment decisions in cancer, but histology-based diagnosis remains challenging. Genomic alterations are highly diagnostic of tumor type, and tumor-type classifiers trained on genomic features have been explored, but the most accurate methods are not clinically feasible, relying on features derived from whole-genome sequencing (WGS), or predicting across limited cancer types. We use genomic features from a data set of 39,787 solid tumors sequenced using a clinically targeted cancer gene panel to develop Genome-Derived-Diagnosis Ensemble (GDD-ENS): a hyperparameter ensemble for classifying tumor type using deep neural networks. GDD-ENS achieves 93% accuracy for high-confidence predictions across 38 cancer types, rivaling the performance of WGS-based methods. GDD-ENS can also guide diagnoses of rare type and cancers of unknown primary and incorporate patient-specific clinical information for improved predictions. Overall, integrating GDD-ENS into prospective clinical sequencing workflows could provide clinically relevant tumor-type predictions to guide treatment decisions in real time. SIGNIFICANCE: We describe a highly accurate tumor-type prediction model, designed specifically for clinical implementation. Our model relies only on widely used cancer gene panel sequencing data, predicts across 38 distinct cancer types, and supports integration of patient-specific nongenomic information for enhanced decision support in challenging diagnostic situations. See related commentary by Garg, p. 906. This article is featured in Selected Articles from This Issue, p. 897.

Humans

The molecular landscape of chordoma: Current frontiers from multi-omics to artificial intelligence.

Chordoma is a rare and aggressive malignant bone tumor of the axial skeleton that has historically challenged clinicians due to its complex anatomical locations and a high recurrence rate of up to 85%. This review synthesizes the most recent advances in chordoma research and offers an overview of how multi-omics, advanced immunology, and artificial intelligence are reshaping the treatment paradigm. Central to its pathogenesis is the T-box transcription factor Brachyury, which this review highlights as both the pathognomonic diagnostic marker and the primary therapeutic vulnerability. Cutting-edge innovations targeting this driver include covalent small-molecule binders, targeted protein degradation, and peptide-centric CAR-T cells designed to attack the intracellular oncoprotein. The tumor immune microenvironment is functionally dynamic, and new dimensions in cellular therapy, such as dual-specific CAR constructs and NK-cell platforms, are being engineered to neutralize immunosuppressive factors. Beyond biological insights, the review emphasizes the role of computational biology, specifically how deep-learning and machine-learning models achieve expert-level precision in tumor segmentation and personalized survival forecasting. By integrating genomic, transcriptomic, epigenomic, and proteomic data, multiomics approaches can fully elucidate chordoma subtypes and underlying resistance mechanisms, ultimately paving the way for more precise and personalized therapeutic strategies.

Humans

Meta-PseU: A meta-classifier for robust prediction of RNA pseudouridine modification sites from long sequences.

BACKGROUND AND OBJECTIVES: Pseudouridine (&#x3a8;) represents one of the most abundant and conserved RNA modifications. &#x3a8; provides an additional hydrogen-bond donor that enhances RNA structural stability and modulates translation. It participates in diverse biological processes, including RNA-protein interactions, splicing, translational control, and stress responses. Aberrant pseudouridylation is implicated in cancer, neurodegenerative disorders, and autoimmune diseases. Despite its biological importance, experimental identification of &#x3a8; sites remains time-consuming and costly, limiting the feasibility of transcriptome-wide profiling. Computational approaches have therefore become essential complements to experimental techniques. However, state-of-the-art machine-learning and deep-learning predictors often suffer from limited generalizability due to small training datasets. To overcome these issues, we aim at constructing new long-sequence datasets and developing a novel &#x3a8; site predictor. METHODS: New long-sequence datasets were constructed as benchmarks for RNA &#x3a8;-site prediction. The &#x3a8; modification sites in RMBase 3.0 were mapped to the reference genomes across three species of human, mouse, and yeast, and the RNA sequences with a length of 201 were generated by extending the upstream and downstream from the mapped, central sites. To eliminate sequence redundancy, the sequences were clustered using CD-HIT with a 70% sequence identity threshold. We developed Meta-PseU, a logistic regression-based meta-classifier that considered 118 machine learning and deep learning classifiers. The datasets and programs are freely accessible at https://github.com/kuratahiroyuki/MetaPseU. RESULTS: By optimizing model configuration, we proposed the Meta-PseU model stacking 32 machine learning and deep learning classifiers out of 118 classifiers. Meta-PseU substantially improved model generalizability, overcoming a key limitation of existing approaches. It greatly outperformed state-of-the-art predictors and achieved increasing accuracy with increasing sequence length. CONCLUSIONS: Long-sequence datasets were newly constructed as benchmarks for RNA &#x3a8;-site prediction. Meta-PseU offers a new framework for robust &#x3a8;-site identification by using long sequences.

Pseudouridine

Colorectal Liver Metastasis Pathomics Model: Integrating Single-Cell and Spatial Transcriptome Analysis With Pathomics for Predicting Liver Metastasis in Colorectal Cancer.

The liver is the primary target organ for hematologic metastasis of colorectal cancer (CRC), and CRC liver metastasis (CRLM) often precludes radical resection, making it the leading cause of death in patients with CRC. To improve the identification and prediction of liver metastasis risk, we identified a cell type of liver metastasis--triggering malignant cells (LMTMCs) through integrating single-cell RNA sequencing and spatial transcriptome analysis. Multiomics cell communication analysis indicated that the interaction between fibroblasts and LMTMCs through the COL1A1-CD44/SDC4 and LAMA4-CD44 signaling axes could promote CRLM. By applying the one-class logistic regression algorithm, we developed a CRLM scoring system in the bulk RNA-sequencing data according to the abundance of LMTMCs in each individual. Using the grouping labels derived from the CRLM scoring system in the bulk data and the corresponding whole-slide images without any manual annotations at the region or pixel level, processed via slide-level weakly supervised learning, a deep-learning model based on the ResNet18 architecture, called Colorectal Liver Metastasis Pathomics Model, was developed to predict the risk of liver metastasis in patients with CRC. The Colorectal Liver Metastasis Pathomics Model achieved an area under the curve of 0.84 at the internal test set of The Cancer Genome Atlas-CRC histology images. In the external independent validation sets, namely the Affiliated Hospital of Southwest Medical University and the Affiliated Traditional Chinese Medicine Hospital of Southwest Medical University cohorts, the areas under the curve were 0.89 and 0.72, respectively, indicating effective classification performances. This study provided new insights and tools for the early identification of CRLM and demonstrated the potential of combining multiomics with deep learning-based pathomics in cancer research.

Humans

Disentangling covariate effects on single-cell-resolved epigenomes with DeepDive.

Understanding the effects of individual biological factors from single-cell-resolved epigenomic data is hindered by multicollinearity, particularly in human cohorts. We introduce DeepDive, a deep-learning framework designed to systematically disentangle known and unknown sources of variation in single-nucleus ATAC-seq data. DeepDive accurately reconstructs chromatin accessibility, outperforms state-of-the-art methods with incomplete covariate information, and robustly recovers true biological signals from even highly entangled covariates, unlocking counterfactual, "what-if," analyses. Applying DeepDive to pancreatic islet cells, we perform counterfactual analyses to prioritize covariates associated with a type 2 diabetes-linked beta-cell subtype and nominate transcription regulators. DeepDive offers a powerful and unbiased tool for mechanistic discovery in complex human disease cohorts.

disentanglement

Unveiling tumor heterogeneity by single cell RNA-sequencing: From basic considerations to clinical applications.

Tumor heterogeneity-encompassing diverse cellular phenotypes, genomic alterations, and microenvironmental contexts-is a principal barrier to effective cancer therapy. Single-cell RNA sequencing (scRNA-seq) has transformed our ability to resolve this complexity by capturing transcriptomes at single-cell resolution. Here, we review the technical foundations required for high-quality scRNA-seq studies. We then trace the evolution of scRNA-seq platforms from manual micromanipulation to high-throughput systems, and describe the computational pipelines that enable reliable data interpretation. The application of scRNA-seq is exemplarily shown in the context of lung cancer, where single-cell profiling has revealed (i) the clonal and sub-clonal architecture of tumors, (ii) extensive remodeling of the immune microenvironment, iii) key mechanisms underlying resistance to targeted agents and immune-checkpoint blockade, and (iv) the dynamics of neo-antigen-specific T-cell responses. Integrating machine-learning techniques-such as deep-learning classifiers and graph-based models-with single-cell transcriptomic data has markedly sped up biomarker discovery, produced more accurate risk-stratification scores, and enabled the generation of patient-specific therapeutic predictions. We surveyed the major trial registry ClinicalTrials.gov and identified &#x223c;380&#xa0;ongoing or completed studies that explicitly incorporate scRNA-seq as a correlative or pharmacodynamic endpoint. Overall, the analysis shows that scRNA-seq becomes an increasingly important component of modern trials, providing high-resolution cellular and molecular readouts that complement conventional imaging and bulk-omics endpoints. While key challenges remain, ranging from costs, scalability and need for rigorous validation before routine clinical deployment, ongoing technological advances continue to expand the potential of scRNA-seq as a cornerstone of precision medicine.

Humans

H&E to recurrence score: A step forward, but not yet a substitute for genomic testing.

Shamai and colleagues developed a multimodal deep-learning model that predicts Oncotype DX recurrence scores from routine H&E slides and clinicopathological variables in hormone receptor&#x2011;positive, HER2&#x2011;negative early breast cancer. Validated across the TAILORx trial and six external cohorts (over 5000 patients), the model achieved an AUC of 0.898 for identifying recurrence score &#x2265;26 and recapitulated genomic assay patterns of chemotherapy benefit. Notably, 31% of clinically high-risk postmenopausal women were downgraded to low risk by AI, suggesting potential to reduce overtreatment. However, several limitations preclude immediate clinical substitution for genomic testing. First, intratumoural heterogeneity leads to discordant predictions with unclear management guidance. Second, the model's chemotherapy benefit estimates rely on TAILORx's age-based menopausal surrogates, which may not reflect real-world hormonal status or LHRH agonist use. Third, predictive value in node-positive disease remains untested in randomised datasets such as RxPONDER. Additionally, calibration uncertainty near risk thresholds and global scalability issues (including IHC requirements and digital pathology infrastructure) persist. While this represents a landmark step toward democratising precision oncology, the AI tool should currently serve as a complementary decision aid, with genomic testing remaining the gold standard for intermediate, borderline, or discordant cases.

Breast cancer

Flexible use of conserved motifs constrains genome access in cell type evolution.

Cell types can be organized into related families, but the regulatory mechanisms that define and maintain these families across deep evolutionary time remain unknown. Here, combining single-nucleus multi-omic sequencing with deep learning to analyse the accessible genomes of two groups of vastly divergent animals including flatworms and vertebrates, we find that hundreds of accessibility-dictating sequence motifs partition into distinct yet conserved sets, or 'vocabularies', each associated with a specific cell type family. However, combinatorial relationships among these motifs preferred by individual cell types are largely species specific. Deep-learning models trained on one species accurately predict family-level chromatin accessibility in distantly related species, albeit frequently rely on different motifs from shared vocabularies to reach convergent predictions. By contrast, models trained on individual cell types within a family lose cross-species predictive power, indicating that the regulatory syntax governing cell type-level identity evolves rapidly. We propose a 'collective maintenance' model in which motif vocabularies defining cell type families are evolutionarily stable, while recombination of these motifs generates cell type-specific regulatory programmes. This suggests that family identity is maintained collectively by large, conserved pools of regulatory factors, analogous to the logic of developmental homology, where character identity persists through network-level conservation despite extensive rewiring.

Journal Article

Accurate somatic small variant discovery for multiple sequencing technologies with DeepSomatic.

Somatic variant detection is an integral part of cancer genomics analysis. While most methods have focused on short-read sequencing, long-read technologies offer potential advantages in repeat mapping and variant phasing. We present DeepSomatic, a deep-learning method for detecting somatic small nucleotide variations and insertions and deletions from both short-read and long-read data. The method has modes for whole-genome and whole-exome sequencing and can run on tumor-normal, tumor-only and formalin-fixed paraffin-embedded samples. To train DeepSomatic and help address the dearth of publicly available training and benchmarking data for somatic variant detection, we generated and make openly available the Cancer Standards Long-read Evaluation (CASTLE) dataset of six matched tumor-normal cell line pairs whole-genome sequenced with Illumina, PacBio HiFi and Oxford Nanopore Technologies, along with benchmark variant sets. Across samples, both cell line and patient-derived, and across short-read and long-read sequencing technologies, DeepSomatic consistently outperforms existing callers.

Humans

CrossAttOmics: multiomics data integration with cross-attention.

MOTIVATION: Advances in high throughput technologies enabled large access to various types of omics. Each omics provides a partial view of the underlying biological process. Integrating multiple omics layers would help have a more accurate diagnosis. However, the complexity of omics data requires approaches that can capture complex relationships. One way to accomplish this is by exploiting the known regulatory links between the different omics, which could help in constructing a better multimodal representation. RESULTS: In this article, we propose CrossAttOmics, a new deep-learning architecture based on the cross-attention mechanism for multiomics integration. Each modality is projected in a lower dimensional space with its specific encoder. Interactions between modalities with known regulatory links are computed in the feature representation space with cross-attention. The results of different experiments carried out in this article show that our model can accurately predict the types of cancer by exploiting the interactions between multiple modalities. CrossAttOmics outperforms other methods when there are few paired training examples. Our approach can be combined with attribution methods like LRP to identify which interactions are the most important. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Sanofi-Public/CrossAttOmics and https://doi.org/10.5281/zenodo.15065928. TCGA data can be downloaded from the Genomic Data Commons Data Portal. CCLE data can be downloaded from the depmap portal.

Humans

SHICEDO: single-cell Hi-C data enhancement with reduced over-smoothing.

MOTIVATION: Single-cell Hi-C (scHi-C) technologies have significantly advanced our understanding of the 3D genome organization. However, scHi-C data are often sparse and noisy, leading to substantial computational challenges in downstream analyses. RESULTS: In this study, we introduce SHICEDO, a novel deep-learning model specifically designed to enhance scHi-C contact matrices by imputing missing or sparsely captured chromatin contacts through a generative adversarial framework. SHICEDO leverages the unique structural characteristics of scHi-C matrices to derive customized features that enable effective data enhancement. Additionally, the model incorporates a channel-wise attention mechanism to mitigate the over-smoothing issue commonly associated with scHi-C enhancement methods. Through simulations and real-data applications, we demonstrate that SHICEDO outperforms the state-of-the-art methods, achieving superior quantitative and qualitative results. Moreover, SHICEDO enhances key structural features in scHi-C data, thus enabling more precise delineation of chromatin structures such as A/B compartments, TAD-like domains, and chromatin loops. AVAILABILITY AND IMPLEMENTATION: SHICEDO is publicly available at https://github.com/wmalab/SHICEDO.

Single-Cell Analysis

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning

PAT: An Image Analysis Tool for Automated Scoring of Pollen in Alexander-Stained Anthers.

Quantitative pollen viability analysis is a critical but labor-intensive step in plant reproductive biology. Existing deep-learning Segment Anything Models (SAM) fail to reliably segment viable pollen in Alexander-stained anthers. To address this, we fine-tuned an existing Cellpose-SAM model for pollen segmentation. We integrated it into PAT (Pollen Analysis Tool), a cross-platform desktop application. PAT features instance segmentation with interactive quality control, an in-app model retraining module, and publication-ready statistical outputs. We deployed PAT in an EMS suppressor screen of semi-sterile Arabidopsis smg7-6 mutants, enabling efficient candidate prioritization for whole-genome sequencing and mapping of the candidate mutation. This screen led to the identification of a point mutation in CAP-D2 (capd2-2), a Condensin I subunit, that rescues the smg7-6 meiotic phenotype. Notably, mutation in a Condensin II subunits (CAP-D3 and CAP-H2) does not confer rescue. Further characterization suggests the capd2-2 allele is hypomorphic, showing no defects in vegetative growth, chromocenter compaction, or transposable element silencing. Collectively, we demonstrate that accessible AI tools have the potential to bridge gaps in plant phenotyping and accelerate the pace of biological discovery.

Alexander staining

Prediction of Atrial Fibrillation From the ECG in the Community Using Deep Learning: A Multinational Study.

BACKGROUND: We aimed to refine and validate a deep neural network model from the ECG to predict atrial fibrillation (AF) risk, using samples from diverse backgrounds: the Framingham Heart Study (FHS), UK Biobank, and Estudo Longitudinal da Sa&#xfa;de do Adulto (ELSA-Brasil). We compared the model's performance to the clinical Cohorts for Heart and Aging Research in Genomic Epidemiology consortium (CHARGE-AF) risk score and evaluated the association with other cardiovascular outcomes. METHODS: The ECG-derived deep-learning prediction of AF (ECG-AF) model was refined using 60% of FHS samples free of AF. Its performance was then tested in the remaining FHS samples, UK Biobank, and ELSA-Brasil, with discrimination assessed by the area under the receiver operating characteristic curve. The association of ECG-AF with cardiovascular outcomes was assessed using Cox proportional hazards models. RESULTS: The study sample included 10&#x2009;097 FHS participants (mean age 53&#xb1;12 years; 54.9% women), 49&#x2009;280 participants from the UK Biobank (mean age 64&#xb1;8 years, 47.9% women), and 12&#x2009;284 participants from ELSA-Brasil (mean age 53&#xb1;8 years, 54.7% women). The ECG-AF model showed moderate discrimination for incident AF (area under the curve, 0.82 [95% CI, 0.80-0.84]) in the FHS, comparable to the CHARGE-AF score (area under the curve, 0.83 [95% CI, 0.81-0.85]), and incremental when combined (area under the curve, 0.85 [95% CI, 0.83-0.87]). In UK Biobank and ELSA-Brasil, combining ECG-AF and CHARGE also improved prediction. Higher ECG-AF scores were associated with increased risks of heart failure, myocardial infarction, stroke, and all-cause mortality in all 3 cohorts. CONCLUSIONS: In multinational cohort studies, the single-input ECG-AF deep neural network model demonstrated good performance in predicting AF and other cardiovascular outcomes, comparable to a multivariable clinical risk score, with improved performance when combined.

Humans

Exploring the use of machine and deep learning in genome-wide association studies: a comprehensive review.

The advent of high-throughput sequencing technologies has generated increasingly large and complex genomic datasets, necessitating analytical approaches capable of capturing high-dimensional and potentially nonlinear genetic interactions. This situation has significantly impacted the entire field of Genome-Wide Association Study (GWAS), whose primary goal is the identification of genomic traits and variants that are statistically associated with the risk of a disease. However, traditional GWAS methods may show reduced performance when applied to highly polygenic and nonlinear genetic architectures. Computational strategies from Artificial Intelligence (AI) and, in particular, from machine- and deep-learning may provide a powerful tool to overcome such limitations, especially by capturing nonlinear interactions and complex hidden regularities in large-scale data, which traditional GWAS approaches might overlook. To date, only a few approaches have been introduced and systematically assessed. In this review, we describe the main characteristics and limitations of standard statistical approaches for GWAS, the main uses of AI methods in computational genomics, and recent attempts to leverage AI strategies in GWAS. Particular attention will be devoted to key issues, such as the interpretability of methods and results, and the curse of dimensionality. More specifically, the review presents 30 methods designed to leverage AI in GWAS, as well as presenting a comprehensive set of evaluation metrics for their performance, also providing references to the most frequently used databases, and biobanks. Overall, this work may serve as a starting point for both dry- and wet-lab researchers, aiming to extract deeper insights from genomic data by moving beyond traditional linear additive assumptions, and leveraging large-scale datasets through AI-driven approaches.

Artificial intelligence

High-Dimensional Sensitivity Analysis for Genomic Studies: An Adversarial Framework for Learning Worst-Case Latent Confounders.

High-dimensional genomics studies are frequently confounded by unmeasured biological processes that obscure disease-specific signals. While existing workflows can estimate these latent confounders, they fail to quantify how robust a discovery is to varying levels of hypothetical confounding. We introduce sensGAN, a deep-learning adversarial framework that systematically explores the confounding spectrum by learning "worst-case" latent variables that nullify the most gene associations under novel predictive-gain constraints. By identifying the minimum confounding strength required to explain away an observed effect, our method shifts the paradigm toward a formal, quantitative sensitivity analysis. In diverse simulations, sensGAN accurately recovers latent structures and outperforms existing methods in identifying confounder-sensitive genes. Applied to human Alzheimer's disease microglia, our framework prioritizes robust disease pathways while successfully isolating signals driven by unmeasured co-occurring neurodegenerative pathologies. Our method is publicly available, deposited at the GitHub repository yifanlinz/ADsensitivityICML.

Journal Article