PubMed HealthSearch

SEARCH · PubMed Health

Results for “Deep Learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8Linked to original sources

Deep generative neural network for accurate drug response imputation.

Drug response differs substantially in cancer patients due to inter- and intra-tumor heterogeneity. Particularly, transcriptome context, especially tumor microenvironment, has been shown playing a significant role in shaping the actual treatment outcome. In this study, we develop a deep variational autoencoder (VAE) model to compress thousands of genes into latent vectors in a low-dimensional space. We then demonstrate that these encoded vectors could accurately impute drug response, outperform standard signature-gene based approaches, and appropriately control the overfitting problem. We apply rigorous quality assessment and validation, including assessing the impact of cell line lineage, cross-validation, cross-panel evaluation, and application in independent clinical data sets, to warrant the accuracy of the imputed drug response in both cell lines and cancer samples. Specifically, the expression-regulated component (EReX) of the observed drug response achieves high correlation across panels. Using the well-trained models, we impute drug response of The Cancer Genome Atlas data and investigate the features and signatures associated with the imputed drug response, including cell line origins, somatic mutations and tumor mutation burdens, tumor microenvironment, and confounding factors. In summary, our deep learning method and the results are useful for the study of signatures and markers of drug response.

Antineoplastic Agents

Refining sequence-to-activity models by increasing model resolution.

Decoding the cis-regulatory syntax that controls gene expression is essential for improving our understanding of cell differentiation and disease. To identify regulatory motifs and their regulatory syntax, deep learning based sequence-to-activity (S2A) models learn transcription factor binding motifs and their combinations from DNA sequence by modeling measured chromatin accessibility. Previously, we developed AI-TAC, a S2A model that predicts chromatin accessibility across various immune cell types in multi-task fashion, effectively decoding the regulatory syntax underlying immune cell differentiation. While ATAC-seq is commonly used to measure regional accessibility, it also provides high-resolution profiles, the distribution of Tn5 insertion sites, that offer additional insights into the precise location and strength of TF binding sites. Here we demonstrate that modeling ATAC-seq profiles alongside accessibility consistently improves predictions of differential chromatin accessibility across cell types. Moreover, we also find that multi-task learning across related immune cell types consistently outperforms single-task models. To understand what additional information bpAITAC learns from ATAC-seq profiles, we systematically compare sequence attributions from models trained with and without ATAC-seq profiles. We identify novel motifs with strong effect sizes that emerge only when profile data is included. Our findings suggest that modeling ATAC-seq at base-pair resolution enables the model to learn a more nuanced and sensitive representation of the cis-regulatory syntax driving immune cell-specific chromatin landscapes.

ATAC-seq

Understanding the sources of performance in deep drug response models reveals insights and improvements.

MOTIVATION: Anti-cancer drug response prediction (DRP) using cancer cell lines (CLs) is crucial in stratified medicine and drug discovery. Recently, new deep learning models for DRP have improved performance over their predecessors. However, different models use different input data types and architectures making it hard to find the source of these improvements. Here we consider published DRP models that report state-of-the-art performance predicting continuous response values. These models take chemical structures of drugs and omics profiles of CLs as input. RESULTS: By experimenting with these models and comparing with our simple baselines, we show that no performance comes from drug features, instead, performance is due to the transcriptomics CL profiles. Furthermore, we show that, depending on the testing type, much of the current reported performance is a property of the training target values. We address these limitations by creating BinaryET and BinaryCB that predict binary drug response values, guided by the hypothesis that this reduces the noise in the drug efficacy data. Thus, better aligning them with biochemistry that can be learnt from the input data. BinaryCB leverages a chemical foundation model, while BinaryET is trained from scratch using a transformer-type architecture. We show that these models learn useful chemical drug features, which is the first time this has been demonstrated for multiple testing types to our knowledge. We further show binarizing the drug response values causes the models to learn useful chemical drug features. We also show that BinaryET improves performance over BinaryCB, and the published models that report state-of-the-art performance. AVAILABILITY AND IMPLEMENTATION: Code is available from https://github.com/Nik-BB/Understanding_DRP_models.

Humans

Nanopore sequencing to detect A-to-I editing sites.

Adenosine-to-inosine (A-to-I) RNA editing, mediated by the ADAR family of enzymes, is pervasive in metazoans and functions as an important mechanism to diversify the proteome and control gene expression. Over the years, there have been multiple efforts to comprehensively map the editing landscape in different organisms and in different disease states. As inosine (I) is recognized largely as guanosine (G) by cellular machineries including the reverse transcriptase, editing sites can be detected as A-to-G changes during sequencing of complementary DNA (cDNA). However, such an approach is indirect and can be confounded by genomic single nucleotide polymorphisms (SNPs) and DNA mutations. Moreover, past studies rely primarily on the Illumina platform, which generates short sequencing reads that can be challenging to map. Recently, nanopore direct RNA sequencing has emerged as a powerful technology to address the issues. Here, we describe the use of the technology together with deep learning models that we have developed, named Dinopore (Detection of inosine with nanopore sequencing), to interrogate the A-to-I editome of any organism.

Inosine

Knowledge-enhanced protein subcellular localization prediction from 3D fluorescence microscope images.

MOTIVATION: Pinpointing the subcellular location of proteins is essential for studying protein function and related diseases. Advances in spatial proteomics have shown that automatic recognition of protein subcellular localization from images could highly facilitate protein translocation analysis and biomarker discovery, but existing machine-learning works have been mostly limited to processing 2D images. By contrast, 3D images have higher spatial resolution and allow researchers to observe cellular structures in their natural context, but currently, there are only a few studies of 3D image processing for protein distribution analysis due to the lack of data and complexity of modeling. RESULTS: We developed a knowledge-enhanced protein subcellular localization model, KE3DLoc, which could recognize distribution patterns in 3D fluorescence microscope images using deep learning methods. The model designs an image feature extraction module that incorporates information from 3D and 2D projected cells and implements asymmetric loss and confidence weights to address data imbalance and weak cell annotation issues. Besides, considering that the biological knowledge in the Gene Ontology (GO) database can provide valuable support for protein location understanding, the KE3DLoc model incorporates a novel knowledge enhancement module that optimizes the protein representation by related knowledge graphs derived from the GO. Since the image module and the knowledge module calculate features from different levels, KE3DLoc designs protein ID aggregation to enhance the consistency of protein features across different cells. Experimental results on three public datasets have demonstrated that the KE3DLoc significantly outperforms existing methods and provides valuable insights for spatial proteomics research. AVAILABILITY AND IMPLEMENTATION: All datasets and codes used in this study are available at GitHub: https://github.com/PRBioimages/KE3DLoc.

Microscopy, Fluorescence

Predictive design of tissue-specific mammalian enhancers that function in the mouse embryo.

Enhancers control tissue-specific gene expression across animals1. Although deep learning2,3 has enabled enhancer prediction and design in mammalian cell lines and non-mammalian model organisms4-10 (reviewed in a previous publication11), it remains unclear whether such approaches can operate within the regulatory complexity of mammalian genomes and tissues in vivo. Here we present a general strategy for designing tissue-specific enhancers that function reliably in mice. We use deep learning to train compact convolutional neural networks on curated chromatin accessibility data and fine-tune them by transfer learning on validated human and mouse enhancers. Guided by these models, we design 15 synthetic enhancers for the heart, limb and central nervous system in mouse embryos, all of which are active in their intended target tissue. These results demonstrate that mammalian enhancer function can be reliably inferred from DNA sequence alone, enabling the predictive de novo design of tissue-specific synthetic enhancers from modest training sets. This work establishes a generalizable framework for programmable control of mammalian gene expression in vivo, opening new avenues in functional genomics, synthetic biology and gene therapy.

Animals

Scalable, generalizable and uncertainty-aware integration of spatial multiomics across diverse modalities and platforms with SCIGMA.

Recent advances in spatial omics technologies have enabled simultaneous profiling of transcriptomic, proteomic, epigenomic, metabolomic and imaging data at high spatial resolution, offering unprecedented opportunities to dissect tissue complexity. However, integrating these diverse and large-scale spatial multimodal datasets remains a major computational challenge. We present SCIGMA, a scalable and generalizable deep learning framework for spatial multiomics integration. SCIGMA introduces an uncertainty-aware contrastive learning objective and multiview graph neural networks to preserve modality-specific signals while learning biologically meaningful joint representations. Unlike previous methods, SCIGMA provides spatially resolved uncertainty estimates, interpretably identifying regions of biological or technical heterogeneity. SCIGMA supports integration of up to five modalities, and its modular framework is extensible to future technologies with even more modalities. It also scales to more than 1 million spatial locations, enabling analysis of high-resolution datasets such as Visium HD and Xenium Prime. We evaluated SCIGMA across 19 datasets spanning 8 modalities, 10 tissues and 9 platforms. On benchmarkable datasets, SCIGMA outperformed other methods in spatial domain detection, modality preservation, feature reconstruction and reproducibility. SCIGMA identifies biologically meaningful structures, refined spatial domains and modality-specific regulatory programs, providing a robust, flexible and future-ready solution for scalable spatial multimodal integration.

Multiomics

stDyer-image improves clustering analysis of spatially resolved transcriptomics and proteomics with morphological images.

MOTIVATION: Spatially resolved transcriptomics (SRT) and spatially resolved proteomics (SRP) data enable the study of gene expression and protein abundances within their precise spatial and cellular contexts in tissues. Certain SRT and SRP technologies also capture corresponding morphology images, adding another layer of valuable information. However, few existing methods developed for SRT data effectively leverage these supplementary images to enhance clustering performance. RESULTS: Here, we introduce stDyer-image, an end-to-end deep learning framework designed for clustering for SRT and SRP datasets with images. Unlike existing methods that utilize images to complement gene expression data, stDyer-image directly links image features to cluster labels. This approach draws inspiration from pathologists, who can visually identify specific cell types or tumor regions from morphological images without relying on gene expression or protein abundances. Benchmarks against state-of-the-art tools demonstrate that stDyer-image achieves superior performance in clustering. Moreover, it is capable of handling large-scale datasets across diverse technologies, making it a versatile and powerful tool for spatial omics analysis. AVAILABILITY AND IMPLEMENTATION: The source code of stDyer-image and detailed tutorials are available at https://github.com/ericcombiolab/stDyer-image.

Proteomics

Essence: A benchmarking-validated transformer framework for early diagnosis of Parkinson's disease using cerebrospinal fluid protein biomarkers.

Parkinson's disease (PD) is a progressive neurodegenerative disorder characterized by motor and non-motor symptoms. The lack of objective molecular biomarkers limits early diagnosis and personalized treatment. Here, we propose Essence, a benchmarking-validated framework integrating cerebrospinal fluid (CSF) proteomics with traditional and deep learning models to identify robust protein signatures for PD. Using data from two independent cohorts, 1266 high-confidence proteins are quantified, among which 178 exhibit differential abundance between PD and healthy controls (HC). Through systematic benchmarking of ten machine learning algorithms and four neural architectures, the Transformer model consistently outperforms alternatives across multiple feature selection strategies, achieving an area under the receiver operating characteristic curve (AUC) of 1.0000 with only 35 features. Functional analyses of the top-ranked 35 proteins reveal enrichment in neuroinflammatory, synaptic, and oxidative stress-related pathways. Importantly, spatial transcriptomic profiling based on the Allen Brain Atlas shows region-specific expression of these biomarkers in PD-relevant brain structures, including the striatum, subthalamic nucleus, hippocampus, and white matter tracts. This anatomical alignment supports the functional relevance of the identified markers and highlights their potential utility in early-stage diagnosis and mechanistic understanding of PD.

Benchmarking

CrossAttOmics: multiomics data integration with cross-attention.

MOTIVATION: Advances in high throughput technologies enabled large access to various types of omics. Each omics provides a partial view of the underlying biological process. Integrating multiple omics layers would help have a more accurate diagnosis. However, the complexity of omics data requires approaches that can capture complex relationships. One way to accomplish this is by exploiting the known regulatory links between the different omics, which could help in constructing a better multimodal representation. RESULTS: In this article, we propose CrossAttOmics, a new deep-learning architecture based on the cross-attention mechanism for multiomics integration. Each modality is projected in a lower dimensional space with its specific encoder. Interactions between modalities with known regulatory links are computed in the feature representation space with cross-attention. The results of different experiments carried out in this article show that our model can accurately predict the types of cancer by exploiting the interactions between multiple modalities. CrossAttOmics outperforms other methods when there are few paired training examples. Our approach can be combined with attribution methods like LRP to identify which interactions are the most important. AVAILABILITY AND IMPLEMENTATION: The code is available at https://github.com/Sanofi-Public/CrossAttOmics and https://doi.org/10.5281/zenodo.15065928. TCGA data can be downloaded from the Genomic Data Commons Data Portal. CCLE data can be downloaded from the depmap portal.

Humans

SHICEDO: single-cell Hi-C data enhancement with reduced over-smoothing.

MOTIVATION: Single-cell Hi-C (scHi-C) technologies have significantly advanced our understanding of the 3D genome organization. However, scHi-C data are often sparse and noisy, leading to substantial computational challenges in downstream analyses. RESULTS: In this study, we introduce SHICEDO, a novel deep-learning model specifically designed to enhance scHi-C contact matrices by imputing missing or sparsely captured chromatin contacts through a generative adversarial framework. SHICEDO leverages the unique structural characteristics of scHi-C matrices to derive customized features that enable effective data enhancement. Additionally, the model incorporates a channel-wise attention mechanism to mitigate the over-smoothing issue commonly associated with scHi-C enhancement methods. Through simulations and real-data applications, we demonstrate that SHICEDO outperforms the state-of-the-art methods, achieving superior quantitative and qualitative results. Moreover, SHICEDO enhances key structural features in scHi-C data, thus enabling more precise delineation of chromatin structures such as A/B compartments, TAD-like domains, and chromatin loops. AVAILABILITY AND IMPLEMENTATION: SHICEDO is publicly available at https://github.com/wmalab/SHICEDO.

Single-Cell Analysis

Accurately Deciphering Tissue Heterogeneity From Spatial Multi-Modal and Multi-Omics With STransformer.

Advances in spatially resolved technologies enable the simultaneous acquisition of diverse data modalities within a tissue slice while preserving critical spatial context, which presents unprecedented opportunities to decipher intricate tissue heterogeneity. However, existing computational approaches lack the intrinsic flexibility to universally process both spatial multi-modal and multi-omics data. Here, we introduce STransformer, a unified deep learning framework designed to seamlessly accommodate a comprehensive landscape of spatial data. By simultaneously capturing short-range cellular interactions and tissue-wide semantic patterns, it extracts robust representations to accurately dissect complex tissue heterogeneity. Systematic evaluations across diverse species, tissue types, and data modalities highlight its profound versatility. For spatial multi-modal data, STransformer delineates intricate anatomical structures in the human cortex, uncovers pathological mechanisms in Alzheimer's disease, and characterizes dynamic spatiotemporal developmental trajectories during chicken cardiogenesis. Scaling to spatial multi-omics data, STransformer synergizes spatial transcriptomic and proteomic profiles to decipher intricate immune microenvironments within the human tonsil, and jointly analyzes spatial epigenomic and transcriptomic data to infer regulatory mechanisms in the mouse embryonic brain. Consequently, STransformer serves as a highly versatile and robust analytical framework for advancing our understanding of tissue heterogeneity and disease pathogenesis.

Multiomics

DiCARN-DNase: enhancing cell-to-cell Hi-C resolution using dilated cascading ResNet with self-attention and DNase-seq chromatin accessibility data.

MOTIVATION: The spatial organization of chromatin is fundamental to gene regulation and essential for proper cellular function. The Hi-C technique remains the leading method for unraveling 3D genome structures, but the limited availability of high-resolution (HR) Hi-C data poses significant challenges for comprehensive analysis. Deep learning models have been developed to predict HR Hi-C data from low-resolution counterparts. Early Convolutional Neural Network (CNN)-based models improved resolution but struggled with issues like blurring and capturing fine details. In contrast, Generative Adversarial Network (GAN)-based methods encountered difficulties in maintaining diversity and generalization. Additionally, most existing algorithms perform poorly in cross-cell line generalization, where a model trained on one cell type is used to enhance HR data in another cell type. RESULTS: In this work, we propose Dilated Cascading Residual Network (DiCARN) to overcome these challenges and improve Hi-C data resolution. DiCARN leverages dilated convolutions and cascading residuals to capture a broader context while preserving fine-grained genomic interactions. Additionally, we incorporate DNase-seq data into our model, providing a robust framework that demonstrates superior generalizability across cell lines in HR Hi-C data reconstruction. AVAILABILITY AND IMPLEMENTATION: DiCARN is publicly available at https://github.com/OluwadareLab/DiCARN.

Chromatin

DeepPlaque: a scalable multimodal platform for Aβ pathology and cell analysis in Alzheimer's disease.

Histological analysis is essential for understanding disease pathology and the microenvironment, particularly in Alzheimer's disease (AD), characterized by beta-amyloid (Aβ) plaques that exist as diffuse, fibrillar, and core species, with distinct toxicity levels. However, accurate classification of Aβ plaque types in postmortem brain tissues and profiling of surrounding cells present significant challenges. To address these challenges, we developed "DeepPlaque", an integrated system featuring "PlaqueNet", a deep learning model for automated classification of Aβ plaque species from diverse imaging platforms. DeepPlaque includes automated workflows for cellular phenotyping and proteomic profiling through targeted laser microdissection. PlaqueNet achieves expert-level accuracy (AUC > 90%) in classifying the 3 major Aβ plaque species, supporting consistent and large-scale annotation. By integrating spatial cellular phenotyping with laser microdissection, DeepPlaque enables high-throughput proteomic analysis of Aβ plaque niches, revealing that microglia are more abundant around core and fibrillar Aβ plaques, with increased expression of apolipoprotein E and amyloid precursor protein in core Aβ plaques. This customizable platform enhances the molecular and cellular characterization of Aβ plaque-associated environments, providing critical insights into AD pathology.

Alzheimer Disease

Artificial Intelligence in Predicting Systemic Complications From Retinal Findings: A New Frontier in Precision Medicine.

Innovations in retinal imaging technologies and growing evidence from retinal imaging of systemic and neurodegenerative diseases have begun to explore the utility of retinal imaging in diagnosing these conditions. Since the retina shares embryological origins with the central nervous system and reflects systemic microvascular characteristics, it is well positioned for noninvasive observation of patients' systemic and neural health. Moreover, accessibility of retinal imaging has improved with the increasing number of ophthalmology clinics. Rapid improvements in various deep learning (DL) tools have also catalyzed the automation of retinal imaging analysis. Systems that utilize DL for retinal imaging are being developed to assist with disease recognition, clinical judgment, and prognostic assessment of systemic health. Various imaging modalities are being integrated with existing genomic and clinical data to estimate an individual's predisposition to certain conditions. Contrary to many existing reviews, the objective of this review is to synthesize the most recent clinical and technological evidence on DL-based diagnostic systems for retinal imaging, with a focus on how different network architectures and their combinations have been developed, validated, and applied across systemic disease detection and prediction. Specifically, this review examines the datasets, model validation approaches, and automated diagnostic systems reported in recent literature. It discusses the extent to which these advancements address existing barriers toward real-time diagnostic application across clinical disciplines. Integrating retinal imaging with DL is an innovative and promising approach to precision medicine and health risk reduction.

artificial intelligence

A neural network model enables worm tracking in challenging conditions and increases signal-to-noise ratio in phenotypic screens.

High-resolution posture tracking of C. elegans has applications in genetics, neuroscience, and drug screening. While classic methods can reliably track isolated worms on uniform backgrounds, they fail when worms overlap, coil, or move in complex environments. Model-based tracking and deep learning approaches have addressed these issues to an extent, but there is still significant room for improvement in tracking crawling worms. Here we train a version of the DeepTangle algorithm developed for swimming worms using a combination of data derived from Tierpsy tracker and hand-annotated data for more difficult cases. DeepTangleCrawl (DTC) outperforms existing methods, reducing failure rates and producing more continuous, gap-free worm trajectories that are less likely to be interrupted by collisions between worms or self-intersecting postures (coils). We show that DTC enables the analysis of previously inaccessible behaviours and increases the signal-to-noise ratio in phenotypic screens, even for data that was specifically collected to be compatible with legacy trackers including low worm density and thin bacterial lawns. DTC broadens the applicability of high-throughput worm imaging to more complex behaviours that involve worm-worm interactions and more naturalistic environments including thicker bacterial lawns.

Caenorhabditis elegans

Lit-OTAR framework for extracting biological evidences from literature.

SUMMARY: The lit-OTAR framework, developed through a collaboration between Europe PMC and Open Targets, leverages deep learning to revolutionize drug discovery by extracting evidence from scientific literature for drug target identification and validation. This novel framework combines named entity recognition for identifying gene/protein (target), disease, organism, and chemical/drug within scientific texts, and entity normalization to map these entities to databases like Ensembl, Experimental Factor Ontology, and ChEMBL. Continuously operational, it has processed over 39 million abstracts and 4.5 million full-text articles and preprints to date, identifying more than 48.5 million unique associations that significantly help accelerate the drug discovery process and scientific research >29.9 m distinct target-disease, 11.8 m distinct target-drug, and 8.3 m distinct disease-drug relationships. AVAILABILITY AND IMPLEMENTATION: The results are accessible through Europe PMC's SciLite web app (https://europepmc.org/) and its annotations API (https://europepmc.org/annotationsapi), as well as via the Open Targets Platform (https://platform.opentargets.org/). The daily pipeline is available at https://github.com/ML4LitS/otar-maintenance, and the Open Targets ETL processes are available at https://github.com/opentargets.

Drug Discovery

PAT: An Image Analysis Tool for Automated Scoring of Pollen in Alexander-Stained Anthers.

Quantitative pollen viability analysis is a critical but labor-intensive step in plant reproductive biology. Existing deep-learning Segment Anything Models (SAM) fail to reliably segment viable pollen in Alexander-stained anthers. To address this, we fine-tuned an existing Cellpose-SAM model for pollen segmentation. We integrated it into PAT (Pollen Analysis Tool), a cross-platform desktop application. PAT features instance segmentation with interactive quality control, an in-app model retraining module, and publication-ready statistical outputs. We deployed PAT in an EMS suppressor screen of semi-sterile Arabidopsis smg7-6 mutants, enabling efficient candidate prioritization for whole-genome sequencing and mapping of the candidate mutation. This screen led to the identification of a point mutation in CAP-D2 (capd2-2), a Condensin I subunit, that rescues the smg7-6 meiotic phenotype. Notably, mutation in a Condensin II subunits (CAP-D3 and CAP-H2) does not confer rescue. Further characterization suggests the capd2-2 allele is hypomorphic, showing no defects in vegetative growth, chromocenter compaction, or transposable element silencing. Collectively, we demonstrate that accessible AI tools have the potential to bridge gaps in plant phenotyping and accelerate the pace of biological discovery.

Alexander staining