PubMed HealthSearch

SEARCH · PubMed Health

Results for “Neural Networks, Computer”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

HallmarkGraph: a cancer hallmark informed graph neural network for classifying hierarchical tumor subtypes.

MOTIVATION: Accurate tumor subtype diagnosis is crucial for precision oncology, yet current methodologies face significant challenges. These include balancing model accuracy with interpretability and the high costs of generating multi-omics data in clinical settings. Moreover, there is a lack of validated models capable of classifying hierarchical tumor subtypes across a comprehensive pan-cancer cohort. RESULTS: We present a graph neural network, HallmarkGraph, the first biologically informed model developed to classify hierarchical tumor subtypes in human cancer. Inspired by cancer hallmarks, the model's architecture integrates transcriptome profiles and gene regulatory interactions to perform multi-label classification. We evaluate the model on a comprehensive pan-cancer cohort comprising 11 476 samples from 26 primary cancers with 405 subtypes up to eight levels. The model demonstrates exceptional performance, achieving 5-fold cross-validation accuracy between 85% and 99% for tumor subtypes labeled with increasing details of genomic information. It also shows good generalizability on a validation dataset of 887 samples, assessed using three metrics that consider tumor subtypes at individual, combined, and sample levels. Benchmarking and ablation experiments show that hallmark-based embeddings slightly influence model performance, while the integrated multilayer perceptron plays a significant role in determining classifier accuracy. Additionally, we use the SHAP method to link cancer hallmarks with genes, identifying key features that influence model decisions. Our findings present a biologically informed machine learning framework capable of tracking tumor transcriptomic trajectories and distinguishing inter- and intra-tumor heterogeneity in pan-cancer. This approach holds promise for enhancing cancer diagnostics. AVAILABILITY AND IMPLEMENTATION: HallmarkGraph is accessible at https://github.com/laixn/HallmarkGraph.

Humans

Development and validation of a deep learning model based on cascade mask regional convolutional neural network to noninvasively and accurately identify human round spermatids.

INTRODUCTION: The difficulty of identifying human round spermatids (hRSs) has impeded applications of the human round spermatid injection (ROSI) technique. RSs can be accurately screened through flow cytometric analysis utilizing the Hoechst fluorescence profile reflecting DNA, but this method is not suitable for isolating hRSs due to the toxicity associated with Hoechst staining. OBJECTIVE: To evaluate the capacity of a deep learning model grounded in a cascade mask region-based convolutional neural network (R-CNN) for the noninvasive and accurate identification of hRSs. METHODS: In this study, we presented the development and validation of a deep learning model for identifying hRSs through the analysis of 3457 optical light microscope images of sorted hRSs obtained via flow cytometric analysis. The model's accuracy and specificity were evaluated by calculating the mean average precision (mAP). Furthermore, a double-blind experiment was conducted to access the reliability of the proposed model in accurately identifying hRSs. It detected the expression of protamine (PRM1) and/or peanut lectin (PNA), which are established markers for RSs. RESULTS: Our deep learning-based model demonstrated a high precision, achieving a mAP of over 0.80 for isolating hRSs in test datasets. The expression of PRM1 and/or PNA was observed in all cells noninvasively selected by our AI model during an independent double-blind test. This phenomenon confirmed the accuracy and effectiveness of the proposed model. The model's capability for noninvasive and accurate isolation of hRSs among spermatogenic cells highlighted its robustness and generalizability for clinical applications. CONCLUSION: The deep learning AI model based on a cascade R-CNN has the ability to accurately identify hRSs among spermatogenic cells. The application of this noninvasive method, which requires no additional procedures in clinical practice, is able to facilitate the widespread implementation of ROSI technique. Therefore, it can provide patients with spermatogenic arrest the opportunity to become biological fathers.

Humans

Detecting Interspecific Positive Selection Using Convolutional Neural Networks.

Traditional statistical methods using maximum likelihood and Bayesian inference can detect positive selection from an interspecific phylogeny and a codon sequence alignment based on model assumptions, but they are prone to false positives due to alignment errors and can lack power. These problems are particularly pronounced when faced with high levels of indels and divergence. To address these issues, we trained and tested convolutional neural network models on simulated data and achieved higher accuracy in detecting selection across a specific range of phylogenetic scenarios and evolutionary modes. This advantage is particularly evident when performing inference on noisy data prone to misalignments. Our method shows some ability to account for these errors, where most statistical frameworks fail to do so in a tractable manner. We explore the generalizability of our convolutional neural network models to unseen evolutionary scenarios and identify future avenues to achieve broader utility. Once trained, our convolutional neural network model is faster at test time, making it a scalable alternative to traditional statistical methods for large-scale, multigene analyses. In addition to binary classification (inference of the presence or absence of positive selection during the evolution of the sequences), we use saliency maps to understand what the model learns and observe how this could be leveraged for sitewise inference of positive selection.

Neural Networks, Computer

Three-dimensional U-Net with transfer learning improves automated whole brain delineation from MRI brain scans of rats, mice, and monkeys.

BACKGROUND: Automated whole-brain delineation (WBD) techniques often struggle to generalize across pre-clinical studies due to variations in animal models, magnetic resonance imaging (MRI) scanners, and tissue contrasts. We developed a 3D U-Net neural network for WBD pre-trained on organophosphate intoxication (OPI) rat brain MRI scans. We used transfer learning (TL) to adapt this OPI-pretrained network to other animal models: rat model of Alzheimer's disease (AD), mouse model of tetramethylenedisulfotetramine (TETS) intoxication, and titi monkey model of social bonding. METHODS: We assessed an OPI-pretrained 3D U-Net across animal models under three conditions: (1) direct application to each dataset; (2) utilizing TL; and (3) training disease-specific U-Net models. For each condition, training dataset size (TDS) was optimized, and output WBDs were compared to manual segmentations for accuracy. RESULTS: The OPI-pretrained 3D U-Net (TDS = 100) achieved the best accuracy [median[min-max]] for the test OPI dataset with a Dice coefficient (DC) = [0.987 [0.977-0.992]] and Hausdorff distance (HD) = [0.86 [0.55-1.27]]mm. TL improved generalization across all models [AD (TDS = 40): DC = 0.987 [0.977-0.992] and HD = 0.72 [0.54-1.00]mm; TETS (TDS = 10): DC = 0.992 [0.984-0.993] and HD = 0.40 [0.31-0.50]mm; Monkey (TDS = 8): DC = 0.977 [0.968-0.979] and HD = 3.03 [2.19-3.91]mm], showing performance comparable to disease-specific networks. CONCLUSIONS: The OPI-pretrained 3D U-Net with TL achieved accuracy comparable to disease-specific networks with reduced training data (TDS ≤ 40 scans) across all models. Future work will focus on developing a multi-region delineation pipeline for pre-clinical MRI brain data, utilizing the proposed WBD as an initial step.

Animals

IGCN: integrative graph convolution networks for patient level insights and biomarker discovery in multi-omics integration.

MOTIVATION: Developing computational tools for integrative analysis across multiple types of omics data has been of immense importance in cancer molecular biology and precision medicine research. While recent advancements have yielded integrative prediction solutions for multi-omics data, these methods lack a comprehensive and cohesive understanding of the rationale behind their specific predictions. To shed light on personalized medicine and unravel previously unknown characteristics within integrative analysis of multi-omics data, we introduce a novel integrative neural network approach for cancer molecular subtype and biomedical classification applications, named Integrative Graph Convolutional Networks (IGCN). RESULTS: To demonstrate the superiority of IGCN, we compare its performance with other state-of-the-art approaches across different cancer subtype and biomedical classification tasks. Our experimental results show that our proposed model outperforms the state-of-the-art and baseline methods. IGCN identifies which types of omics data receive more emphasis for each patient when predicting a specific class. Additionally, IGCN has the capability to pinpoint significant biomarkers from a range of omics data types. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/bozdaglab/IGCN.

Humans

Transfer learning with multiomics integration and deep neural networks reveals drug resistance mechanisms in cancer.

Drug resistance remains one of the primary challenges in effective cancer therapy. In this study, we employed a deep neural network (DNN)-based transfer learning (TL) approach to predict drug response and uncover drug resistance mechanisms. We integrated gene expression, somatic mutation, and copy number aberration (CNA) data with drug response profiles using multi-omics integration (MI). We used the Genomics of Drug Sensitivity in Cancer (GDSC) data for training and incorporated drugs with same pathways into the training models. We then evaluated drug response predictions on independent in-vivo PDX Encyclopedia (PDX) and ex-vivo the Cancer Genome Atlas (TCGA) datasets. In addition, we conducted pathway enrichment analyses to elucidate the mechanisms underlying drug resistance for paclitaxel, 5-fluorouracil (5-FU), gemcitabine, and cetuximab. We also applied Fisher's exact test (FET) to assess potential associations between drug resistance and the presence of mutations or CNAs. Our pan-drug models outperformed other methods based on the area under the precision-recall curve (AUCPR). Our pathway enrichment analyses revealed LDHB-mediated pyruvate metabolism and FYN-mediated focal adhesion might have pivotal roles in paclitaxel resistance, while PINK1-mediated mitophagy might be critical in 5-FU resistance. In addition to transcriptional activation, FET suggested that CNAs in LDHB and PINK1 may also be associated with resistance to paclitaxel and 5-FU, respectively. Furthermore, enrichment results for paclitaxel and cetuximab indicated shared resistance mechanisms between the two drugs. Importantly, our findings are consistent with prior experimental studies, providing literature-based validation of our results. Overall, our DNN-based TL approach achieved strong predictive performance across PDX & TCGA datasets and enrichment analyses provided valuable biological insights into drug resistance mechanisms.

Humans

Artificial neural network data fusion-mediated dual-mode sensor based on Fe3O4@PdIr for Salmonellatyphimurium detection in food.

Salmonella Typhimurium (S. typhimurium) is a major foodborne pathogen that poses a serious threat to public health. In this study, a colorimetric/electrochemical dual-mode biosensor assisted by artificial neural network (ANN) was developed for the sensitive detection of S. typhimurium. Fe3O4@PdIr nanocomposites with enhanced peroxidase-like activity and electrochemical performance were prepared and conjugated with an aptamer specific to S. typhimurium to obtain Fe3O4@PdIr-Apt. Through the sandwich binding of Fe3O4@PdIr-Apt and Apt to the target, the nanocomposites were attached to microplates or Au electrodes, thereby generating colorimetric and electrochemical signals. The ANN model deeply resolved the complex nonlinear relationship between the dual signals, enabling mutual correction and ultimately performing data fusion to output a single detection result, which significantly reduced the mean square error while improving detection sensitivity and reliability. This sensor exhibited a wide linear range of 2.7-2.7 × 108 CFU/mL and a low detection limit of 1.66 CFU/mL. Additionally, this method was successfully applied to the detection of S. typhimurium in pork and milk, with a recovery rate of 95.19% ∼ 104.07%. It indicated that the constructed sensor holds great practical potential for S. typhimurium detection.

Neural Networks, Computer

Network methods for diagonal integration of unpaired single-cell multiomics data: a review.

MOTIVATION: Advances in single-cell sequencing have enabled multiomics profiling at unprecedented resolution; however, mass spectrometry-based single-cell proteomics (scMS) remains inherently destructive, precluding simultaneous transcriptomic capture. Unlike antibody-based methods such as CITE-seq, which permit paired profiling but are restricted to targeted protein panels, scMS provides unbiased, genome-scale coverage of the intracellular proteome yet necessitates post hoc integration of unpaired datasets. This diagonal integration challenge, where transcriptomes and proteomes are measured in separate cells lacking shared anchors, remains underserved by existing reviews, which focus predominantly on vertical integration strategies enabled by non-destructive assays. RESULTS: We survey the complete computational pipeline for constructing mechanistic proteogenomic networks from unpaired single-cell data, covering: (i) unimodal network inference such as knowledge-based approaches, probabilistic graphical models, temporal directionality inference, and generative and foundation model strategies that establish the transcriptomic scaffold; (ii) cross-modal integration architectures such as network propagation, graph neural networks (scMRDR, scmFormer, scCotag), and consensus frameworks designed explicitly for the unpaired proteomics setting; and (iii) benchmarking paradigms spanning network reconstruction (BEELINE, GRETA, CausalBench) and multi-task integration evaluation (scMultiBench, SCMMIB), with guidance on metric selection under network sparsity and class imbalance. We identify three principal axes of future development: generative proteomic translation from transcriptomic precursors, inductive prior embedding in next-generation architectures, and perturbation-based causal benchmarking. AVAILABILITY AND IMPLEMENTATION: This is a review article; no novel software is distributed. A curated benchmark resource table, methods starter guide, and per-method bottleneck annotations are provided in the Supplementary Material.

Multiomics

Chromatin structures from integrated AI and polymer physics model.

The physical organization of the genome in three-dimensional space regulates many biological processes, including gene expression and cell differentiation. Three-dimensional characterization of genome structure is critical to understanding these biological processes. Direct experimental measurements of genome structure are challenging; computational models of chromatin structure are therefore necessary. We develop an approach that combines a particle-based chromatin polymer model, molecular simulation, and machine learning to efficiently and accurately estimate chromatin structure from indirect measures of genome structure. More specifically, we introduce a new approach where the interaction parameters of the polymer model are extracted from experimental Hi-C data using a graph neural network (GNN). We train the GNN on simulated data from the underlying polymer model, avoiding the need for large quantities of experimental data. The resulting approach accurately estimates chromatin structures across all chromosomes and across several experimental cell lines despite being trained almost exclusively on simulated data. The proposed approach can be viewed as a general framework for combining physical modeling with machine learning, and it could be extended to integrate additional biological data modalities. Ultimately, we achieve accurate and high-throughput estimations of chromatin structure from Hi-C data, which will be necessary as experimental methodologies, such as single-cell Hi-C, improve.

Chromatin

GiantHost: a domain-adaptive and uncertainty-aware framework for giant virus host prediction.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) play crucial roles in global ecosystems. Although metagenomics has vastly accelerated the discovery of novel NCLDVs, predicting their hosts from fragmented contigs remains a critical bottleneck, with no dedicated end-to-end computational tools currently available. Addressing this gap requires overcoming three fundamental challenges: the extreme scarcity of labeled reference genomes, the severe domain shift between laboratory isolates and diverse environmental metagenomes, and the inability of traditional deterministic models to quantify prediction uncertainty-a crucial requirement for reliable ecological profiling where novel, divergent viruses are prevalent. RESULTS: We present GiantHost, the first NCLDV host prediction tool with domain adaptation and uncertainlty awareness. GiantHost employs a dual-tower neural network to integrate dense genome traits and sparse GVOG profiles, allowing better integration of heterogeneous features. To overcome label scarcity and domain shift, we leverage 1400 environmental viral genomes (GVMAGs) via semi-supervised multi-task learning and Domain Adversarial Neural Networks (DANN), effectively bridging the distributional gap between RefSeq and environmental data. Additionally, GiantHost incorporates Conformal Prediction (CP) to output statistically guaranteed prediction sets rather than overconfident single labels. Evaluated under rigorous genome-level cross-validation, GiantHost demonstrates robust predictive power. Applied to the Tara Ocean dataset, GiantHost successfully captured the vertical stratification of NCLDV hosts-revealing a depth-dependent decline of phytoplankton-infecting viruses and a relative enrichment of Amoebozoa-infecting viruses in the mesopelagic zone. AVAILABILITY: The source code of GiantHost is available via: https://github.com/FuchuanQu/GiantHost.

Giant Viruses

Comparative analysis of convolutional neural network models for the histopathological differentiation of acinic cell carcinoma and secretory carcinoma.

OBJECTIVE: Although artificial intelligence tools show promise for enhancing the diagnosis of head and neck lesions, few studies have tested these resources for the microscopic diagnosis of salivary gland tumors. Specifically, the microscopic differentiation between acinic cell carcinoma and secretory carcinoma has never been addressed in this context. Therefore, this exploratory study aimed to comparatively evaluate the feasibility of applying convolutional neural networks for the microscopic differentiation between acinic cell carcinomas and secretory carcinomas. METHODS: A cross-sectional study using whole-slide images from 46 patients with acinic cell carcinoma (n = 26) or secretory carcinoma (n = 20) was conducted. Eight CNNs (ResNet-50, InceptionV3, VGG16, Xception, MobileNet, DenseNet121, EfficientNetB0, and EfficientNetV2B0) were trained and evaluated for accuracy, sensitivity, specificity, F1-score, and AUC. Performance was measured in training, validation, and test subsets. Accuracy and loss curves were also presented. RESULTS: InceptionV3 demonstrated the best overall performance, with the lowest loss (1.39), highest accuracy (0.81), sensitivity (0.90), and F1-score (0.81). VGG16 achieved the highest AUC (0.86) and precision (0.77). DenseNet121 showed the lowest performance in terms of accuracy (0.65) and F1-Score (0.52), but the highest specificity (0.85). CONCLUSION: This proof-of-concept study suggests that convolutional neural networks may be feasible tools to support the microscopic differentiation between acinic cell carcinoma and secretory carcinoma. The performance of these models critically depends on the size of the dataset and the quality of annotations. The findings should be interpreted cautiously given the limited dataset and potential sources of bias. Further validation with larger, multicenter datasets is needed before any clinical application can be considered.

Humans

A multi-modal transformer for cell type-agnostic regulatory predictions.

Sequence-based deep learning models have emerged as powerful tools for deciphering the cis-regulatory grammar of the human genome but cannot generalize to unobserved cellular contexts. Here, we present EpiBERT, a multi-modal transformer that learns generalizable representations of genomic sequence and cell type-specific chromatin accessibility through a masked accessibility-based pre-training objective. Following pre-training, EpiBERT can be fine-tuned for gene expression prediction, achieving accuracy comparable to the sequence-only Enformer model, while also being able to generalize to unobserved cell states. The learned representations are interpretable and useful for predicting chromatin accessibility quantitative trait loci (caQTLs), regulatory motifs, and enhancer-gene links. Our work represents a step toward improving the generalization of sequence-based deep neural networks in regulatory genomics.

Humans

Deep generative neural network for accurate drug response imputation.

Drug response differs substantially in cancer patients due to inter- and intra-tumor heterogeneity. Particularly, transcriptome context, especially tumor microenvironment, has been shown playing a significant role in shaping the actual treatment outcome. In this study, we develop a deep variational autoencoder (VAE) model to compress thousands of genes into latent vectors in a low-dimensional space. We then demonstrate that these encoded vectors could accurately impute drug response, outperform standard signature-gene based approaches, and appropriately control the overfitting problem. We apply rigorous quality assessment and validation, including assessing the impact of cell line lineage, cross-validation, cross-panel evaluation, and application in independent clinical data sets, to warrant the accuracy of the imputed drug response in both cell lines and cancer samples. Specifically, the expression-regulated component (EReX) of the observed drug response achieves high correlation across panels. Using the well-trained models, we impute drug response of The Cancer Genome Atlas data and investigate the features and signatures associated with the imputed drug response, including cell line origins, somatic mutations and tumor mutation burdens, tumor microenvironment, and confounding factors. In summary, our deep learning method and the results are useful for the study of signatures and markers of drug response.

Antineoplastic Agents

AI-driven multi-omics modeling of myalgic encephalomyelitis/chronic fatigue syndrome.

Myalgic encephalomyelitis/chronic fatigue syndrome (ME/CFS) is a chronic illness with a multifactorial etiology and heterogeneous symptomatology, posing major challenges for diagnosis and treatment. Here we present BioMapAI, a supervised deep neural network trained on a 4-year, longitudinal, multi-omics dataset from 249 participants, which integrates gut metagenomics, plasma metabolomics, immune cell profiling, blood laboratory data and detailed clinical symptoms. By simultaneously modeling these diverse data types to predict clinical severity, BioMapAI identifies disease- and symptom-specific biomarkers and classifies ME/CFS in both held-out and independent external cohorts. Using an explainable AI approach, we construct a unique connectivity map spanning the microbiome, immune system and plasma metabolome in health and ME/CFS adjusted for age, gender and additional clinical factors. This map uncovers altered associations between microbial metabolism (for example, short-chain fatty acids, branched-chain amino acids, tryptophan, benzoate), plasma lipids and bile acids, and heightened inflammatory responses in mucosal and inflammatory T cell subsets (MAIT, γδT) secreting IFN-γ and GzA. Overall, BioMapAI provides unprecedented systems-level insights into ME/CFS, refining existing hypotheses and hypothesizing unique mechanisms-specifically, how multi-omics dynamics are associated to the disease's heterogeneous symptoms.

Humans

Adversarial attack of sequence-free enhancer prediction identifies chromatin architecture.

MOTIVATION: The wide range of cellular complexity created by multicellular organisms is due in large part to the intricate and synergistic interplay of regulatory complexes throughout the eukaryotic genome. These regulatory elements "enhance" specific gene programs and have been shown to operate in diverse networks that are distinct across cell states of the same organism. Attempts to characterize and predict enhancers have typically focused on leveraging information-dense DNA sequence in parallel with epigenomic assays. We examined the viability of enhancer prediction using only a minimal set of epigenomic datasets without direct DNA information. RESULTS: We demonstrate that chromatin datasets are sufficient to identify enhancers genome-wide with high accuracy. By training networks leveraging data from multiple cell types simultaneously, we generated a cell-type invariant enhancer prediction platform that utilized only the patterns of protein binding for inference. We also showed the utility of swarm-based adversarial attacks [adversarial particle swarm optimization (APSO)] to deconvolute trained genomic neural networks for the first time. Critically, unlike saliency mapping or other game-theory based approaches, APSO is completely network-architecture independent and can be applied to any prediction engine to derive the features that drive inference. AVAILABILITY AND IMPLEMENTATION: All software and code for data downloading, processing, enhancer inference, eXplainable AI (XAI), and complete figure generation are publicly available on GitHub at https://github.com/EpiGenomicsCode/ChromEnhancer and Zenodo at https://doi.org/10.5281/zenodo.15652797.

Enhancer Elements, Genetic

DiCARN-DNase: enhancing cell-to-cell Hi-C resolution using dilated cascading ResNet with self-attention and DNase-seq chromatin accessibility data.

MOTIVATION: The spatial organization of chromatin is fundamental to gene regulation and essential for proper cellular function. The Hi-C technique remains the leading method for unraveling 3D genome structures, but the limited availability of high-resolution (HR) Hi-C data poses significant challenges for comprehensive analysis. Deep learning models have been developed to predict HR Hi-C data from low-resolution counterparts. Early Convolutional Neural Network (CNN)-based models improved resolution but struggled with issues like blurring and capturing fine details. In contrast, Generative Adversarial Network (GAN)-based methods encountered difficulties in maintaining diversity and generalization. Additionally, most existing algorithms perform poorly in cross-cell line generalization, where a model trained on one cell type is used to enhance HR data in another cell type. RESULTS: In this work, we propose Dilated Cascading Residual Network (DiCARN) to overcome these challenges and improve Hi-C data resolution. DiCARN leverages dilated convolutions and cascading residuals to capture a broader context while preserving fine-grained genomic interactions. Additionally, we incorporate DNase-seq data into our model, providing a robust framework that demonstrates superior generalizability across cell lines in HR Hi-C data reconstruction. AVAILABILITY AND IMPLEMENTATION: DiCARN is publicly available at https://github.com/OluwadareLab/DiCARN.

Chromatin

SpatialRNA: a Python package for easy application of Graph Neural Network models on single-molecule spatial transcriptomics dataset.

SUMMARY: Image-based spatial transcriptomics (iST) deliver gene expression measurements of RNA transcripts in tissue slices with single-molecule resolution and spatial context preserved. Modern Graph Neural Network (GNN) models are promising methods for capturing the complex molecular and cellular phenotypes in tissues at single-transcript and single-cell levels. A key application of GNNs is the detection of spatial domains or niches, that is, groups of molecules and/or cells that collaboratively work together to produce complex phenotypes. Due to the vast number of detected transcripts in (iST) dataset, applying GNNs on RNA molecule graphs is not trivial. We present a Python package, SpatialRNA, for easy (sub)graph generation from tissue samples and provide comprehensive tutorials for convenient and efficient application of Graph Neural Network models under the PyG framework. This highly scalable tool comprehensively segments tissue into spatial domains, aiding in biological interpretation of iST data and its underlying molecular microenvironments. AVAILABILITY AND IMPLEMENTATION: The SpatialRNA package is freely accessible from online repository https://github.com/ruqianl/spatialrna and can be installed via pip. Comprehensive tutorials, guidance on parameter selection, and complete workflows of case studies are available from the documentation website https://ruqianl.github.io/spatialrna_docs/, and uploaded on Zenodo with a DOI 10.5281/zenodo.17339575.

Neural Networks, Computer

Genome- and peak-informed two-stage framework for scATAC-seq cell type identification.

MOTIVATION: Accurate cell type annotation is essential in scATAC-seq analysis, as it underpins the characterization of cellular heterogeneity, the identification of regulatory elements, and downstream biological discovery. However, current annotation methods still face major challenges. First, although some approaches attempt to integrate genomic sequence information, they typically rely on shallow sequence representations and thus fail to capture the long-range dependencies and regulatory signals encoded in DNA. Second, substantial batch effects introduced by different platforms, sequencing batches, or tissue sources remain insufficiently addressed. Existing models often lack robust distribution alignment and domain generalization capabilities, leading to confounding non-biological variation and reduced annotation accuracy across datasets. RESULTS: To overcome these limitations, we propose seqAlignATAC, a two-stage intra-modality annotation framework that integrates sequence-derived embeddings with domain adaptation. In the first stage, we employ a large-scale pretrained nucleotide language model to extract low-dimensional, biologically informative representations from the genomic sequences of chromatin-accessible peaks. In the second stage, these embeddings are fed into a supervised neural network equipped with an adaptive alignment module to mitigate batch effects and harmonize feature distributions between labeled reference and unlabeled target datasets. Extensive experiments across multiple settings demonstrate that seqAlignATAC achieves competitive accuracy and robustness, effectively leveraging genome-level information while alleviating batch-induced distributional discrepancies. AVAILABILITY AND IMPLEMENTATION: The source code of seqAlignATAC is available at: https://github.com/BioCS-Lab/seqAlignATAC.

Humans