PubMed HealthSearch

SEARCH · PubMed Health

Results for “Deep sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

HyLnc: a hybrid deep learning and feature-based approach for long non-coding RNA prediction.

Long non-coding RNAs (lncRNAs) play important roles in gene regulation, development and disease, yet accurate identification of lncRNAs from transcriptomic data remains a major computational challenge. Existing methods often rely either on handcrafted sequence features or deep learning approaches, each with their inherent limitations in capturing the full complexity of RNA sequences. In this study, we proposed HyLnc, a computational framework that integrates transformer-based contextual embeddings with biologically meaningful sequence features for improved lncRNA prediction. A custom BERT-based model was first pre-trained on a large corpus of metazoan RNA sequences using a masked language modelling strategy to learn contextual nucleotide dependencies. The model was subsequently fine-tuned on curated datasets of lncRNAs and protein-coding transcripts and 256-dimensional deep sequence embeddings were extracted. Parallelly, 348 handcrafted features, including ORF characteristics, untranslated region (UTR) properties, nucleotide composition and Fickett scores, were computed. A multi-stage feature selection strategy was applied to identify the most informative features, resulting in optimized hybrid feature sets. Multiple machine learning classifiers were evaluated, with the RF model achieving the best performance. The proposed framework attained an accuracy of 91.30%, F1-score of 91.23% and MCC of 82.60 on an independent validation dataset, outperforming several existing lncRNA prediction tools. Thus, HyLnc demonstrates that integrating deep contextual representations with biologically interpretable features enhances lncRNA prediction. This approach provides a robust and scalable solution for large-scale transcriptome annotation and can be extended to other sequence-based prediction.

RNA, Long Noncoding

Transfer RNA genes from the hyperthermophilic Archaeon, Methanopyrus kandleri.

Genes encoding the Leu (GAG), Ser (UGA), Gln (UUG) and Lys (UUU) tRNAs have been cloned and sequenced from the deep sea hyperthermophilic Archaeon, Methanopyrus kandleri. Sequences conforming to the TATA box element established for methanogen promoters are located upstream of the tRNA(Gln) and tRNA(Lys) genes. All four of the tRNA genes appear to encode the 3' terminal CCA residues of the mature tRNA. These methanogen tRNAs are predicted to contain most, but not all, invariant residues and are characterized by a high level of G + C base pairing, consistent with the 98 degrees C optimum growth temperature of M. kandleri.

Base Sequence

Molecular evolution of the 5'-terminal domain of large-subunit rRNA from lower eukaryotes. A broad phylogeny covering photosynthetic and non-photosynthetic protists.

This paper summarizes the present status of an analysis of protist phylogeny using rapid partial sequencing of 28S rRNA. Data from 12 protistan phyla are now available and have been used to construct a tentative dendrogram based on a distance matrix method. The tree is robust and has considerable internal consistency. The following salient points are observed: a number of flagellate groups (particularly Euglenozoa) emerge very early among eukaryotes, whereas ciliates and dinoflagellates emerge late, suggesting that some characteristics that had been considered as primitive may in fact be derived. Both chlorophytic and chromophytic photosynthetic protists emerge very late in the tree, close to the Metazoa-Metaphyta-Fungi radiation, suggesting relatively late occurrence of the photosynthetic symbiosis. Taxonomic and phylogenetic information is also obtained within a phylum where rRNA of enough species are sequenced. A deep trichotomy is thus observed within the ciliates. The data are discussed with respect to classical protist phylogenies.

Animals

A deep metagenomic atlas of Qinghai-Xizang Plateau lakes reveals their microbial diversity and salinity adaptation mechanisms.

The Qinghai-Xizang Plateau (QXP), harboring the planet's highest density of plateau lakes, offers an exceptional biogeographic environment for studying extremophilic microbial communities and their adaptation to salinity. Through deep metagenomic sequencing, we construct the Qinghai-Xizang Lake Sediment Genome (QXLSG) catalog, a high-resolution genomic catalog comprising 5,866 metagenome-assembled genomes (MAGs), 58.16 million non-redundant protein encoding genes, and 19,008 biosynthetic gene clusters. Notably, 80.78% of the 2,742 species-level MAGs represent undescribed taxa, significantly expanding the known microbial diversity. Salinity emerges as the primary environmental factor influencing microbial community. Functional annotation highlights that the "salt-out" strategy, particularly the uptake of glycine betaine, is the main mechanism for salinity tolerance. This strategy is prevalent in both hypersaline lake communities and the dominant microbial phyla. Overall, this study provides a crucial genetic resource for future bioprospecting and deepens our understanding of the fundamental mechanisms of microbial adaptation to extreme saline environments.

Lakes

Phyllosphere microbiomes in grassland plants harbor a vast reservoir of novel antimicrobial peptides and biosynthetic diversity.

INTRODUCTION: The phyllosphere microorganisms colonizing plant surface harbor capacities to synthesize diverse specialized metabolites that mediate communication and interactions with environment and host. However, most known metabolites are derived from a few culturable microorganisms, and the genomic diversity and biosynthetic potential of the vast majority of bacteria associated with plants remain largely unexplored. OBJECTIVES: Here, we aim to explore the genome architecture, biosynthetic ability, and host specific adaptability of grassland ecosystems, uncovering new perspectives on grassland phyllosphere microbial resources. METHODS: We employed ultra-deep metagenomic sequencing, functional analysis, host-associated characterization, and bioactivity assays to explore the phyllosphere microbiome across 221 grassland plant samples representing 45 families. This approach revealed host preference in biosynthetic gene clusters (BGCs) and validated the antimicrobial efficacy of phyllosphere-derived antimicrobial peptides (AMPs). RESULTS: Grassland plant phyllosphere microbiomes encode diverse BGCs. We identified 885,396 potential AMPs from over 68 million non-redundant gene sequences. Then, we reconstructed hundreds of near-complete genomes from phyllosphere metagenomes, and 32.61 % of reconstructed genomes were identified as unclassified genomes, primarily within Pseudomonadota, Actinomycetota, Bacillota and Bacteroidota phyla. Of the near-complete genomes, 91.97 % of the BGCs and 99.76 % of the identified AMPs were previously uncharacterized. Host phylogenetic analysis revealed functional divergence. Poaceae-associated Pseudomonas genomes contain an average of 28 BGCs, significantly higher than those in Asteraceae-associated genomes (mean = 14.76, P = 0.033). Similarly, Poaceae-associated Pantoea genomes carried an average of 9 BGCs, exhibiting significant enrichment compared to genomes from Asteraceae (mean = 7.13, P = 6.1e-05), Lamiaceae (mean = 7, P = 0.015), Ranunculaceae (mean = 8.22, P = 0.0053), and Rosaceae (mean = 7.75, P = 0.00069). ParaFit analyses further confirmed that host phylogeny significantly structures microbial functional repertoires, with intra-family hosts sharing more KEGG pathways than inter-family hosts. These results suggest that host evolutionary relationships are associated with metabolic specialization in phyllosphere microbiomes. All 13 AMPs synthesized via solid-phase peptide synthesis demonstrated antimicrobial activity, inhibiting the growth of at least one tested bacterial strain. CONCLUSION: This study demonstrates the promise of grassland plant phyllosphere microbiome as a rich source for novel antimicrobial agents.

Antimicrobial Peptides

NovoBoard: A Comprehensive Framework for Evaluating the False Discovery Rate and Accuracy of De Novo Peptide Sequencing.

De novo peptide sequencing is one of the most fundamental research areas in mass spectrometry-based proteomics. Many methods have often been evaluated using a couple of simple metrics that do not fully reflect their overall performance. Moreover, there has not been an established method to estimate the false discovery rate (FDR) of de novo peptide-spectrum matches. Here we propose NovoBoard, a comprehensive framework to evaluate the performance of de novo peptide-sequencing methods. The framework consists of diverse benchmark datasets (including tryptic, nontryptic, immunopeptidomics, and different species) and a standard set of accuracy metrics to evaluate the fragment ions, amino acids, and peptides of the de novo results. More importantly, a new approach is designed to evaluate de novo peptide-sequencing methods on target-decoy spectra and to estimate and validate their FDRs. Our FDR estimation provides valuable information to assess the reliability of new peptides identified by de novo sequencing tools, especially when no ground-truth information is available to evaluate their accuracy. The FDR estimation can also be used to evaluate the capability of de novo peptide sequencing tools to distinguish between de novo peptide-spectrum matches and random matches. Our results thoroughly reveal the strengths and weaknesses of different de novo peptide-sequencing methods and how their performances depend on specific applications and the types of data.

Peptides

A Robust, Self-Digestion-Resistant LysN with Superior Activity and Cleavage Fidelity for Advanced Proteomic Workflows.

LysN is a valuable protease in proteomics because it cleaves peptide bonds N-terminal to lysine, generating peptides with physicochemical properties complementary to those produced by LysC and trypsin. However, the broader adoption of LysN in proteomic workflows has been limited by the lack of commercially available enzymes that combine high activity, low missed-cleavage rates, and sufficient stability under practical sample-processing conditions. Here, we report the recombinant production and proteomic characterization of a self-digestion-resistant and highly active LysN from Shewanella loihica (SL-LysN). Using terminomics, we mapped the mature N- and C-termini of the enzyme and established the primary structure of the active protease. We further developed a high-density fermentation, refolding, and purification workflow to obtain highly purified recombinant SL-LysN. Biochemical and proteomic benchmarking showed that SL-LysN displayed 3.3-fold higher specific activity than commercial LysN and reduced missed cleavages by approximately 80%. Notably, SL-LysN retained high activity in the presence of 8 M urea or 1% SDS and showed strong resistance to autolysis, indicating exceptional robustness for proteomic sample preparation. In complex mammalian proteome digests, SL-LysN achieved >95% cleavage specificity and a missed-cleavage rate of only 5.9%. These features address a long-standing bottleneck in N-terminal proteolysis and establish SL-LysN as a high-performance enzymatic tool for advanced proteomic workflows, including deep protein sequencing, quantitative proteomics, terminomics, de novo sequencing and analyses requiring efficient digestion under denaturing conditions.

Shewanella

The periphery of nuclear speckles defines a spatially and temporally regulated compartment of long-lived intron-retained RNAs that resolves during mitosis.

RNA localization adds a fundamental layer to gene expression by determining when and where translation-ready mRNAs become available, yet how this timing is coordinated with nuclear architecture and cell-cycle progression remains unclear. Here we identify a subnuclear RNA niche at the nuclear speckle periphery that couples intron retention to cell-cycle-timed RNA release. Using compartment-resolved transcriptional inhibition, sequence-based deep learning and single-molecule and super-resolution RNA imaging in human pluripotent stem cells, we define a class of nuclear RNAs with long-lived retained introns that persist for hours and are enriched in transcripts encoding regulators of genome maintenance and mitosis, including centromere and kinetochore assembly, DNA repair and telomere maintenance. Long-lived retained introns exhibit elevated GC content, predicted structural stability and enrichment for nuclear speckle-associated RNA-binding proteins. In interphase, these RNAs localize to a distinct nuclear speckle-peripheral RNA niche in a spatial arrangement conserved across cell types. During mitotic remodelling, they undergo coordinated, kinase-dependent splicing and are released into the cytoplasm of early G1 daughter cells. Together, these findings link cis-encoded intronic features, subnuclear organization and mitotic remodelling to temporal control of RNA fate.

Mitosis

Distinct modes of evolution drive HIV escape from two broadly neutralizing antibodies.

Broadly neutralizing antibodies (bNAbs) show promise for HIV treatment and prevention, but are vulnerable to resistance evolution. Comprehensively understanding in vivo viral escape from individual bNAbs is necessary to design bNAb combinations that will provide durable responses. We characterize viral escape from two such bNAbs, 10-1074 and 3BNC117, using deep, longitudinal sequencing of full length HIV envelope (env) genes from study participants treated with bNAb monotherapy. Improved sequencing depth and computational evolutionary analyses permit us to identify in vivo routes and parallelism underlying HIV escape from each bNAb, providing new insights into this evolutionary process: 10-1074 escape is restricted to a small number of previously documented pathways, but these escape mutations 1) pre-exist in intra-host viral populations before therapy, 2) are not all equally preferred, and 3) emerge with a high degree of genetic parallelism within and across viral populations. In contrast, 3BNC117 escape follows background-specific patterns in which specific escape mutations present in one population rarely emerge or spread in other populations, but often still exhibit parallel evolutionary responses within their host. That bNAbs elicit starkly different in vivo escape profiles depending on their Env target exposes the limitations of generalizing escape patterns across therapies and highlights the substantial challenges in predicting a viral population's bNAb susceptibility from genetic diversity alone.

Journal Article

Intracranial recordings of endogenous ERPs in humans.

Target detection and stimulus omission tasks of the type used to elicit scalp P300 and related potentials were studied in a group of 40 patients in whom intracranial electrodes had been implanted during evaluation for epilepsy surgery. Two distinct task-related intracranial ERP patterns have been identified, one in the medial temporal lobe and the other in the frontal lobe. These patterns overlap in time with each other and with scalp P300. The temporal lobe pattern consists of positive potentials dorsal and posterior to the hippocampus, sharp negative potentials within and medial to the hippocampus, and positive potentials in the vicinity of the amygdala. This 3-part pattern has been observed for counted targets in auditory, somatic, and visual modalities and for counted stimulus omissions with latencies that covary with scalp P300. This pattern is absent or greatly attenuated in ignore tasks when targets were not counted. The frontal pattern consists of a widespread negative-positive-negative sequence at deep sites which in some patients inverts in polarity at superficial sites and on the scalp. This pattern is consistent with a source or sources within the frontal lobe. Differences in shape and onset latency between the frontal and medial temporal lobe ERP patterns indicate that the former are not simply a distant recording of the latter. These data strongly suggest multiple contributions to scalp P300.

Acoustic Stimulation

abCRISPR: deep learning-based design of abasic gRNA sequences for specific CRISPR-Cas9 genome editing.

SUMMARY: CRISPR-Cas9 has become a widely used tool for genome editing. However, its off-target cleavage caused by partial sequence matches with guide RNAs (gRNAs) remains a critical limitation. Recently, abasic gRNAs (ØXØ) have been developed to enhance target specificity, but their effects vary depending on the positional sequence context. Here, we present abCRISPR, a deep neural network (DNN) framework for the rational design of ØXØ sequences with minimized off-target activity. abCRISPR leverages informative few-shot training with paired datasets of abasic and unmodified gRNAs, using high-quality random mismatch target libraries, exhaustively sequenced for mismatched off-target substrates (n = 97583) in in vitro CRISPR-Cas9 cleavage experiments. Predicted off-target activities for both abasic and unmodified gRNAs showed strong correlation with experimental data (r ≥ 0.95, 10-fold cross-validation). Notably, these comprehensive training sets provide robust ground-truth negatives, enabling accurate and sensitive prediction of off-targets. For unmodified gRNAs, abCRISPR (AUC = 0.98) was validated to outperform existing deep learning-based methods (AUC = 0.45-0.68). When applied to the human genome, abCRISPR generated ØXØ sequences, covering 58 875 004 potent CRISPR-targetable sites with improved target specificity. Together, this work provides a comprehensive bioinformatics resource for safe and precise CRISPR-Cas9 genome editing. AVAILABILITY AND IMPLEMENTATION: The source code for abCRISPR and training data are available at https://doi.org/10.5281/zenodo.20398246. abCRISPR results for the human genome are available at http://clip.korea.ac.kr/abCRISPR/.

Deep Learning

In vitro reconstitution of chromatin replication recapitulates symmetric histone recycling.

Symmetric histone recycling is vital for maintaining epigenetic inheritance upon eukaryotic DNA replication. Recent genome-wide studies have uncovered key determinants of this process, but how these factors collectively support parental histone transfer remains incompletely understood. Here, we successfully reconstitute histone recycling with 24 purified proteins and analyze the products digested by Micrococcal nuclease with Repli-pore-seq, the newly developed pipeline combining nanopore sequencing and deep-learning-based classification. As a result, we identify histones symmetrically recycled as tetrasomes or hexasomes on nucleosome-favorable sequences. We also observe the discordance of the recycled position between lagging and leading strands on the GC-rich DNA sequences. Moreover, removal of Pol δ, Pol32, Dpb3/4, Ctf4, Csm3/Tof1, or Mrc1 disrupts the balance of histone recycling between the two daughter strands, whereas removal of Ctf4, Csm3/Tof1, or Mrc1 additionally alters the positions at which histones were recycled. Furthermore, addition of the lagging-strand maturation factors Fen1 and Cdc9 enhances histone recycling to the lagging strand. These findings provide critical insights into the molecular players and mechanisms underlying symmetric histone recycling.

Histones

African populations and the evolution of human mitochondrial DNA.

The proposal that all mitochondrial DNA (mtDNA) types in contemporary humans stem from a common ancestor present in an African population some 200,000 years ago has attracted much attention. To study this proposal further, two hypervariable segments of mtDNA were sequenced from 189 people of diverse geographic origin, including 121 native Africans. Geographic specificity was observed in that identical mtDNA types are shared within but not between populations. A tree relating these mtDNA sequences to one another and to a chimpanzee sequence has many deep branches leading exclusively to African mtDNAs. An African origin for human mtDNA is supported by two statistical tests. With the use of the chimpanzee and human sequences to calibrate the rate of mtDNA evolution, the age of the common human mtDNA ancestor is placed between 166,000 and 249,000 years. These results thus support and extend the African origin hypothesis of human mtDNA evolution.

Africa

DeepES: deep learning-based enzyme screening to identify orphan enzyme genes.

MOTIVATION: Progress in sequencing technology has led to determination of large numbers of protein sequences, and large enzyme databases are now available. Although many computational tools for enzyme annotation were developed, sequence information is unavailable for many enzymes, known as orphan enzymes. These orphan enzymes hinder sequence similarity-based functional annotation, leading gaps in understanding the association between sequences and enzymatic reactions. RESULTS: Therefore, we developed DeepES, a deep learning-based tool for enzyme screening to identify orphan enzyme genes, focusing on biosynthetic gene clusters and reaction class. DeepES uses protein sequences as inputs and evaluates whether the input genes contain biosynthetic gene clusters of interest by integrating the outputs of the binary classifier for each reaction class. The validation results suggested that DeepES can capture functional similarity between protein sequences, and it can be implemented to explore orphan enzyme genes. By applying DeepES to 4744 metagenome-assembled genomes, we identified candidate genes for 236 orphan enzymes, including those involved in short-chain fatty acid production as a characteristic pathway in human gut bacteria. AVAILABILITY AND IMPLEMENTATION: DeepES is available at https://github.com/yamada-lab/DeepES. Model weights and the candidate genes are available at Zenodo (https://doi.org/10.5281/zenodo.11123900).

Deep Learning

Sequential development of connections between striate and extrastriate visual cortical areas in the rat.

In these experiments we have asked whether the projection from the rat's primary visual cortex, area 17, to the extrastriate visual cortical area 18a is formed in a sequence and whether that sequence resembles the pattern of inside-out cortical neurogenesis. For this purpose fluorescent retrograde tracers were injected into area 18a at different postnatal ages (P1, P5, adult). Animals survived until 3-4 weeks of age, after migration is complete and neurons have arrived at their final laminar location. In the ipsilateral cortex, P1 injections retrogradely labeled cells in layers 5 and 6 of area 17. Labeling after P5 injections extended into more superficial layers and included the bottom of layer 2/3 and layers 4-6. After P5, more labeled cells were found at the top of layer 2/3, producing the adult laminar pattern, where the projection originates predominantly from layer 2/3. A similar sequence of laminar labeling was observed in the transcallosal connection of area 18a. This sequence of labeling, deep layers before superficial, resembles the pattern in which cortical neurons are born and indicates that axons arrive at their cortical targets in the order the cells were generated.

Animals

A deep intronic IFT172 variant causing pseudoexon inclusion identified by whole-genome sequencing in nephronophthisis.

Nephronophthisis is an autosomal recessive ciliopathy and a major genetic cause of end-stage kidney disease in children and young adults. Although next-generation sequencing panels have improved diagnostic yield, some patients remain genetically unresolved, partly due to deep intronic variants that disrupt pre-mRNA splicing and are not captured by exon-focused approaches. We report a 13-year-old boy who presented with advanced kidney dysfunction, small renal cysts, and kidney histopathology consistent with nephronophthisis. Targeted gene panel sequencing failed to identify causative pathogenic variants beyond a missense variant of uncertain significance. Whole-genome sequencing subsequently revealed compound heterozygous variants in IFT172 (NM_015662.3): a missense variant (c.4696C > T, p.Arg1566Cys) and a deep intronic variant (c.4915-94A > G). In silico analysis predicted activation of cryptic splice sites leading to inclusion of an 86-bp pseudoexon, which was confirmed by a minigene splicing assay. These findings established a molecular diagnosis of IFT172-related nephronophthisis. To our knowledge, this is the first report demonstrating pseudoexon inclusion in IFT172, thereby expanding its mutational spectrum. Our case underscores the importance of evaluating deep intronic regions using whole-genome sequencing and functional validation in genetically unresolved nephronophthisis.

Humans

SIMS: A deep-learning label transfer tool for single-cell RNA sequencing analysis.

Cell atlases serve as vital references for automating cell labeling in new samples, yet existing classification algorithms struggle with accuracy. Here we introduce SIMS (scalable, interpretable machine learning for single cell), a low-code data-efficient pipeline for single-cell RNA classification. We benchmark SIMS against datasets from different tissues and species. We demonstrate SIMS's efficacy in classifying cells in the brain, achieving high accuracy even with small training sets (<3,500 cells) and across different samples. SIMS accurately predicts neuronal subtypes in the developing brain, shedding light on genetic changes during neuronal differentiation and postmitotic fate refinement. Finally, we apply SIMS to single-cell RNA datasets of cortical organoids to predict cell identities and uncover genetic variations between cell lines. SIMS identifies cell-line differences and misannotated cell lineages in human cortical organoids derived from different pluripotent stem cell lines. Altogether, we show that SIMS is a versatile and robust tool for cell-type classification from single-cell datasets.

Single-Cell Analysis

Refining sequence-to-activity models by increasing model resolution.

Decoding the cis-regulatory syntax that controls gene expression is essential for improving our understanding of cell differentiation and disease. To identify regulatory motifs and their regulatory syntax, deep learning based sequence-to-activity (S2A) models learn transcription factor binding motifs and their combinations from DNA sequence by modeling measured chromatin accessibility. Previously, we developed AI-TAC, a S2A model that predicts chromatin accessibility across various immune cell types in multi-task fashion, effectively decoding the regulatory syntax underlying immune cell differentiation. While ATAC-seq is commonly used to measure regional accessibility, it also provides high-resolution profiles, the distribution of Tn5 insertion sites, that offer additional insights into the precise location and strength of TF binding sites. Here we demonstrate that modeling ATAC-seq profiles alongside accessibility consistently improves predictions of differential chromatin accessibility across cell types. Moreover, we also find that multi-task learning across related immune cell types consistently outperforms single-task models. To understand what additional information bpAITAC learns from ATAC-seq profiles, we systematically compare sequence attributions from models trained with and without ATAC-seq profiles. We identify novel motifs with strong effect sizes that emerge only when profile data is included. Our findings suggest that modeling ATAC-seq at base-pair resolution enables the model to learn a more nuanced and sensitive representation of the cis-regulatory syntax driving immune cell-specific chromatin landscapes.

ATAC-seq