PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “deep sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Longitudinal analysis of high-risk HPV infections reveals within-host viral genome changes over time.

Persistent infection with high-risk (HR)-HPV causes cervical cancer, however, it is unclear why most infections resolve while a minority progress. We deep sequenced the HPV genomes of 1,228 HR-HPV-positive serial samples from 351 women with persistent infections (2-10 serial samples per woman over 1-8 years), including 279 controls and 72 precancer/cancer cases, to assess HR-HPV genome changes during infection and relation to infection outcomes. Seventy-seven percent of persistent infections (45-97% by HPV type) were infections with the same exact viral genome isolate; for HPV16, only 52% were persistent with the same isolate. This may suggest some infections include a type-specific isolate switch or new isolate infection during persistence. We additionally observed within-host change to the HPV genome estimated as gradual changes to intrahost single nucleotide variant (iSNV) frequency, and changes varied by HPV type, with HPV33 infections showing the most iSNV changes. Cases exhibited fewer viral genome changes during infection compared to controls (OR = 0.31, 95% CI = 0.1 - 0.86, p = 0.019), suggesting a more stable and clonal viral genome in cases. By viral gene, E7 had fewer nonsynonymous mutations in the cases compared to controls that cleared within 2 years of infection (p = 0.012), which confirms the importance of E7 conservation and suggests mutations to E7 reduce persistence associated with progression. There was a similar pattern in E4 (p = 0.013), while E5 had more changes in the cases (p = 0.008). A subset of 28 infections had an intervening HPV-negative sample between HPV-positive visits; 93% of these infections had the same exact viral genome isolate in the samples before and after the negative, consistent with subclinical persistence and subsequent re-detection. Our data suggests that HR-HPV type-persistence can include a collection of viral isolates, and viral mutations during infection, particularly in E7, reduce HR-HPV persistence and thus carcinogenic potential.

Humans↗

The use of next-generation sequencing in personalized medicine.

The revolutionary progress in development of next-generation sequencing (NGS) technologies has made it possible to deliver accurate genomic information in a timely manner. Over the past several years, NGS has transformed biomedical and clinical research and found its application in the field of personalized medicine. Here we discuss the rise of personalized medicine and the history of NGS. We discuss current applications and uses of NGS in medicine, including infectious diseases, oncology, genomic medicine, and dermatology. We provide a brief discussion of selected studies where NGS was used to respond to wide variety of questions in biomedical research and clinical medicine. Finally, we discuss the challenges of implementing NGS into routine clinical use.

High-throughput sequencing↗

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500 m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed > 99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071ᵀ (= ATCC 10145ᵀ), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33 Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8 kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~ 22 kb, ~ 17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family↗

Widespread occurrence of a novel division of bacteria identified by 16S rRNA gene sequences originally found in deep marine sediments.

Phylogenetic analysis of 16S rRNA gene sequences from deep marine sediments identified a deeply branching clade, designated candidate division JS1. Primers for PCR amplification of partial 16S rRNA genes that target the JS1 division were developed and used to detect JS1 sequences in DNA extracted from various sedimentary environments, including, for the first time, coastal marine and brackish sediments.

Bacteria↗

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning↗

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis↗

Genome sequence of the deep-sea gamma-proteobacterium Idiomarina loihiensis reveals amino acid fermentation as a source of carbon and energy.

We report the complete genome sequence of the deep-sea gamma-proteobacterium, Idiomarina loihiensis, isolated recently from a hydrothermal vent at 1,300-m depth on the Loihi submarine volcano, Hawaii. The I. loihiensis genome comprises a single chromosome of 2,839,318 base pairs, encoding 2,640 proteins, four rRNA operons, and 56 tRNA genes. A comparison of I. loihiensis to the genomes of other gamma-proteobacteria reveals abundance of amino acid transport and degradation enzymes, but a loss of sugar transport systems and certain enzymes of sugar metabolism. This finding suggests that I. loihiensis relies primarily on amino acid catabolism, rather than on sugar fermentation, for carbon and energy. Enzymes for biosynthesis of purines, pyrimidines, the majority of amino acids, and coenzymes are encoded in the genome, but biosynthetic pathways for Leu, Ile, Val, Thr, and Met are incomplete. Auxotrophy for Val and Thr was confirmed by in vivo experiments. The I. loihiensis genome contains a cluster of 32 genes encoding enzymes for exopolysaccharide and capsular polysaccharide synthesis. It also encodes diverse peptidases, a variety of peptide and amino acid uptake systems, and versatile signal transduction machinery. We propose that the source of amino acids for I. loihiensis growth are the proteinaceous particles present in the deep sea hydrothermal vent waters. I. loihiensis would colonize these particles by using the secreted exopolysaccharide, digest these proteins, and metabolize the resulting peptides and amino acids. In summary, the I. loihiensis genome reveals an integrated mechanism of metabolic adaptation to the constantly changing deep-sea hydrothermal ecosystem.

Amino Acids↗

Recovery and phylogenetic analysis of novel archaeal rRNA sequences from a deep-sea deposit feeder.

In 1992, two independent reports based on small-subunit rRNA gene (SSU rDNA) cloning revealed the presence of novel Archaea among marine bacterioplankton. Here, we report the presence of further novel Archaea SSU rDNA sequences recovered from the midgut contents of a deep-sea marine holothurian. Phylogenetic analyses show that these abyssal Archaea are a paraphyletic component of a highly divergent clade that also includes some planktonic sequences. Our data confirm that this clade is a deep-branching lineage in the tree of life.

Animals↗

Arabidopsis to rice. Applying knowledge from a weed to enhance our understanding of a crop species.

Although Arabidopsis is well established as the premiere model species in plant biology, rice (Oryza sativa) is moving up fast as the second-best model organism. In addition to the availability of large sets of genetic, molecular, and genomic resources, two features make rice attractive as a model species: it represents the taxonomically distinct monocots and is a crop species. Plant structural genomics was pioneered on a genome-scale in Arabidopsis and the lessons learned from these efforts were not lost on rice. Indeed, the sequence and annotation of the rice genome has been greatly accelerated by method improvements made in Arabidopsis. For example, the value of full-length cDNA clones and deep expressed sequence tag resources, obtained in Arabidopsis primarily after release of the complete genome, has been recognized by the rice genomics community. For rice >250,000 expressed sequence tags and 28,000 full-length cDNA sequences are available prior to the completion of the genome sequence. With respect to tools for Arabidopsis functional genomics, deep sequence-tagged lines, inexpensive spotted oligonucleotide arrays, and a near-complete whole genome Affymetrix array are publicly available. The development of similar functional genomics resources for rice is in progress that for the most part has been more streamlined based on lessons learned from Arabidopsis. Genomic resource development has been essential to set the stage for hypothesis-driven research, and Arabidopsis continues to provide paradigms for testing in rice to assess function across taxonomic divisions and in a crop species.

Arabidopsis↗

A multi-modal transformer for cell type-agnostic regulatory predictions.

Sequence-based deep learning models have emerged as powerful tools for deciphering the cis-regulatory grammar of the human genome but cannot generalize to unobserved cellular contexts. Here, we present EpiBERT, a multi-modal transformer that learns generalizable representations of genomic sequence and cell type-specific chromatin accessibility through a masked accessibility-based pre-training objective. Following pre-training, EpiBERT can be fine-tuned for gene expression prediction, achieving accuracy comparable to the sequence-only Enformer model, while also being able to generalize to unobserved cell states. The learned representations are interpretable and useful for predicting chromatin accessibility quantitative trait loci (caQTLs), regulatory motifs, and enhancer-gene links. Our work represents a step toward improving the generalization of sequence-based deep neural networks in regulatory genomics.

Humans↗

Sequence evolution of mitochondrial tRNA genes and deep-branch animal phylogenetics.

Mitochondrial DNA sequences are often used to construct molecular phylogenetic trees among closely related animals. In order to examine the usefulness of mtDNA sequences for deep-branch phylogenetics, genes in previously reported mtDNA sequences were analyzed among several animals that diverged 20-600 million years ago. Unambiguous alignment was achieved for stem-forming regions of mitochondrial tRNA genes by virtue of their conservative secondary structures. Sequences derived from stem parts of the mitochondrial tRNA genes appeared to accumulate much variation linearly for a long period of time: nearly 100 Myr for transition differences and more than 350 Myr for transversion differences. This characteristic could be attributed, in part, to the structural variability of mitochondrial tRNAs, which have fewer restrictions on their tertiary structure than do nonmitochondrial tRNAs. The tRNA sequence data served to reconstruct a well-established phylogeny of the animals with 100% bootstrap probabilities by both maximum parsimony and neighbor-joining methods. By contrast, mitochondrial protein genes coding for cytochrome b and cytochrome oxidase subunit I did not reconstruct the established phylogeny or did so only weakly, although a variety of fractions of the protein gene sequences were subjected to tree-building. This discouraging phylogenetic performance of mitochondrial protein genes, especially with respect to branches originating over 300 Myr ago, was not simply due to high randomness in the data. It may have been due to the relative susceptibility of the protein genes to natural selection as compared with the stem parts of mitochondrial tRNA genes. On the basis of these results, it is proposed that mitochondrial tRNA genes may be useful in resolving deep branches in animal phylogenies with divergences that occurred some hundreds of Myr ago. For this purpose, we designed a set of primers with which mtDNA fragments encompassing clustered tRNA genes were successfully amplified from various vertebrates by the polymerase chain reaction.

Animals↗

Translating functional molecular knowledge into crop-breeding success.

Historical plant breeding, which optimizes phenotypes through selective crossing guided by phenotypic evaluation and molecular markers, is limited by evolutionary constraints that hinder rapid crop improvement. A new paradigm, precision breeding, circumvents these limitations by targeting genetic variants through functional molecular knowledge. To generate this knowledge at scale, sequence-based deep learning leverages high-quality genome sequence data to predict variant effects at base-pair resolution. When linked to agronomically important traits, these predictions enable breeders to prioritize variants for precision selection or editing. Although it is still in the early stages of development, we foresee three key applications for this approach: introgressing genes from distant breeding pools, purging deleterious mutations and designing new plant ideotypes. Looking ahead, refined computational models will facilitate targeted editing and the systematic redesign of complex physiological processes to address emerging breeding goals under shifting environmental conditions.

Crops, Agricultural↗

N-terminal amino acid sequence of the deep-sea tube worm haemoglobin remarkably resembles that of annelid haemoglobin.

The deep-sea giant tube worm Lamellibrachia, belonging to the phylum Vestimentifera, contains two extracellular haemoglobins, an Mr 3,000,000 haemoglobin and an Mr 440,000 haemoglobin. The former has a hexagonal bilayer structure and consists of six polypeptide chains (AI-VI); a study of its haem content shows that not all of the chains contain haem. The Mr 440,000 haemoglobin consists of four haem-containing chains (BI-IV). We isolated most of the chains by reverse-phase chromatography and determined the amino acid sequences of the 21-45 N-terminal residues. Eight chains (AI-IV and BI-IV) showed significant homology with haem-containing chains of annelid giant haemoglobin. The highest homology was found between Lamellibrachia chain AI and Tylorrhynchus chain I; surprisingly, 18 out of the 20 N-terminal residues are identical. On the other hand, chain AV, with an unusual Mr of 32,000, showed a rather different sequence and is likely to be a non-haem chain which might act as a linker protein in the assembly of the haem-containing chains. From these results, we conclude that the tube worm Mr 3,000,000 haemoglobin is highly homologous with annelid haemoglobin.

Amino Acid Sequence↗

Nucleotide sequence and expression of a deep-sea ribulose-1,5-bisphosphate carboxylase gene cloned from a chemoautotrophic bacterial endosymbiont.

The gene coding for ribulose-1,5-bisphosphate carboxylase [RuBisCO; 3-phospho-D-glycerate carboxy-lyase (dimerizing), EC 4.1.1.39] was cloned from a sulfur-oxidizing chemoautotrophic bacterium that resides as an endosymbiont within the gill tissues of Alvinoconcha hessleri, a gastropod inhabiting deep-sea hydrothermal vents. Nucleotide sequence analysis of the cloned fragment demonstrated that the genes encoding the large (RbcL) and small (RbcS) subunits of the symbiont RuBisCO were organized similarly to the RuBisCO operons of free-living photo- and chemoautotrophic prokaryotes. The symbiont rbcL gene shared the highest degree of nucleotide sequence identity with the cyanobacterium Anabaena (69%) while the rbcS nucleotide sequence shared 61% identity with that of the green alga Chlamydomonas reinhardtii. Comparison with a 153-nucleotide partial rbcL sequence from a symbiont of the bivalve Solemya reidi indicated that the two symbiont sequences shared 85% sequence identity at the nucleotide level and 93% at the amino acid level, suggesting a relatively recent common origin. Escherichia coli transformed with a plasmid carrying the RuBisCO operon of the gastropod symbiont in the proper orientation for transcription from the plasmid lac promoter expressed catalytically active RuBisCO. The presence of enzyme activity suggests the proper assembly of the subunits of this deep-sea RuBisCO into the holoenzyme.

Amino Acid Sequence↗

Flexible use of conserved motifs constrains genome access in cell type evolution.

Cell types can be organized into related families, but the regulatory mechanisms that define and maintain these families across deep evolutionary time remain unknown. Here, combining single-nucleus multi-omic sequencing with deep learning to analyse the accessible genomes of two groups of vastly divergent animals including flatworms and vertebrates, we find that hundreds of accessibility-dictating sequence motifs partition into distinct yet conserved sets, or 'vocabularies', each associated with a specific cell type family. However, combinatorial relationships among these motifs preferred by individual cell types are largely species specific. Deep-learning models trained on one species accurately predict family-level chromatin accessibility in distantly related species, albeit frequently rely on different motifs from shared vocabularies to reach convergent predictions. By contrast, models trained on individual cell types within a family lose cross-species predictive power, indicating that the regulatory syntax governing cell type-level identity evolves rapidly. We propose a 'collective maintenance' model in which motif vocabularies defining cell type families are evolutionarily stable, while recombination of these motifs generates cell type-specific regulatory programmes. This suggests that family identity is maintained collectively by large, conserved pools of regulatory factors, analogous to the logic of developmental homology, where character identity persists through network-level conservation despite extensive rewiring.

Journal Article↗

Insights from human/mouse genome comparisons.

Large-scale public genomic sequencing efforts have provided a wealth of vertebrate sequence data poised to provide insights into mammalian biology. These include deep genomic sequence coverage of human, mouse, rat, zebrafish, and two pufferfish ( Fugu rubripes and Tetraodon nigroviridis) (Aparicio et al. 2002; Lander et al. 2001; Venter et al. 2001; Waterston et al. 2002). In addition, a high-priority has been placed on determining the genomic sequence of chimpanzee, dog, cow, frog, and chicken (Boguski 2002). While only recently available, whole genome sequence data have provided the unique opportunity to globally compare complete genome contents. Furthermore, the shared evolutionary ancestry of vertebrate species has allowed the development of comparative genomic approaches to identify ancient conserved sequences with functionality. Accordingly, this review focuses on the initial comparison of available mammalian genomes and describes various insights derived from such analysis.

Animals↗