PubMed HealthSearch

SEARCH · PubMed Health

Results for “Deep sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

High throughput screening of eukaryotic release factor 1 variants to enhance noncanonical amino acid incorporation.

Noncanonical amino acids (ncAAs) enable diversification of protein functions, but the efficiency of genetic code expansion (GCE) in eukaryotes is hindered by competition between suppressor tRNAs and release factors. Prior work has identified eukaryotic release factor 1 (eRF1) mutants that improve ncAA incorporation, suggesting that screens for improved variants may lead to further enhancements. Here, we developed a high-throughput system to screen eRF1 mutants in Saccharomyces cerevisiae where eRF1 mutants are coexpressed on a plasmid alongside genomically encoded, wild-type eRF1. This strategy enabled recovery of live cells expressing eRF1 variants that enhance ncAA incorporation, even with mutants known to severely affect cell viability in the absence of WT eRF1 expression. We prepared and screened a million-member library of randomly mutated eRF1 variants for clones exhibiting improved ncAA integration phenotypes. Deep sequencing revealed a diverse set of enriched mutations across all three major domains of eRF1. Interestingly, several enriched mutations identified here are also found in naturally occurring eRF1 homologs from species that recode canonical stop codons. When eRF1 variants were combined with yeast knockout strains also known to enhance ncAA incorporation, this resulted in further improvements to efficiency, highlighting the complementarity of release factor engineering to other GCE enhancement strategies. This work demonstrates that high-throughput engineering of the eukaryotic translational apparatus is a powerful approach to identify previously unknown solutions for enhancing ncAA incorporation, with implications for elucidating and precisely manipulating the molecular functions of essential translational machinery.

Noncanonical amino acids

Principles of bacterial genome organization, a conformational point of view.

Bacterial chromosomes are large molecules that need to be highly compacted to fit inside the cells. Chromosome compaction must facilitate and maintain key biological processes such as gene expression and DNA transactions (replication, recombination, repair, and segregation). Chromosome and chromatin 3D-organization in bacteria has been a puzzle for decades. Chromosome conformation capture coupled to deep sequencing (Hi-C) in combination with other "omics" approaches has allowed dissection of the structural layers that shape bacterial chromosome organization, from DNA topology to global chromosome architecture. Here we review the latest findings using Hi-C and discuss the main features of bacterial genome folding.

Genome, Bacterial

Central conducting lymphatic anomaly: from bench to bedside.

Central conducting lymphatic anomaly (CCLA) is a complex lymphatic anomaly characterized by abnormalities of the central lymphatics and may present with nonimmune fetal hydrops, chylothorax, chylous ascites, or lymphedema. CCLA has historically been difficult to diagnose and treat; however, recent advances in imaging, such as dynamic contrast magnetic resonance lymphangiography, and in genomics, such as deep sequencing and utilization of cell-free DNA, have improved diagnosis and refined both genotype and phenotype. Furthermore, in vitro and in vivo models have confirmed genetic causes of CCLA, defined the underlying pathogenesis, and facilitated personalized medicine to improve outcomes. Basic, translational, and clinical science are essential for a bedside-to-bench and back approach for CCLA.

Cell-Free Nucleic Acids

A custom library construction method for super-resolution ribosome profiling in Arabidopsis.

BACKGROUND: Ribosome profiling, also known as Ribo-seq, is a powerful technique to study genome-wide mRNA translation. It reveals the precise positions and quantification of ribosomes on mRNAs through deep sequencing of ribosome footprints. We previously optimized the resolution of this technique in plants. However, several key reagents in our original method have been discontinued, and thus, there is an urgent need to establish an alternative protocol. RESULTS: Here we describe a step-by-step protocol that combines our optimized ribosome footprinting in plants with available custom library construction methods established in yeast and bacteria. We tested this protocol in 7-day-old Arabidopsis seedlings and evaluated the quality of the sequencing data regarding ribosome footprint length, mapped genomic features, and the periodic properties corresponding to actively translating ribosomes through open resource bioinformatic tools. We successfully generated high-quality Ribo-seq data comparable with our original method. CONCLUSIONS: We established a custom library construction method for super-resolution Ribo-seq in Arabidopsis. The experimental protocol and bioinformatic pipeline should be readily applicable to other plant tissues and species.

3-nt periodicity

Characterization of the brain virome in human immunodeficiency virus infection and substance use disorder.

Viruses can infect the brain in individuals with and without HIV-infection: however, the brain virome is poorly characterized. Metabolic alterations have been identified which predispose people to substance use disorder (SUD), but whether these could be triggered by viral infection of the brain is unknown. We used a target-enrichment, deep sequencing platform and bioinformatic pipeline named "ViroFind", for the unbiased characterization of DNA and RNA viruses in brain samples obtained from the National Neuro-AIDS Tissue Consortium. We analyzed fresh frozen post-mortem prefrontal cortex from 72 individuals without known viral infection of the brain, including 16 HIV+/SUD+, 20 HIV+/SUD-, 16 HIV-/SUD+, and 20 HIV-/SUD-. The average age was 52.3 y and 62.5% were males. We identified sequences from 26 viruses belonging to 11 viral taxa. These included viruses with and without known pathogenic potential or tropism to the nervous system, with sequence coverage ranging from 0.03 to 99.73% of the viral genomes. In SUD+ people, HIV-infection was associated with a higher total number of viruses, and HIV+/SUD+ compared to HIV-/SUD+ individuals had an increased frequency of Adenovirus (68.8 vs 0%; p<0.001) and Epstein-Barr virus (EBV) (43.8 vs 6.3%; p=0.037) as well as an increase in Torque Teno virus (TTV) burden. Conversely, in HIV+ people, SUD was associated with an increase in frequency of Hepatitis C virus, (25 in HIV+/SUD+ vs 0% in HIV+/SUD-; p=0.031). Finally, HIV+/SUD- compared to HIV-/SUD- individuals had an increased frequency of EBV (50 vs 0%; p<0.001) and an increase in TTV viral burden, but a decreased Adenovirus viral burden. These data demonstrate an unexpectedly high variety in the human brain virome, identifying targets for future research into the impact of these taxa on the central nervous system. ViroFind could become a valuable tool for monitoring viral dynamics in various compartments, monitoring outbreaks, and informing vaccine development.

Male

Longitudinal analysis of high-risk HPV infections reveals within-host viral genome changes over time.

Persistent infection with high-risk (HR)-HPV causes cervical cancer, however, it is unclear why most infections resolve while a minority progress. We deep sequenced the HPV genomes of 1,228 HR-HPV-positive serial samples from 351 women with persistent infections (2-10 serial samples per woman over 1-8 years), including 279 controls and 72 precancer/cancer cases, to assess HR-HPV genome changes during infection and relation to infection outcomes. Seventy-seven percent of persistent infections (45-97% by HPV type) were infections with the same exact viral genome isolate; for HPV16, only 52% were persistent with the same isolate. This may suggest some infections include a type-specific isolate switch or new isolate infection during persistence. We additionally observed within-host change to the HPV genome estimated as gradual changes to intrahost single nucleotide variant (iSNV) frequency, and changes varied by HPV type, with HPV33 infections showing the most iSNV changes. Cases exhibited fewer viral genome changes during infection compared to controls (OR = 0.31, 95% CI&#x2009;=&#x2009;0.1 - 0.86, p&#x2009;=&#x2009;0.019), suggesting a more stable and clonal viral genome in cases. By viral gene, E7 had fewer nonsynonymous mutations in the cases compared to controls that cleared within 2 years of infection (p&#x2009;=&#x2009;0.012), which confirms the importance of E7 conservation and suggests mutations to E7 reduce persistence associated with progression. There was a similar pattern in E4 (p&#x2009;=&#x2009;0.013), while E5 had more changes in the cases (p&#x2009;=&#x2009;0.008). A subset of 28 infections had an intervening HPV-negative sample between HPV-positive visits; 93% of these infections had the same exact viral genome isolate in the samples before and after the negative, consistent with subclinical persistence and subsequent re-detection. Our data suggests that HR-HPV type-persistence can include a collection of viral isolates, and viral mutations during infection, particularly in E7, reduce HR-HPV persistence and thus carcinogenic potential.

Humans

The use of next-generation sequencing in personalized medicine.

The revolutionary progress in development of next-generation sequencing (NGS) technologies has made it possible to deliver accurate genomic information in a timely manner. Over the past several years, NGS has transformed biomedical and clinical research and found its application in the field of personalized medicine. Here we discuss the rise of personalized medicine and the history of NGS. We discuss current applications and uses of NGS in medicine, including infectious diseases, oncology, genomic medicine, and dermatology. We provide a brief discussion of selected studies where NGS was used to respond to wide variety of questions in biomedical research and clinical medicine. Finally, we discuss the challenges of implementing NGS into routine clinical use.

High-throughput sequencing

Genome mining reveals an architecturally expanded pyoluteorin-associated biosynthetic gene cluster and a divergent flavin-dependent halogenase-like sequence in deep-sea Pseudomonas Aeruginosa from the Gulf of Guinea.

BACKGROUND: Marine deep-sea environments harbour microorganisms with extraordinary biosynthetic potential, yet their secondary metabolite repertoires remain largely uncharacterised. RESULTS: This study reports the isolation, phenotypic characterisation, and whole-genome analysis of Pseudomonas aeruginosa strain E1, recovered from deep Atlantic seawater (Gulf of Guinea, ~2500&#xa0;m depth), which exhibits antifungal activity against multidrug-resistant Candida parapsilosis. Three presumptive P. aeruginosa isolates (E1, E17, and E44) showed >&#x2009;99% 16S rRNA gene sequence identity to P. aeruginosa reference sequences, while whole-genome dDDH analysis of strain E1 yielded 95.2% (95% CI: 93.6-96.4%; formula d4) relative to the P. aeruginosa type strain DSM 50071&#x1d40; (=&#x2009;ATCC 10145&#x1d40;), supporting its species-level assignment. Antifungal screening and PCR-based detection of flavin-dependent halogenase genes identified strain E1 as the primary candidate for genomic investigation. Illumina whole-genome sequencing produced a 6.33&#xa0;Mb draft genome assembly (113 contigs, 5862 protein-coding genes, 66.4% GC content). Genome mining with antiSMASH 8.0 identified 27 biosynthetic gene clusters (BGCs) spanning nonribosomal peptide synthetase (NRPS), polyketide synthase (PKS), phenazine, terpene, and metallophore pathways. Region 7.1 of strain E1 harbours a predicted 50.8&#xa0;kb pyoluteorin-associated BGC, comprising 34 genes, substantially larger than its terrestrial counterpart (~&#x2009;22&#xa0;kb, ~&#x2009;17 genes), and featuring nine transport genes and three regulatory elements. Phylogenetic analysis resolved three halogenase genes: ctg7_146 showed 98.7% amino acid identity to PltA, and ctg7_149 showed 99.2% amino acid identity to PltM, supporting their annotation as PltA-like and PltM-like components of the predicted pyoluteorin biosynthetic pathway. Among the characterised reference enzymes included in this analysis, ctg7_143 showed the highest amino acid identity to PltM from P. fluorescens Pf-5. However, the identity remained low at approximately 30.4%, supporting its placement as a divergent FDH-like sequence rather than a close PltM orthologue. CONCLUSION: This study provides the first comprehensive genomic characterisation of a pyoluteorin-BGC-harbouring marine P. aeruginosa strain, demonstrating conservation of the core biosynthetic machinery alongside an expanded transport architecture and a divergent FDH-like sequence that may represent a candidate for future biochemical investigation. These findings expand current knowledge of FDH-like sequence diversity in deep-sea bacteria and support further investigation of Gulf of Guinea microorganisms as a potential source of biosynthetic and enzymatic diversity.

Multigene Family

DeepGeSeq: deep learning library for genomic sequence modeling and analysis.

MOTIVATION: Deep learning methods have demonstrated significant potential in genomics, enabling broad applications such as sequence activity prediction, regulatory rule identification, and variant effect quantification. However, their widespread adoption is often hindered by the steep computational learning curve required for model construction, training, and downstream biological interpretation. Here, we introduce DeepGeSeq, a user-friendly Deep-learning library tailored for Genomic Sequence modeling and analysis. RESULTS: By integrating state-of-the-art architectural modules, DeepGeSeq streamlines the entire deep learning workflow, requiring minimal user input via a simple configuration file and an intuitive agentic skill. We comprehensively validate the efficacy of DeepGeSeq through diverse case studies, encompassing pipeline verification using synthetic datasets, the reproduction and application of established models, and model fine-tuning coupled with biological interpretation on user-defined data. Furthermore, we demonstrate DeepGeSeq's versatility in domain-specific applications, including single-cell ATAC-seq modeling for cell-type clustering, and MPRA data modeling coupled with in silico saturation mutagenesis to dissect cis-regulatory elements. Ultimately, DeepGeSeq bridges the gap between computational complexity and biological discovery, providing an accessible resource that facilitates the development and broad application of deep learning methods in genomics research. AVAILABILITY AND IMPLEMENTATION: https://github.com/JiaqiLi1024/DeepGeSeq.

Deep Learning

Deep generative models in biological sequence and structure analysis and design.

Deep generative models have transformed biological sequence modeling from predictive analysis toward increasingly controllable design. Early biological applications of Variational Autoencoders (VAEs) and Generative Adversarial Networks (GANs) established latent representation learning and sequence synthesis, while recent advances in transformer-based language models, discrete diffusion, flow-matching, and multimodal generative frameworks have substantially expanded the scope of biological design. This review examines generative models for DNA, RNA, and protein sequence design, emphasizing how different model classes represent biological constraints, operate over discrete and continuous spaces, and integrate sequence, structure, and function. We compare VAEs, GANs, autoregressive and masked language models, diffusion models, and flow-based approaches across genomics, transcriptomics, and proteomics, with particular attention to controllability, long-range dependency modeling, structural grounding, generalization, and experimental utility. We further examine evaluation strategies, out-of-distribution generalization, and closed-loop design-build-test-learn workflows that connect in silico generation with empirical validation. We distinguish fundamental modality-dependent constraints including sequence discreteness, context length, structural coupling, and physical or thermodynamic requirements from architecture-dependent advantages that reflect the current state of the field. Current studies suggest that long-context models are particularly useful for genome-scale representation and sequence modeling, whereas structure-aware diffusion, flow-based, and inverse-folding approaches provide better frameworks for geometry-constrained RNA and protein design. This perspective provides a critical framework for understanding the present capabilities, limitations, and convergence of generative approaches toward reliable and experimentally grounded biological design.

Biological sequence analysis

A multi-modal transformer for cell type-agnostic regulatory predictions.

Sequence-based deep learning models have emerged as powerful tools for deciphering the cis-regulatory grammar of the human genome but cannot generalize to unobserved cellular contexts. Here, we present EpiBERT, a multi-modal transformer that learns generalizable representations of genomic sequence and cell type-specific chromatin accessibility through a masked accessibility-based pre-training objective. Following pre-training, EpiBERT can be fine-tuned for gene expression prediction, achieving accuracy comparable to the sequence-only Enformer model, while also being able to generalize to unobserved cell states. The learned representations are interpretable and useful for predicting chromatin accessibility quantitative trait loci (caQTLs), regulatory motifs, and enhancer-gene links. Our work represents a step toward improving the generalization of sequence-based deep neural networks in regulatory genomics.

Humans

Translating functional molecular knowledge into crop-breeding success.

Historical plant breeding, which optimizes phenotypes through selective crossing guided by phenotypic evaluation and molecular markers, is limited by evolutionary constraints that hinder rapid crop improvement. A new paradigm, precision breeding, circumvents these limitations by targeting genetic variants through functional molecular knowledge. To generate this knowledge at scale, sequence-based deep learning leverages high-quality genome sequence data to predict variant effects at base-pair resolution. When linked to agronomically important traits, these predictions enable breeders to prioritize variants for precision selection or editing. Although it is still in the early stages of development, we foresee three key applications for this approach: introgressing genes from distant breeding pools, purging deleterious mutations and designing new plant ideotypes. Looking ahead, refined computational models will facilitate targeted editing and the systematic redesign of complex physiological processes to address emerging breeding goals under shifting environmental conditions.

Crops, Agricultural

N-terminal amino acid sequence of the deep-sea tube worm haemoglobin remarkably resembles that of annelid haemoglobin.

The deep-sea giant tube worm Lamellibrachia, belonging to the phylum Vestimentifera, contains two extracellular haemoglobins, an Mr 3,000,000 haemoglobin and an Mr 440,000 haemoglobin. The former has a hexagonal bilayer structure and consists of six polypeptide chains (AI-VI); a study of its haem content shows that not all of the chains contain haem. The Mr 440,000 haemoglobin consists of four haem-containing chains (BI-IV). We isolated most of the chains by reverse-phase chromatography and determined the amino acid sequences of the 21-45 N-terminal residues. Eight chains (AI-IV and BI-IV) showed significant homology with haem-containing chains of annelid giant haemoglobin. The highest homology was found between Lamellibrachia chain AI and Tylorrhynchus chain I; surprisingly, 18 out of the 20 N-terminal residues are identical. On the other hand, chain AV, with an unusual Mr of 32,000, showed a rather different sequence and is likely to be a non-haem chain which might act as a linker protein in the assembly of the haem-containing chains. From these results, we conclude that the tube worm Mr 3,000,000 haemoglobin is highly homologous with annelid haemoglobin.

Amino Acid Sequence

Nucleotide sequence and expression of a deep-sea ribulose-1,5-bisphosphate carboxylase gene cloned from a chemoautotrophic bacterial endosymbiont.

The gene coding for ribulose-1,5-bisphosphate carboxylase [RuBisCO; 3-phospho-D-glycerate carboxy-lyase (dimerizing), EC 4.1.1.39] was cloned from a sulfur-oxidizing chemoautotrophic bacterium that resides as an endosymbiont within the gill tissues of Alvinoconcha hessleri, a gastropod inhabiting deep-sea hydrothermal vents. Nucleotide sequence analysis of the cloned fragment demonstrated that the genes encoding the large (RbcL) and small (RbcS) subunits of the symbiont RuBisCO were organized similarly to the RuBisCO operons of free-living photo- and chemoautotrophic prokaryotes. The symbiont rbcL gene shared the highest degree of nucleotide sequence identity with the cyanobacterium Anabaena (69%) while the rbcS nucleotide sequence shared 61% identity with that of the green alga Chlamydomonas reinhardtii. Comparison with a 153-nucleotide partial rbcL sequence from a symbiont of the bivalve Solemya reidi indicated that the two symbiont sequences shared 85% sequence identity at the nucleotide level and 93% at the amino acid level, suggesting a relatively recent common origin. Escherichia coli transformed with a plasmid carrying the RuBisCO operon of the gastropod symbiont in the proper orientation for transcription from the plasmid lac promoter expressed catalytically active RuBisCO. The presence of enzyme activity suggests the proper assembly of the subunits of this deep-sea RuBisCO into the holoenzyme.

Amino Acid Sequence

Flexible use of conserved motifs constrains genome access in cell type evolution.

Cell types can be organized into related families, but the regulatory mechanisms that define and maintain these families across deep evolutionary time remain unknown. Here, combining single-nucleus multi-omic sequencing with deep learning to analyse the accessible genomes of two groups of vastly divergent animals including flatworms and vertebrates, we find that hundreds of accessibility-dictating sequence motifs partition into distinct yet conserved sets, or 'vocabularies', each associated with a specific cell type family. However, combinatorial relationships among these motifs preferred by individual cell types are largely species specific. Deep-learning models trained on one species accurately predict family-level chromatin accessibility in distantly related species, albeit frequently rely on different motifs from shared vocabularies to reach convergent predictions. By contrast, models trained on individual cell types within a family lose cross-species predictive power, indicating that the regulatory syntax governing cell type-level identity evolves rapidly. We propose a 'collective maintenance' model in which motif vocabularies defining cell type families are evolutionarily stable, while recombination of these motifs generates cell type-specific regulatory programmes. This suggests that family identity is maintained collectively by large, conserved pools of regulatory factors, analogous to the logic of developmental homology, where character identity persists through network-level conservation despite extensive rewiring.

Journal Article

The Use of Next-Generation Sequencing in Personalized Medicine.

The revolutionary progress in development of next-generation sequencing (NGS) technologies has made it possible to deliver accurate genomic information in a timely manner. Over the past several years, NGS has transformed biomedical and clinical research and found its application in the field of personalized medicine. Here we discuss the rise of personalized medicine and the history of NGS. We discuss current applications and uses of NGS in medicine, including infectious diseases, oncology, genomic medicine, and dermatology. We provide a brief discussion of selected studies where NGS was used to respond to wide variety of questions in biomedical research and clinical medicine. Finally, we discuss the challenges of implementing NGS into routine clinical use.

Humans

Identification of Freezing-Responsive microRNAs and Their Targets in Chinese Jujube by Small RNA and Degradome Sequencing.

The jujube tree fruit remains a primary fruit in northern China, yet its geographical distribution and yield are significantly constrained by freezing stress during winter. Numerous studies have highlighted the pivotal regulatory function of microRNAs (miRNAs) in plant responses to low-temperature stress. Nevertheless, the specific miRNAs involved in the response to low temperatures and their associated gene networks in Ziziphus jujuba Mill are not well understood. In this investigation, we utilized high-throughput sequencing to analyze small RNA libraries from branches subjected to temperatures of 4 &#xb0;C and -30 &#xb0;C. Our analysis identified a total of 342 miRNAs, comprising 123 known miRNAs and 219 novel miRNAs. The differential expression analysis revealed that under low-temperature conditions, 177 miRNAs underwent significant changes. Among them, specific upregulation of miR319 in the less cold-resistant variety and miR6483 in sensitive variety was observed. By employing degradome sequencing, we identified a total of 1551 target genes corresponding to 3059 unique miRNA target interaction pairs involving 299 miRNAs. Functional analysis using Gene Ontology (GO) and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways indicated that these target genes are primarily associated with transcriptional regulation, metabolic pathways, and genetic information processing. Through a comprehensive analysis, we pinpointed 11 genes corresponding to 9 miRNAs that are implicated in jujube tree cold stress, and 7 target genes of 7 miRNAs were confirmed by 5'-RACE analysis. These miRNAs are likely to exert crucial regulatory functions in the context of jujube tree cold stress. This study is the first to systematically identify miRNAs and their target genes in the response of Ziziphus jujuba Mill to low-temperature stress, which provides important resources for in-depth analysis of the molecular mechanism of jujube tree cold resistance and for cold-resistant breeding.

Ziziphus

Microbial diversity and metabolic pathways linked to benzene degradation in petrochemical-polluted groundwater.

The rapid advance in shotgun metagenome sequencing has enabled us to identify uncultivated functional microorganisms in polluted environments. While aerobic petrochemical-degrading pathways have been extensively studied, the anaerobic mechanisms remain less explored. Here, we conducted a study at a petrochemical-polluted groundwater site in Henan Province, Central China. A total of twelve groundwater monitoring wells were installed to collect groundwater samples. Benzene appeared to be the predominant pollutant, detected in 10 out of 12 samples, with concentrations ranging from 1.4&#xa0;&#x3bc;g/L to 5,280&#xa0;&#x3bc;g/L. Due to the low aquifer permeability, pollutant migration occurred slowly, resulting in relatively low benzene concentrations downstream within the heavily polluted area. Deep metagenome sequencing revealed Proteobacteria as the dominant phylum, accounting for over 63&#xa0;% of total abundances. Microbial &#x3b1;-diversity was low in heavily polluted samples, with community compositions substantially differing from those in lightly polluted samples. dmpK encoding the phenol/toluene 2-monooxygenase was detected across all samples, while the dioxygenase bedC1 was not detected, suggesting that aerobic benzene degradation might occur through monooxygenation. Sequence assembly and binning yielded 350 high-quality metagenome-assembled genomes (MAGs), with 30 MAGs harboring functional genes associated with aerobic or anaerobic benzene degradation. About 80&#xa0;% of MAGs harboring functional genes associated with anaerobic benzene degradation remained taxonomically unclassified at the genus level, suggesting that our current database coverage of anaerobic benzene-degrading microorganisms is very limited. Furthermore, two genes integral to anaerobic benzene metabolism, i.e, benzoyl-CoA reductase (bamB) and glutaryl-CoA dehydrogenase (acd), were not annotated by metagenome functional analyses but were identified within the MAGs, signifying the importance of integrating both contig-based and MAG-based approaches. Together, our efforts of functional annotation and metagenome binning generate a robust blueprint of microbial functional potentials in petrochemical-polluted groundwater, which is crucial for designing proficient bioremediation strategies.

Groundwater