PubMed HealthSearch

SEARCH · PubMed Health

Results for “DNA sequencing”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Embed-Search-Align: DNA sequence alignment using Transformer models.

MOTIVATION: DNA sequence alignment, an important genomic task, involves assigning short DNA reads to the most probable locations on an extensive reference genome. Conventional methods tackle this challenge in two steps: genome indexing followed by efficient search to locate likely positions for given reads. Building on the success of Large Language Models in encoding text into embeddings, where the distance metric captures semantic similarity, recent efforts have encoded DNA sequences into vectors using Transformers and have shown promising results in tasks involving classification of short DNA sequences. Performance at sequence classification tasks does not, however, guarantee sequence alignment, where it is necessary to conduct a genome-wide search to align every read successfully, a significantly longer-range task by comparison. RESULTS: We bridge this gap by developing a "Embed-Search-Align" (ESA) framework, where a novel Reference-Free DNA Embedding (RDE) Transformer model generates vector embeddings of reads and fragments of the reference in a shared vector space; read-fragment distance metric is then used as a surrogate for sequence similarity. ESA introduces: (i) Contrastive loss for self-supervised training of DNA sequence representations, facilitating rich reference-free, sequence-level embeddings, and (ii) a DNA vector store to enable search across fragments on a global scale. RDE is 99% accurate when aligning 250-length reads onto a human reference genome of 3 gigabases (single-haploid), rivaling conventional algorithmic sequence alignment methods such as Bowtie and BWA-Mem. RDE far exceeds the performance of six recent DNA-Transformer model baselines such as Nucleotide Transformer, Hyena-DNA, and shows task transfer across chromosomes and species. AVAILABILITY AND IMPLEMENTATION: Please see https://anonymous.4open.science/r/dna2vec-7E4E/readme.md.

Sequence Analysis, DNA

Use of 3D chaos game representation to quantify DNA sequence similarity with applications for hierarchical clustering.

A 3D chaos game is shown to be a useful way for encoding DNA sequences. Since matching subsequences in DNA converge in space in 3D chaos game encoding, a DNA sequence's 3D chaos game representation can be used to compare DNA sequences without prior alignment and without truncating or padding any of the sequences. Two proposed methods inspired by shape-similarity comparison techniques show that this form of encoding can perform as well as alignment-based techniques for building phylogenetic trees. The first method uses the volume overlap of intersecting spheres and the second uses shape signatures by summarizing the coordinates, oriented angles, and oriented distances of the 3D chaos game trajectory. The methods are tested using: (1) the first exon of the beta-globin gene for 11 species, (2) mitochondrial DNA from four groups of primates, and (3) a set of synthetic DNA sequences. Simulations show that the proposed methods produce distances that reflect the number of mutation events; additionally, on average, distances resulting from deletion mutations are comparable to those produced by substitution mutations.

Animals

esloco: simulation-based estimation of local coverage in long-read DNA sequencing.

SUMMARY: Long-read DNA sequencing is increasingly applied for whole-genome studies, yet experimental planning often lacks reliable estimates of target region coverage, leading to costly and time-consuming pilot studies and replicates. We present esloco, a Monte Carlo-based simulation framework for estimating local coverage in long-read sequencing experiments, including scenarios with unknown target regions (e.g. viral integration, CRISPR-Cas9) or PCR-free designs (e.g. base modifications). By modeling coverage as a function of sequencing depth and read length distribution, esloco enables informed predictions of local sequencing outcomes. Benchmarking across a 45-gene panel demonstrated close agreement with empirical data, underscoring the framework's reliability. AVAILABILITY AND IMPLEMENTATION: esloco is a Python package available on PyPI (https://pypi.org/project/esloco/), GitHub (https://github.com/aweich/esloco), and Zenodo (https://doi.org/10.5281/zenodo.17776161).

Sequence Analysis, DNA

Prediction and functional interpretation of inter-chromosomal genome architecture from DNA sequence with TwinC.

Three-dimensional nuclear DNA architecture comprises well-studied intra-chromosomal (cis) folding and less characterized inter-chromosomal (trans) interfaces. Current predictive models of 3D genome folding can effectively infer pairwise cis-chromatin interactions from the primary DNA sequence but generally ignore trans contacts. There is an unmet need for robust models of trans-genome organization that provide insights into their underlying principles and functional relevance. We present TwinC, an interpretable convolutional neural network model that reliably predicts trans contacts measurable through proximity ligation-dependent (in situ and intact Hi-C) and independent (DNA SPRITE) genome-wide chromatin conformation assays. . TwinC uses a paired sequence design from replicate Hi-C experiments to learn single base pair relevance in trans interactions across two stretches of DNA. The method achieves high predictive accuracy (AUROC=0.80) on a cross-chromosomal test set from in situ and intact Hi-C experiments in heart tissue. Furthermore, we train TwinC using in situ Hi-C data from the widely used GM12878 cell line and validate its performance with orthogonal DNA SPRITE assay in the same cell type. Mechanistically, the neural network learns the importance of compartments, chromatin accessibility, clustered transcription factor binding and G-quadruplexes in forming trans contacts. In summary, TwinC models and interprets trans genome architecture, shedding light on this poorly understood aspect of gene regulation.

Journal Article

Evaluating the analytical validity of circulating tumor DNA sequencing assays for precision oncology.

Circulating tumor DNA (ctDNA) sequencing is being rapidly adopted in precision oncology, but the accuracy, sensitivity and reproducibility of ctDNA assays is poorly understood. Here we report the findings of a multi-site, cross-platform evaluation of the analytical performance of five industry-leading ctDNA assays. We evaluated each stage of the ctDNA sequencing workflow with simulations, synthetic DNA spike-in experiments and proficiency testing on standardized, cell-line-derived reference samples. Above 0.5% variant allele frequency, ctDNA mutations were detected with high sensitivity, precision and reproducibility by all five assays, whereas, below this limit, detection became unreliable and varied widely between assays, especially when input material was limited. Missed mutations (false negatives) were more common than erroneous candidates (false positives), indicating that the reliable sampling of rare ctDNA fragments is the key challenge for ctDNA assays. This comprehensive evaluation of the analytical performance of ctDNA assays serves to inform best practice guidelines and provides a resource for precision oncology.

Circulating Tumor DNA

ELYS associates with distinct DNA sequence environments during post-mitotic nuclear pore reassembly.

Nuclear pore complexes (NPCs) contribute to genome organization and cell identity, yet how post-mitotic NPC assembly is coordinated with chromatin architecture remains unclear. Here, we show that the nucleoporin ELYS preferentially associates with chromatin regions displaying distinct intrinsic DNA sequence features that are not explained by the repressive histone marks examined here. ELYS-bound regions are enriched for AT-rich sequences, whereas ELYS binding at super-enhancer-associated loci shift toward GC-rich sequence composition, revealing distinct sequence environments. These findings indicate that ELYS localization is associated with distinct intrinsic DNA sequence features and suggest a mechanism by which nuclear pore-associated architecture restores transcriptional programs after mitosis.

Journal Article

Accumulation of numerous cellular T-DNA sequences in the genus Diospyros by multiple rounds of natural transformation.

Horizontal gene transfer (HGT) is an important phenomenon in the evolutionary history of plants. Natural transformation by Agrobacterium is a special case of HGT and leads to the insertion of cellular T-DNA (cT-DNA) sequences, for example, in Diospyros lotus. The genus Diospyros contains about 795 species with economically important members, like different types of persimmon (D. kaki, D. lotus, and D. virginiana) and ebony (e.g., D. ebenum). Whole genome sequences (WGS) from D. kaki, D. oleifera, D. lotus, and D. virginiana were investigated for cT-DNAs. These four species belong to one clade and contain 15 different cT-DNAs (DiTA to DiTO). The hexaploid species D. kaki cv. "Xiaoguo-tianshi" contains seven types of cT-DNA (DiTA to DiTG) on 27 of 42 homeologs, adding up to 628 kb of cT-DNA. Five of these seven cT-DNAs are non-fixed, as shown by empty chromosomal insertion sites. The evolutionary history of the Diospyros cT-DNAs was reconstructed using the divergence of their inverted repeats. Insert age varied from 3 to 12 million years. Partial cT-DNA sequences were detected in 35 additional species from five Diospyros clades. Our data highlight the unexpectedly large scale of natural Agrobacterium transformation in Diospyros and demonstrate the necessity of whole genome approaches for studies on the origin and evolution of cT-DNAs.

Diospyros

MCALIGN: stochastic alignment of noncoding DNA sequences based on an evolutionary model of sequence evolution.

A method is described for performing global alignment of noncoding DNA sequences based on an evolutionary model parameterized by the frequency distribution of lengths of insertion/deletion events (indels) and their rate relative to nucleotide substitutions. A stochastic hill-climbing algorithm is used to search for the most probable alignment between a pair of sequences or three sequences of known phylogenetic relationship. The performance of the procedure, parameterized according to the empirical distribution of indel lengths in noncoding DNA of Drosophila species, is investigated by simulation. We show that there is excellent agreement between true and estimated alignments over a wide range of sequence divergences, and that the method outperforms other available alignment methods.

Algorithms

sedimix: a workflow for the analysis of hominin nuclear DNA sequences from sediments.

SUMMARY: Sediment DNA-the recovery of genetic material from archaeological sediments-is an exciting new frontier in ancient DNA research, offering the potential to study individuals at a given archaeological site without destructive sampling. In recent years, several studies have demonstrated the promise of this approach by extracting hominin DNA from prehistoric sediments, including those dating back to the Middle or Late Pleistocene. However, a lack of open-source workflows for analysis of hominin sediment DNA samples poses a challenge for data processing and reproducibility of findings across studies. Here, we introduce a snakemake workflow, sedimix, for processing genomic sequences from archaeological sediment DNA samples to identify hominin sequences and generate relevant summary statistics to assess the reliability of the pipeline. By performing simulations and comparing our results to two published studies with human DNA from ∼25,000 years ago (including shotgun data from a sediment sample and capture data from touch DNA recovered from a deer tooth pendant) we demonstrate that sedimix yields accurate and reliable inferences. sedimix offers a reliable and adaptable framework to aid in the analysis of sediment DNA datasets and improve reproducibility across studies. AVAILABILITY AND IMPLEMENTATION: sedimix is available as an open-source software with the associated code, example data, and user manual with installation instructions available at https://github.com/jierui-cell/sedimix. A permanent archived version of this release is available via Zenodo: https://doi.org/10.5281/zenodo.17244854.

Animals

RadiSeq: a single- and bulk-cell whole-genome DNA sequencing simulator for radiation-damaged cell models.

Objective.To build and validate a simulation framework to perform single-cell and bulk-cell whole genome sequencing simulation of radiation-exposed Monte Carlo (MC) cell models to assist radiation genomics studies.Approach.Sequencing the genomes of radiation-damaged cells can provide useful insight into radiation action for radiobiology research. However, carrying out post-irradiation sequencing experiments can often be challenging, expensive, and time-consuming. Although computational simulations have the potential to provide solutions to these experimental challenges, and aid in designing optimal experiments, the absence of tools currently limits such application. MC toolkits exist to simulate radiation exposures of cell models but there are no tools to simulate single- and bulk-cell sequencing of cell models containing radiation-damaged DNA. Therefore, we aimed to develop a MC simulation framework to address this gap by designing a tool capable of simulating sequencing processes for radiation-damaged cells. Main results.We developed RadiSeq-a multi-threaded whole-genome DNA sequencing simulator written in C++. RadiSeq can be used to simulate Illumina sequencing of radiation-damaged cell models produced by MC simulations. RadiSeq has been validated through comparative analysis, where simulated data were matched against experimentally obtained data, demonstrating reasonable agreement between the two. Additionally, it comes with numerous features designed to closely resemble actual whole-genome sequencing. RadiSeq is also highly customizable with a single input parameter file.Significance.RadiSeq enables the research community to perform complex simulations of radiation-exposed DNA sequencing, supporting the optimization, planning, and validation of costly and time-intensive radiation biology experiments. This framework provides a powerful tool for advancing radiation genomics research.

Monte Carlo Method

Sassy: fuzzy searching DNA sequences using SIMD.

MOTIVATION: Approximate string matching (ASM) is the problem of finding all occurrences of a pattern in a text while allowing up to k errors. Many modern methods use seed-chain-extend, which is fast in practice, but does not guarantee finding all matches with ≤k errors. However, applications such as CRISPR off-target detection require exhaustive results. RESULTS: We introduce Sassy, a library and tool for ASM of short patterns in long texts. Sassy splits the text into four parts that are searched in parallel, and uses bitvectors in the text direction rather than the pattern direction. This has complexity O(k⌈n/W⌉) when searching a random text of length n, where W=256 is the SIMD width, and provides significant speedups for small k. Separately, we allow matches of the pattern to extend beyond the text for an overhang cost of, e.g. α=0.5 per character, to find matches near contig or read ends.Sassy is 4× to 15× faster than Edlib for patterns ≤1000 bp, and can search text with a throughput near 2 Gbp/s. Likewise, Sassy is over 100× faster than parasail. We apply Sassy to CRISPR off-target detection by searching 61 guide sequences in a human genome. Sassy is 100× faster than SWOffinder and only slightly slower (for k≤3) than CHOPOFF, for which building its index takes 20 min. Sassy also scales well to larger k, unlike CHOPOFF whose index took over 10 h to build for k=5. AVAILABILITY AND IMPLEMENTATION: Sassy is available as library and binary at https://github.com/RagnarGrootKoerkamp/sassy, and archived at swh:1:dir:e884758dce5777a441bc2799dc8824e563c5f97b.

Sequence Analysis, DNA

Long-read DNA sequencing resolves a rare case of alloimmune hemolysis mimicking autoimmune hemolysis.

BACKGROUND: Immune hemolytic anemia poses a significant challenge in transfusion medicine, as identification of underlying alloantibodies can be masked by warm and/or cold autoantibodies. This increases the risk of transfusing incompatible blood, which can precipitate or exacerbate hemolysis. Identifying alloantibodies in the presence of autoantibodies remains difficult with standard serologic and genotypic methods, often delaying accurate diagnosis and appropriate transfusion strategies. CASE REPORT: We describe a 63-year-old woman with autoimmune hemolytic anemia who suffered near-fatal hemolysis following transfusion. Despite extensive serologic and genotypic testing, the cause of her hemolytic transfusion reactions remained elusive. Given her clinical course and transfusion history, we hypothesized that her acute hemolytic transfusion reactions could be due to immune sensitization to a high-incidence RBC antigen. Research whole-genome long-read sequencing (LRS) revealed homozygosity for a rare KEL*02N.16 allele, consistent with a rare Ko phenotype, which was validated by Sanger sequencing. Retrospective serologic testing with Ko RBCs further confirmed alloimmunization within the Kell system. CONCLUSION: This case highlights the limitations of conventional serologic and genotypic methods in detecting rare blood group phenotypes, and emphasizes the diagnostic power of long-read sequencing in transfusion medicine. Early molecular testing in complex hemolytic cases can facilitate targeted transfusion strategies, reduce the risk of severe hemolysis, and improve patient outcomes. As sequencing technologies become more accessible, they have the potential to revolutionize blood group typing and alloimmunization risk assessment in clinical practice.

Humans

Characterizing the regulatory logic of transcriptional control at the DNA sequence level by ensembles of thermodynamic models.

MOTIVATION: Understanding how the genome encodes the regulatory logic of transcription is a main challenge of the post-genomic era, and can be overcome with the aid of customized computational tools. RESULTS: We report an automated framework for analyzing an ensemble of fits to data of a thermodynamics-based sequence-level model for transcriptional regulation. The fits are clustered accordingly with their intrinsic regulatory logic. A multiscale analysis enables visualization of quantitative features resulting from the deconvolution of the regulatory profile provided by multiple transcription factors interacting with the locus of a gene. Quantitative experimental data on reporters driven by the whole locus of the even-skipped gene in the blastoderm of Drosophila embryos was used for validating our approach. A few clusters of highly active DNA binding sites within the enhancers collectively modulate even-skipped gene transcription. Analysis of variable enhancers' length shows the importance of bound protein-protein interactions for transcriptional regulation. The interplay between activation and quenching enables function conservation of enhancers despite length variations. AVAILABILITY AND IMPLEMENTATION: The transcription factor level data used for performing the reported study is accessible in the input files in Zenodo and GitHub as well the full code. Additional data from formerly FlyEx database will be available under request.

Thermodynamics

DNA sequencing for microbial surveillance in cystic fibrosis airways: advances, challenges, and clinical translation.

SUMMARYDNA sequencing has revolutionized microbial surveillance in cystic fibrosis (CF), transforming pathogen identification from culture-dependent to total microbial community identification using molecular-based approaches. Techniques such as 16S rRNA gene sequencing have uncovered the complexity of the CF airway microbiome, while shotgun metagenomics, metatranscriptomics, and viromics now provide strain-level, functional, and viral insights beyond bacterial identification. Despite these advances, key technical and logistical challenges remain, including the processing of high-viscosity sputum samples, overwhelming host DNA contamination, managing large data sets, and the integration of complex bioinformatic outputs into clinical workflows. Emerging innovations such as host DNA depletion protocols, targeted enrichment panels, and adaptive sampling on Oxford Nanopore platforms are helping to overcome these barriers, improving microbial recovery and sequencing efficiency. As cystic fibrosis transmembrane conductance regulator (CFTR) modulator therapies are changing the lives of people with cystic fibrosis (pwCF), sequencing offers an unprecedented opportunity to track potential microbial adaptation in response. This review investigates current advances, limitations, and translational opportunities in DNA sequencing for CF airway microbiome surveillance, highlighting how these technologies can help reshape research and clinical microbiology in the post-modulator era.

Cystic Fibrosis

A novel relationship between time offsets in capillary electrophoresis and DNA sequence variations in short tandem repeats.

Next-generation sequencing (NGS) provides increased discriminatory power in forensic DNA analysis due to the detection of isoalleles. Differences in sequences between alleles allow for a second layer of differentiation between DNA contributors beyond the number of short tandem repeat (STR) repeat units. However, because NGS is a more time and resource-intensive analysis than conventional capillary electrophoresis (CE), laboratories may benefit from indicators that suggest NGS is likely to provide added value. This study examined whether CE migration offsets, measured as residuals in the OSIRIS analysis software, can differ significantly among STR isoalleles. Residuals represent the time offset between a sample allele peak and its corresponding allelic ladder peak. Paired CE and NGS data from 95 single source samples were analyzed for CE-based residual differences, as the NGS data provided the sequence information of the corresponding isoalleles. Residual values differed significantly among isoalleles at several STR loci. Statistically significant differences were identified at D16S539 and D3S1358, as well as at specific allele lengths within D12S391, D13S317, and D8S1179. These findings demonstrate that CE residual variation can reflect underlying STR sequence differences between contributors. In practice, residual-based metrics could help laboratories to identify casework reference samples where NGS is likely to provide additional discrimination, without the need for processing outside of a routine CE workflow. Due to the potentially large number of isoalleles, community wide efforts to aggregate CE residual differences versus isoallele sequences may be useful in the validation and implementation of this approach to add value to forensic DNA analyses.

Electrophoresis, Capillary

MRD-2 in the GHSG HD21 trial assessed by a validated circulating tumor DNA sequencing assay.

Beyond cure, major goals in patients with Hodgkin lymphoma (HL) are tailoring treatment to a patient's individual risk for relapse to reduce acute and late toxicities, identifying candidates for early incorporation of novel agents, and making treatment affordable on a global level. Minimal residual disease (MRD) assessment by circulating tumor DNA (ctDNA) sequencing emerged as a promising strategy to achieve these goals; however, previous studies differed in sampling time points, assay validation, and definitions for MRD negativity. Here, we applied LymphoVista, a validated ctDNA sequencing assay for genotyping and MRD monitoring in lymphoma, to samples obtained from the German Hodgkin Study Group (GHSG) HD21 trial after 2 cycles of treatment (MRD-2) using a case-cohort design. Patients with positive MRD-2 result were at higher risk for relapse, progression, or death compared with MRD-2-negative patients (4-year progression-free survival [PFS], 36.7% vs 82.2%; hazard ratio, 5.3; 95% confidence interval, 2.0-13.8; P = .0008). After inverse probability weighting accounting for the number of events in the full reference set, patients with positive and negative MRD-2 results had 4-year PFS rates of 72.2% vs 95.3%. Combining MRD-2 with positron emission tomography after 2 cycles of BrECADD/eBEACOPP (PET-2) can identify patients at very low and patients at very high risk of relapse, progression, or death. In summary, these results suggest that MRD-2 assessment by LymphoVista allows for early outcome prognostication in patients with HL and could be used as a tool to improve treatment guidance on its own or in conjunction with PET-2.

Humans

High-throughput DNA extraction and cost-effective miniaturized metagenome and amplicon library preparation of soil samples for DNA sequencing.

Reductions in sequencing costs have enabled widespread use of shotgun metagenomics and amplicon sequencing, which have drastically improved our understanding of the microbial world. However, large sequencing projects are now hampered by the cost of library preparation and low sample throughput, comparatively to the actual sequencing costs. Here, we benchmarked three high-throughput DNA extraction methods: ZymoBIOMICS™ 96 MagBead DNA Kit, MP BiomedicalsTM FastDNATM-96 Soil Microbe DNA Kit, and DNeasy® 96 PowerSoil® Pro QIAcube® HT Kit. The DNA extractions were evaluated based on length, quality, quantity, and the observed microbial community across five diverse soil types. DNA extraction of all soil types was successful for all kits, however DNeasy® 96 PowerSoil® Pro QIAcube® HT Kit excelled across all performance parameters. We further used the nanoliter dispensing system I.DOT One to miniaturize Illumina amplicon and metagenomic library preparation volumes by a factor of 5 and 10, respectively, with no significant impact on the observed microbial communities. With these protocols, DNA extraction, metagenomic, or amplicon library preparation for one 96-well plate are approx. 3, 5, and 6 hours, respectively. Furthermore, the miniaturization of amplicon and metagenome library preparation reduces the chemical and plastic costs from 5.0 to 3.6 and 59 to 7.3 USD pr. sample. This enhanced efficiency and cost-effectiveness will enable researchers to undertake studies with greater sample sizes and diversity, thereby providing a richer, more detailed view of microbial communities and their dynamics.

Metagenome

scooby: Modeling multi-modal genomic profiles from DNA sequence at single-cell resolution.

Understanding how regulatory DNA elements shape gene expression across individual cells is a fundamental challenge in genomics. Joint RNA-seq and epigenomic profiling provides opportunities to build unifying models of gene regulation capturing sequence determinants across steps of gene expression. However, current models, developed primarily for bulk omics data, fail to capture the cellular heterogeneity and dynamic processes revealed by single-cell multi-modal technologies. Here, we introduce scooby, the first framework to model scRNA-seq coverage and scATAC-seq insertion profiles along the genome from sequence at single-cell resolution. For this, we leverage the pre-trained multi-omics profile predictor Borzoi as a foundation model, equip it with a cell-specific decoder, and fine-tune its sequence embeddings. Specifically, we condition the decoder on the cell position in a precomputed single-cell embedding resulting in strong generalization capability. Applied to a hematopoiesis dataset, scooby recapitulates cell-specific expression levels of held-out genes, and identifies regulators and their putative target genes through in silico motif deletion. Moreover, accurate variant effect prediction with scooby allows for breaking down bulk eQTL effects into single-cell effects and delineating their impact on chromatin accessibility and gene expression. We anticipate scooby to aid unraveling the complexities of gene regulation at the resolution of individual cells.

Journal Article