PubMed HealthSearch

SEARCH · PubMed Health

Results for “barcode”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Performance comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of Poliovirus Viral Protein 1 (VP1) amplicons.

Accurate and timely sequencing of poliovirus is critical for global eradication efforts, particularly for molecular epidemiology based on the typing region of the genome, viral protein 1 (VP1). While Oxford Nanopore Technologies (ONT) sequencing has expanded capabilities for poliovirus surveillance, the relative performance of different ONT library preparation methods, including ligation-based (Native Barcoding) and transposase-based (Rapid Barcoding) approaches, has not been systematically evaluated. In this study, we compared rapid barcoding and native barcoding workflows for sequencing VP1 amplicons from 17 type 2 poliovirus-positive samples, each processed in triplicate. Native barcoding generated significantly more sequencing output, producing approximately 2.3-fold greater total read yield than rapid barcoding, and demonstrated higher run-to-run reproducibility (R2 = 0.979-0.998 vs. 0.847-0.929, respectively; p&#x202f;<&#x202f;0.001). In addition, native barcoding generated 80% of the total yield achieved by rapid barcoding within approximately 7&#x202f;h, whereas rapid barcoding required approximately 40&#x202f;h to reach the same output. Despite these differences, both methods produced identical VP1 consensus sequences across all samples, with comparable read quality (median per-base Q-scores of approximately Q17-Q18). Rapid barcoding provided substantial practical advantages, reducing hands-on library preparation time (55 vs. 200&#x202f;min) and per-sample cost ($12.82 vs. $16.54), while simplifying workflow and reducing technical complexity. These findings indicate that sequencing yield may not be a determinant of downstream analytical outcomes for poliovirus VP1 ONT sequencing. Rapid barcoding therefore represents a cost-effective and efficient approach for routine poliovirus surveillance, whereas native barcoding remains advantageous in applications requiring rapid data generation or maximal sequencing depth.

Poliovirus

A hybrid and cost-efficient barcoding strategy for full-length 16S rRNA gene nanopore sequencing of environmental samples.

BACKGROUND: Accurate species-level identification of bacteria in complex environmental samples is essential for applications in biotechnology, ecological monitoring, and clinical diagnostics. Short-read platforms such as Illumina frequently truncate the 16S rRNA gene, limiting taxonomic resolution. In this work, we applied Oxford Nanopore Technology (ONT) long-read sequencing to full-length 16S rRNA amplicon in samples from natural soil amended with lignocellulosic biomass and a simplified microbial community derived from cultures grown on selective and differential carboxymethyl cellulose (CMC)-based substrates, with the aim to evaluate the difference in performance between a real, complex community and a less complex system. To reduce consumable costs, we substituted the standard ONT Barcoding kits with an in-house hybrid barcoding workflow. Specifically, PacBio PCR-based barcoding protocol was used for sample indexing, followed by library preparation using the ONT Ligation Sequencing Kit. This simplified approach retained compatibility with MinION and Flongle flow cells and supported accurate downstream demultiplexing while lowering barcode costs substantially. Additionally, a new bioinformatic workflow tailored to ONT data was implemented. RESULTS: Overall, the hybrid protocol significantly reduced per-sample barcoding costs while preserving high sequencing quality and throughput. The sequencing run yielded over 5 Gb of quality-filtered data (Q-score &#x2265; 10). Furthermore, the new bioinformatic workflow allowed taxonomic assignment at the species level for 49.38% of annotated taxa, compared to just 4.59% using Illumina NovaSeq sequencing of the V3-V4 region. ONT also recovered 2.3 times more genera and 1.3 times more families. Although 16S rRNA gene sequencing often cannot distinguish between closely related species, particularly within taxonomically complex groups, in this work, full-length reads substantially improved both taxonomic resolution and database matching. CONCLUSIONS: These results show that full-length 16S rRNA sequencing with ONT, paired with a low-cost barcoding strategy, enhanced taxonomic resolution compared to short-read workflows. This approach also offers a scalable and cost-effective option for high-resolution microbiome profiling in research and applied settings.

RNA, Ribosomal, 16S

Is There a Fly in My Soup? To What Extent Do Metabarcoding and Individual Barcoding Tell the Same Story?

Metabarcoding has become the method of choice for characterizing complex arthropod communities. The extent to which metabarcoded bulk samples will recover the same community composition as individual sequencing of all individuals in the sample remains poorly quantified. Biases such as unequal extraction of DNA from different taxa, primer mismatches and non-random PCR may cause the selective drop-out of species from metabarcoding data. At the same time, DNA metabarcoding may reveal arthropod taxa present not as individuals, but as DNA residues on the surface or in the gut of insects. To quantify the consistency in sample contents established by different means, we metabarcoded 45 bulk insect samples, then extracted all arthropods and sequenced them individually. Metabarcoding targeted 418&#x2009;bp at the 3' end of the Folmer barcoding region, while individual barcodes captured the entire 658&#x2009;bp Folmer region. The metabarcoding workflow, including PCR amplification, sequencing and bioinformatics, was performed in three replicates from three separate lysate aliquots per sample. For the main analyses, sequences were assigned to Barcode Index Numbers (BINs) as identical taxonomic categories across data types, thereby allowing the detection of even rare but biologically true taxa. Since such reference-based validation will be unavailable to any researcher dealing with metabarcoding data alone, we validated our key findings through an alternative workflow, i.e., de novo clustering of sequences. We found that metabarcoding is replicable, as different replicates of the same sample recover similar species richness and composition. Individual barcoding and metabarcoding provide similar impressions of relative differences in community structure: species-rich vs. species-poor samples rank similarly among data types (Spearman's &#x2374;&#x2009;=&#x2009;0.88-0.99) as do differences in relative dissimilarity between sample pairs (Spearman's &#x2374;&#x2009;=&#x2009;0.55-0.90). Dissimilarity between data types varies with BIN richness in the sample, but this relationship reflects nestedness rather than turnover: metabarcoding recovers the same set of core species as individual barcoding but adds hundreds of species on top. Any BIN recovered as an individual occurred with high probability in the metabarcoding data, and any BIN found in high read abundances by metabarcoding was likely found as an individual (p&#x2009;>&#x2009;0.8). In terms of abundances, the number of individual insects per BIN was well predicted by the number of metabarcoding reads (R2&#x2009;>&#x2009;0.68 for a model including taxonomy as a random effect). Our analysis suggests that metabarcoding data will be informative of the sample contents in terms of arthropod species richness, composition and taxon-specific abundances. Taxa recovered in low copy numbers in metabarcoding sequence data will likely represent DNA left as residues from past biotic interactions. Barring sequencing errors, both types of data yield biologically relevant insights into the taxa present in the source community.

Animals

Paired Single-Cell Transcriptome and DNA Barcode Detection in Zebrafish Using ScarTrace.

ScarTrace is a CRISPR/Cas9-based genetic lineage tracing method that allows for uniquely barcoding the DNA of single cells at a target GFP sequence during developing zebrafish embryos. Single cells from barcoded adult zebrafish can be isolated from various tissues (e.g., marrow, brain, eyes, fins), and their transcriptome and barcode sequences are captured by single-cell cDNA amplification and genomic DNA nested PCR, respectively. Computationally, cell type and barcode identification permit clone tracing and lineage tree reconstruction of tissues to unravel fate decisions during embryogenesis.

Animals

DNA barcoding in diverse educational settings: five case studies.

Despite 250 years of modern taxonomy, there remains a large biodiversity knowledge gap. Most species remain unknown to science. DNA barcoding can help address this gap and has been used in a variety of educational contexts to incorporate original research into school curricula and informal education programmes. A growing body of evidence suggests that actively conducting research increases student engagement and retention in science. We describe case studies in five different educational settings in Canada and the USA: a programme for primary and secondary school students (ages 5-18), a year-long professional development programme for secondary school teachers, projects embedding this research into courses in a post-secondary 2-year institution and a degree-granting university, and a citizen science project. We argue that these projects are successful because the scientific content is authentic and compelling, DNA barcoding is conceptually and technically straightforward, the workflow is adaptable to a variety of situations, and online tools exist that allow participants to contribute high-quality data to the international research effort. Evidence of success includes the broad adoption of these programmes and assessment results demonstrating that participants are gaining both knowledge and confidence. There are exciting opportunities for coordination among educational projects in the future.This article is part of the themed issue 'From DNA barcodes to biomes'.

Biodiversity

Barcoded mutant library enables high-throughput functional genomics in a filamentous fungus.

Advances in sequencing technology enabling rapid and inexpensive whole-genome sequencing highlight how few genes are functionally characterized. This problem is particularly acute in filamentous fungi, where even in the best studied organisms upward of half of genes are poorly characterized or unannotated. High-throughput tools to identify gene function exist for single-celled organisms, like yeast and bacteria. However, filamentous fungi present challenges to high-throughput gene characterization, including low transformation efficiency and multinucleate cells. Filamentous fungi are critical components of nutrient cycling in ecosystems, form symbioses with plants that improve nutrient uptake, and are devastating human, plant, and animal pathogens causing millions of deaths and substantial crop loss each year. Thus, it is critical to overcome challenges to rapid gene characterization in filamentous fungi. We generated a library of hundreds of millions of uniquely barcoded plasmids containing a broad host-range drug resistance marker for ectopic insertion into filamentous fungal genomes by Agrobacterium tumefaciens. We then optimized A. tumefaciens mediated transformation of the biocontrol agent Trichoderma atroviride and made an insertional mutagenesis library containing 83,311 barcoded insertions, disrupting 5,331 of 11,863 predicted genes. This library enables high-throughput screens to rapidly connect genotype to phenotype. Quantifying relative barcode abundance in the pooled library before and after exposure to experimental conditions identified candidate genes and recovered known pathway components in amino acid biosynthetic, fructose utilization, and xylose utilization pathways. This resource establishes a scalable platform for high-throughput functional genomics in filamentous fungi, enabling investigations of fungal biology to improve medical outcomes, biotechnology, and sustainable agriculture.

Genomics

Comparative chloroplast genomics of six Bupleurum (Apiaceae) accessions: candidate barcodes, phylogeny based on available plastomes, and candidate RNA-editing sites.

INTRODUCTION: Bupleurum L. (Apiaceae), a taxonomically intricate genus of about 190 species and a source of Radix Bupleuri (Chai Hu), is difficult to discriminate because of convergent morphology, infraspecific variation, and limited genomic sampling. This study aimed to characterize plastome variation, identify and validate candidate molecular markers, reconstruct plastid phylogenetic relationships, and assess candidate plastid RNA-editing sites in Bupleurum. METHODS: We assembled six plastomes from subgenus Bupleurum, screened 51 Bupleurum plastomes for diagnostic loci, reconstructed whole-plastome and partitioned protein-coding-sequence phylogenies, and predicted plastid C-to-U RNA-editing candidates across the six newly assembled plastomes using a PREP-Cp-compatible workflow. Candidate barcode performance was evaluated against the reference plastome phylogenies, and codon-based models were used to test for positive selection. RESULTS: The plastomes were 154,496-155,778 bp with the canonical quadripartite structure and GC contents of 37.67-37.73%. Gene content was stable (131-132 genes; 86-87 protein-coding genes); B. falcatum subsp. cernuum lacked ycf15 but contained an additional inverted-repeat-associated ycf1 annotation. A/U-ending synonymous codons were favoured. Finite pairwise Ka/Ks estimates were below 1 for most genes, and site-specific codon models detected no positive selection. Each plastome contained 55-61 pure microsatellites, dominated by A/T mononucleotide motifs. MarkerSeek ranked 265 features and identified atpF-atpH, petA-psbJ, rpl32-trnL-UAG, and ycf1 as leading candidate barcodes. ycf1 recovered 38 of 41 nodes strongly supported by both reference trees, whereas a partitioned four-locus analysis recovered 40 of 41 and distinguished all 51 accession sequences. However, only one of seven multi-accession operational binomial groups was monophyletic, and only one showed a positive local barcode gap. The whole-plastome phylogeny recovered Bupleurum as monophyletic relative to Chamaesium. The two sampled Penninervia accessions occupied early-diverging positions without forming an exclusive clade. B. falcatum subsp. cernuum was sister to B. ranunculoides, with B. ranunculoides subsp. telonense sister to that pair. A partitioned 74-CDS analysis recovered the same key relationships and 45 of 50 internal bipartitions. Across the six newly assembled plastomes, 57-63 nonsynonymous C-to-U candidates were predicted per accession (367 total) in 21-22 genes; 269 affected the second codon position and 98 the first. DISCUSSION: Bupleurum plastomes are structurally conservative but retain localised divergence useful for marker development. Concordant whole-plastome and CDS genealogies support genus monophyly, whereas sparse Penninervia sampling and maternal plastid inheritance preclude rejecting traditional subgeneric classification. The predicted RNA-editing sites represent candidates for future experimental validation rather than an established Bupleurum editome. These genomic resources support authentication, conservation, and evolutionary research in Bupleurum.

Apiaceae

Synthetic DNA barcodes identify singlets in scRNA-seq datasets and evaluate doublet&#xa0;algorithms.

Single-cell RNA sequencing (scRNA-seq) datasets contain true single cells, or singlets, in addition to cells that coalesce during the protocol, or doublets. Identifying singlets with high fidelity in scRNA-seq is necessary to avoid false negative and false positive discoveries. Although several methodologies have been proposed, they are typically tested on highly heterogeneous datasets and lack a priori knowledge of true singlets. Here, we leveraged datasets with synthetically introduced DNA barcodes for a hitherto unexplored application: to extract ground-truth singlets. We demonstrated the feasibility of our framework, "singletCode," to evaluate existing doublet detection methods across a range of contexts. We also leveraged our ground-truth singlets to train a proof-of-concept machine learning classifier, which outperformed other doublet detection algorithms. Our integrative framework can identify ground-truth singlets and enable robust doublet detection in non-barcoded datasets.

Algorithms

Identification of a robust promoter in mouse and human hepatocytes by in vivo biopanning of a barcoded AAV library.

Recombinant adeno-associated viruses (AAVs) are leading vectors for in vivo human gene therapy. An integral vector element is promoters, which control transgene expression in either a ubiquitous or cell-type-selective manner. Identifying optimal capsid-promoter combinations is challenging, especially when considering on- versus off-target expression. Here, we report a pipeline for in vivo promoter biopanning in AAV building on our AAV capsid barcoding technology and illustrate its potential by screening 53 promoters in 16 murine tissues using an AAV9 vector. Surprisingly, the 2.2-kb human glial fibrillary acidic protein (GFAP) promoter was the top hit in the liver, where it outperformed robust benchmarks such as the human &#x3b1;-1-antitrypsin promoter or the clinically used liver-specific promoter 1 (LP1). Analysis of hepatic cell populations revealed preferred GFAP promoter activity in hepatocytes. Notably, the GFAP promoter also surpassed the LP1 and cytomegalovirus promoters in human hepatocytes engrafted in an immune-deficient mouse. These findings establish the GFAP promoter as an exciting alternative for research and clinical applications requiring efficient and specific transgene expression in hepatocytes. Our pipeline expands the arsenal of technologies for high-throughput in vivo screening of viral vector components and is compatible with capsid barcoding, facilitating the combinatorial interrogation of complex AAV libraries.

Dependovirus

Doblin: inferring dominant clonal lineages from high-resolution DNA barcoding time series.

MOTIVATION: The lineage dynamics and history of cells in a population reflect the interplay of evolutionary forces they experience, including mutation, drift, and selection. When the population is polyclonal, lineage dynamics also manifest the extent of clonal competition among co-existing mutational variants. If the population exists in a community of other species, the lineage dynamics could also reflect the population's ecological interaction with the rest of the community. Recent advances in high-resolution lineage tracking via DNA barcoding, coupled with next-generation sequencing of bacteria, yeast, and mammalian cells, allow for precise quantification of clonal dynamics in these organisms. RESULTS: In this work, we introduce Doblin, an R suite for identifying dominant barcode lineages based on high-resolution lineage tracking data. We first benchmarked Doblin's accuracy using lineage data from evolutionary simulations, showing that it recovers the clones' identity and relative fitness in the simulation. Next, we applied Doblin to analyze clonal dynamics in laboratory evolutions of Escherichia coli populations undergoing antibiotic treatment and in colonization experiments of the gut microbial community. Doblin's versatility allows it to be applied to lineage time-series data across different experimental setups. AVAILABILITY AND IMPLEMENTATION: Doblin is available on CRAN (https://CRAN.R-project.org/package=doblin) and Github (https://github.com/dagagf/doblin).

DNA Barcoding, Taxonomic

Species identification, discovery, and biomonitoring: Strategic priorities for DNA barcoding in Europe, set in a global context.

The International Barcode of Life (iBOL) initiative is building a globally accessible DNA-based system for species identification and discovery. This paper outlines the mission and strategic priorities for the iBOL community in Europe (iBOL Europe), set in a global context. The mission of iBOL Europe is to produce, curate, and provide access to a complete DNA barcode reference library of European eukaryotic biodiversity, catalyzing species discovery and enabling comprehensive, harmonized species identification and biomonitoring, and supporting the global iBOL program. Immediate objectives include completing reference libraries for priority taxa, democratizing access to sequencing technologies, and strengthening a distributed community of practice. Key actions identified span five thematic areas: community building, sample collection and taxonomic verification, sequencing infrastructure, data management, and mainstreaming DNA-based approaches to meet societal needs. The strategy emphasizes integration with European research infrastructures to ensure long-term sustainability and resilience for biodiversity genomics in Europe.

DNA barcoding

Barcoded oligonucleotide system (BOLT) for targeted organ delivery.

The therapeutic potential of oligonucleotides (oligos) is limited by insufficient delivery to extrahepatic tissues. In vitro assays often fail to accurately predict in vivo behavior, while testing each oligo candidate in animals remains inherently low throughput. Here, we conceive a barcoded oligonucleotide system (BOLT), a platform that enables high-throughput in vivo evaluations of small-molecule ligands and identifies tissue-specific oligo delivery. BOLT integrates rational design of oligo barcodes, modular conjugation chemistry, and next-generation sequencing (NGS)-based quantification, allowing simultaneous evaluation of many chemically diverse ligand-oligo conjugates within a single animal. Notably, this platform is applicable in both mice and nonhuman primates (NHPs). Using BOLT, we discovered ligands with tropism for tissues such as the brain, lung, and muscle. Collectively, these results indicate that the BOLT platform can accelerate the discovery of tissue-targeting ligands for broad oligo therapeutics.

Journal Article

Dosa: A method to covalently barcode proteins for high throughput biochemistry.

Deep mutational scanning couples a protein's activity to DNA sequencing for high throughput assessment of the effects of all single amino acid substitutions, but it largely uses indirect assays, like growth, as proxy for protein activity. Here, we covalently link variant proteins in vivo to an RNA barcode by fusing them to E. coli tRNA (m5U54) methyltransferase TrmA (E358Q), which forms a covalent bond with a tRNA stem-loop. Following cell lysis, variant proteins are separated in vitro according to their biochemical properties and identified by their barcodes. We use this method, Dosa, to analyze a large pool of FLAG epitope variants for binding to an anti-FLAG antibody, to profile the cleavage preferences of variants of enteropeptidase and human rhinovirus 3C protease, and to measure the solubility of several hundred A&#x3b2;(1-42) variants. This method should be amenable to numerous biochemical assays with proteins produced in E. coli or mammalian cells.

Protein display

Performance Profiles of Short DNA Barcode Segments for Family Level Detection of Asteraceae Within Asterales.

Short DNA barcodes may facilitate sequence recovery from degraded material, but their ability to retain target-family identity while excluding related taxa varies among genomic regions. We computationally evaluated 16 nuclear, plastid, and mitochondrial marker regions from 11 Asterales families using 279,956 NCBI locus-record matches and an accession-disjoint discovery/test design. Thirty-one candidate segments of 50-200 bp (mean, 98.55 bp) were screened in discovery data and evaluated for within-Asteraceae sequence recall, differentiation from non-Asteraceae Asterales, in silico primer behavior, phylogenetic placement, and exploratory matching across 808 metadata-defined metagenomic samples. Conserved regions such as matR and rbcL showed high within-Asteraceae identity, whereas ITS1, ITS, and trnH-psbA showed larger differences from related-family backgrounds; ITS2 and ycf1 showed intermediate profiles. Candidate segments were placed within or immediately adjacent to Asteraceae reference branches in segment-specific maximum-likelihood analyses, although support and topology varied among regions. Metadata-defined target-containing groups had higher mean query coverage and identity than background groups; because target presence was not independently verified and no classifier was fitted, these comparisons were descriptive and did not estimate diagnostic accuracy. Definitionally linked sequence statistics were interpreted as structural associations rather than evidence of causal evolutionary mechanisms. These results provide a family-level computational comparison of candidate short segments for Asteraceae detection within Asterales. Species identification, operational marker combinations, threshold robustness, and laboratory performance require validation using taxonomically dense, voucher-linked, and experimentally characterized datasets.

Asteraceae

An in vivo barcoded CRISPR-Cas9 screen identifies Ncoa4-mediated ferritinophagy as a dependence in Tet2-deficient hematopoiesis.

TET2 is among the most commonly mutated genes in both clonal hematopoiesis and myeloid malignancies; thus, the ability to identify selective dependencies in TET2-deficient cells has broad translational significance. Here, we identify regulators of Tet2 knockout (KO) hematopoietic stem and progenitor cell (HSPC) expansion using an in vivo CRISPR-Cas9 KO screen, in which nucleotide barcoding enabled large-scale clonal tracing of Tet2-deficient HSPCs in a physiologic setting. Our screen identified candidate genes, including Ncoa4, that are selectively required for Tet2 KO clonal outgrowth compared with wild type. Ncoa4 targets ferritin for lysosomal degradation (ferritinophagy), maintaining intracellular iron homeostasis by releasing labile iron in response to cellular demands. In Tet2-deficient HSPCs, increased mitochondrial adenosine triphosphate production correlates with increased cellular iron requirements and, in turn, promotes Ncoa4-dependent ferritinophagy. Restricting iron availability reduces Tet2 KO stem cell numbers, revealing a dependency in TET2-mutated myeloid neoplasms.

CRISPR-Cas Systems

Short-read genome skimming enables molecular barcoding of old myxomycete collections.

This study evaluates the effectiveness of Illumina-based genome skimming for barcoding myxomycete herbarium collections ranging from 29 to 91 years in age. We successfully retrieved partial sequences of the standard marker gene (nucSSU) in all cases, as well as additional markers (mtSSU, EF1a, and COI) for certain collections. Altogether, 28 genes were recognized in the studied material. In a 33-year-old specimen of Lindbladia tubulina, the assembly reached an N50 of 4.19 kb, enabling the recovery of extended functional loci. The input genomic DNA quantity emerges as the primary determinant of sequencing success. Samples with high DNA yields provide representative amounts of contigs coming confirmedly (matching sequences in the NCBI nucleotide database) or potentially (no-hit fraction) from myxomycetes, regardless of specimen age. In addition to target DNA, we revealed distinct signals of both anthropogenic contamination (human DNA and skin microflora) and natural substrate inhabitants, including oribatid mites and bacteria from dead wood, soil, and grass litter. Thus, even in old collections, metagenomic data still carry information regarding the substrate upon which the myxomycete developed. The results demonstrate that short-read genome skimming may help to integrate historical type material of myxomycetes into contemporary phylogenetic research. This method overcomes the length-dependent limitations of traditional Sanger sequencing, thus providing a roadmap for the future of museomics in myxomycetology.

Amoebozoa

Comprehensive plastome variation and RNA editing in Mentha: insights into phylogenetic relationships and candidate DNA barcodes.

INTRODUCTION: Mentha is an economically and medicinally important genus in Lamiaceae, but its taxonomy and species delimitation remain challenging because of frequent hybridization, polyploidy, and marked morphological plasticity. METHODS: In this study, we comparatively analyzed 12 plastomes representing major Mentha species, hybrid taxa, and unresolved accessions, including four newly assembled genomes, to characterize plastome structure, repeat composition, sequence divergence, phylogenetic relationships, and plastid RNA editing. The M. arvensis plastome and RNA-seq datasets originated from independent Swiss and Indian accessions, respectively. RESULTS: The plastomes were highly conserved in overall organization, ranging from 151,824 to 152,154 bp and displaying the typical quadripartite structure. Gene content and order were largely stable across taxa, with only minor variation likely associated with annotation differences at IR/SC boundary regions. Codon usage analysis revealed a clear bias toward A/U-ending synonymous codons, and most shared protein-coding genes showed low Ka/Ks ratios, indicating predominant purifying selection. Repeat analyses showed that simple sequence repeats were mainly composed of A/T-rich mononucleotide motifs, whereas long repeats were concentrated in the 30-40 bp size class. Comparative analyses identified six hypervariable regions, namely ccsA-ndhD, ycf1, ndhD, rpl32-trnL-UAG, rbcL-accD, and petA-psbJ, which represent promising candidate plastid markers for species discrimination. Phylogenetic analysis based on complete plastomes provided strong support for relationships among the sampled taxa and recovered a close affinity among M. aquatica, M. arvensis, and M. canadensis. In addition, RNA-seq analysis of M. arvensis identified 17 candidate plastid RNA editing sites, most of which were C-to-U conversions and nonsynonymous events. DISCUSSION: Together, these results expand plastid genomic resources for Mentha and provide a useful framework for phylogenetic inference, species identification, and future germplasm utilization.

RNA editing

Shedding dynamics of a DNA virus population during acute and long-term persistent infection.

Although much is known of the molecular mechanisms of virus infection within cells, substantially less is understood about within-host infection. Such knowledge is key to understanding how viruses take up residence and transmit infectious virus, in some cases throughout the life of the host. Here, using murine polyomavirus (muPyV) as a tractable model, we monitor parallel infections of thousands of differentially barcoded viruses within a single host. In individual mice, we show that numerous viruses (>2600) establish infection and are maintained for long periods post-infection. Strikingly, a low level of many different barcodes is shed in urine at all times post-infection, with a minimum of at least 80 different barcodes present in every sample throughout months of infection. During the early acute phase, bulk shed virus genomes derive from numerous different barcodes. This is followed by long term persistent infection detectable in diverse organs. Consistent with limited productive exchange of virus genomes between organs, each displays a unique pattern of relative barcode abundance. During the persistent phase, constant low-level shedding of typically hundreds of barcodes is maintained but is overlapped with rare, punctuated shedding of high amounts of one or a few individual barcodes. In contrast to the early acute phase, these few infrequent highly shed barcodes comprise the majority of bulk shed genomes observed during late times of persistent infection, contributing to a stark decrease in bulk barcode diversity that is shed over time. These temporally shifting patterns, which are conserved across hosts, suggest that polyomaviruses balance continuous transmission potential with reservoir-driven high-level reactivation. This offers a mechanistic basis for polyomavirus ubiquity and long-term persistence, which are typical of many DNA viruses.

Animals