PubMed HealthSearch

SEARCH · PubMed Health

Results for “DNA barcoding”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

Performance comparison of rapid and native barcoding methods for Oxford Nanopore sequencing of Poliovirus Viral Protein 1 (VP1) amplicons.

Accurate and timely sequencing of poliovirus is critical for global eradication efforts, particularly for molecular epidemiology based on the typing region of the genome, viral protein 1 (VP1). While Oxford Nanopore Technologies (ONT) sequencing has expanded capabilities for poliovirus surveillance, the relative performance of different ONT library preparation methods, including ligation-based (Native Barcoding) and transposase-based (Rapid Barcoding) approaches, has not been systematically evaluated. In this study, we compared rapid barcoding and native barcoding workflows for sequencing VP1 amplicons from 17 type 2 poliovirus-positive samples, each processed in triplicate. Native barcoding generated significantly more sequencing output, producing approximately 2.3-fold greater total read yield than rapid barcoding, and demonstrated higher run-to-run reproducibility (R2 = 0.979-0.998 vs. 0.847-0.929, respectively; p&#x202f;<&#x202f;0.001). In addition, native barcoding generated 80% of the total yield achieved by rapid barcoding within approximately 7&#x202f;h, whereas rapid barcoding required approximately 40&#x202f;h to reach the same output. Despite these differences, both methods produced identical VP1 consensus sequences across all samples, with comparable read quality (median per-base Q-scores of approximately Q17-Q18). Rapid barcoding provided substantial practical advantages, reducing hands-on library preparation time (55 vs. 200&#x202f;min) and per-sample cost ($12.82 vs. $16.54), while simplifying workflow and reducing technical complexity. These findings indicate that sequencing yield may not be a determinant of downstream analytical outcomes for poliovirus VP1 ONT sequencing. Rapid barcoding therefore represents a cost-effective and efficient approach for routine poliovirus surveillance, whereas native barcoding remains advantageous in applications requiring rapid data generation or maximal sequencing depth.

Poliovirus

Fluctuating DNA methylation tracks cancer evolution at clinical scale.

Cancer development and response to treatment are evolutionary processes1,2, but characterizing evolutionary dynamics at a clinically meaningful scale has remained challenging3. Here we develop a new methodology called EVOFLUx, based on natural DNA methylation barcodes fluctuating over time4, that quantitatively infers evolutionary dynamics using only a bulk tumour methylation profile as input. We apply EVOFLUx to 1,976 well-characterized lymphoid cancer samples spanning a broad spectrum of diseases and show that initial tumour growth rate, malignancy age and epimutation rates vary by orders of magnitude across disease types. We measure that subclonal selection occurs only infrequently within bulk samples and detect occasional examples of multiple independent primary tumours. Clinically, we observe faster initial tumour growth in more aggressive disease subtypes, and that evolutionary histories are strong independent prognostic factors in two series of chronic lymphocytic leukaemia. Using EVOFLUx for phylogenetic analyses of aggressive Richter-transformed chronic lymphocytic leukaemia samples detected that the seed of the transformed clone existed decades before presentation. Orthogonal verification of EVOFLUx inferences is provided using additional genetic data, including long-read nanopore sequencing, and clinical variables. Collectively, we show how widely available, low-cost bulk DNA methylation data precisely measure cancer evolutionary dynamics, and provides new insights into cancer biology and clinical behaviour.

Humans

Metschnikowia maris comb. nov., a large-spored yeast species endemic to Serra do Mar Atlantic Rainforest biome, Sao Paulo State, Brazil.

Two yeast isolates from passion flowers were sampled in the southern part of the Serra do Mar Atlantic Rainforest in Sao Paulo State, Brazil. Barcode sequencing and mating experiments showed them to be representatives of Metschnikowia matae var. maris, thus originally named due to the availability of only a single isolate and uncertainties regarding reproductive isolation. The two new isolates being of the complementary mating type to the previously known strain, intravarietal crosses were performed. They yielded a preponderance of two-spored asci, unlike crosses with M. matae var. matae, which led to largely sterile asci. We therefore elevate the variety maris to the rank of species, with the name Metschnikowia maris comb. nov. The holotype is UFMG-CM-Y397T (MAT&#x3b1;). Strain UFMG-CM-Y7613A (MAT a) is designated as allotype. The new combination is registered as MB 859665.

Brazil

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100&#x2009;000 and 1&#x2009;000&#x2009;000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

GenBank mining reveals novel insights into Rhizobium phylogeny: Identical 16S rRNA sequences are mainly uncoupled from species designation, host plant, and geographic origin: How this search suggested the definition of a direct 'microbial h-index'.

16S rDNA is the historical gold standard for bacterial identification, particularly in metabarcoding approaches reliant on sequence similarity thresholds. We analyzed 6,660 Rhizobium 16S rRNA gene sequences from GenBank to examine the relationship between sequence identity and three metadata: species name, host plant, and geographic origin. Using an iterative BLAST-based pipeline, we detected 116,069 pairwise matches and assessed concordance among sequences (average length 1,328 bp) sharing 100% identity. For those in which the organism name, host plant and country of isolation were present in the record, surprisingly, 66.59% of identical sequence pairs showed full discordance across all three metadata, while only 1.40% shared the same name, host, and country. The most widespread sequence, detected 371 times, was associated with over 56 different host plants across 25 countries and bore multiple species name designations. These results highlight a striking mismatch between the 16S barcode and the taxonomic, ecological, and phenotypic variability it is assumed to reflect, likely arising from the slow evolution of rRNA genes contrasted with the mobility of ecologically relevant genes via horizontal transfer on plasmids, transposons, and phages. Our findings further challenge the limitations of relying on 16S rRNA alone for fine-scale taxonomic and metadata-based inference in capturing the true functional and ecological diversity of bacteria, endorsing the critical importance of polyphasic taxonomic approaches that integrate genomic, phenotypic, and ecological data. An interesting byproduct of the analysis was to realize the possibility of treating these data as if they were 'citations.' The more one finds the same query sequence, the more that sequence can be considered biologically 'cited', i.e., re-proposed elsewhere in the world. Thus, one can also analyze the h-index of such a ranking. In our Rhizobium dataset, we calculated an h-index&#x2009;=&#x2009;201, meaning the sequence ranked 201st had 202 identical homologues in GenBank. Although the research effort on given species is directly connected with it, this number provides a quantitative indicator of a taxon's sequence recurrence and distribution within public databases, independent of nomenclatural inconsistencies, offering a novel framework for assessing bacterial representation across global datasets.

RNA, Ribosomal, 16S

Recurrent and niche-specific functional bacteriome of maize hybrid revealed by integrated metabarcoding and culturomics.

The plant microbiome plays a pivotal role in plant survival in natural habitats by facilitating nutrient acquisition, stress adaptation, and disease suppression, while also offering opportunities to enhance crop productivity and climate resilience. However, the distribution of persistent and culturable bacteriome across maize-associated niches and their functional potential remain poorly resolved. This study integrated metagenomic next-generation sequencing (mNGS-based metabarcoding) and culturomics to characterise the maize-associated bacteriome of bulk soil, rhizoplane, phylloplane, and cob of the maize hybrid PHM-1 under contrasting cropping and tillage systems, and to identify recurrent and agriculturally promising bacteriome components. The bacteriome exhibited pronounced niche-specific structuring, whereas overall bacterial community composition did not differ significantly across cropping and tillage treatments (ANOSIM, R&#x2009;=&#x2009;0.038, p&#x2009;=&#x2009;0.306). Proteobacteria predominated in the culturable bacteriome (69-84%; mean, 76.2%) but accounted for only 1% of the total bacteriome, whereas Patescibacteria and Firmicutes were relatively enriched. Niche-specific dominance was evident, with Pantoea accounting for 40.79% of the total and 56.27% of the culturable phylloplane bacteriome under cereal monocropping, while Serratia represented 31.59% and 59.40% of the total and culturable cob bacteriomes, respectively. Across niches, mNGS captured substantially greater bacteriome diversity, particularly uncultured and unidentified taxa in soil-associated compartments, whereas culturomics recovered a narrower but functionally accessible fraction. Culturomics yielded 99 isolates representing 32 species across 12 genera, including six genera shared with the mNGS-derived recurrent bacteriome: Bacillus, Enterobacter, Pantoea, Pseudomonas, Serratia, and Stenotrophomonas. Functional screening identified strong biocontrol and plant-beneficial traits among core-associated isolates. Pseudomonas oryzihabitans ZM-DL-PA10 inhibited Rhizoctonia solani, Macrophomina phaseolina, and Bipolaris maydis by up to 40.6%, 43.9%, and 45.2%, respectively, through secreted and volatile metabolites; exhibited P, K, and Zn solubilisation; and produced IAA and siderophores. It also recorded the lowest B. maydis disease index (ADI) of 1.00. Pantoea ananatis ZM-BH-EA4 showed 52.4% and 68.5% inhibition of R. solani and B. maydis, respectively, through volatile metabolites. Collectively, the integration of mNGS and culturomics revealed a strongly compartmentalised maize bacteriome and identified recurrent, culturable, and functionally promising bacterial taxa, providing a targeted resource for microbiome-based crop protection and climate-resilient maize production.

Zea mays

Common to rare transfer learning (CORAL) enables inference and prediction for a quarter million rare Malagasy arthropods.

DNA-based biodiversity surveys result in massive-scale data, including up to millions of species-of which, most are rare. Making the most of such data for inference and prediction requires modeling approaches that can relate species occurrences to environmental and spatial predictors, while incorporating information about their taxonomic or phylogenetic placement. Even if the scalability of joint species distribution models to large communities has greatly advanced, incorporating hundreds of thousands of species has not been feasible to date, leading to compromised analyses. Here we present a 'common to rare transfer learning' (CORAL) approach, based on borrowing information from the common species to enable statistically and computationally efficient modeling of both common and rare species. We illustrate that CORAL leads to much improved prediction and inference in the context of DNA metabarcoding data from Madagascar, comprising 255,188 arthropod species detected in 2,874 samples.

Animals

NanoASV: a snakemake workflow for reproducible field-based Nanopore full-length 16S metabarcoding amplicon data analysis.

SUMMARY: NanoASV is a conda environment and snakemake-based workflow using state-of-the-art bioinformatics software to process full-length SSU rRNA (16S/18S) amplicons acquired with Oxford Nanopore Sequencing technology. Its strength lies in reproducibility, portability, and the possibility to run offline, allowing in-field analysis. It can be installed on the Nanopore MK1C sequencing device and process data locally. AVAILABILITY AND IMPLEMENTATION: Source code and documentation are freely available at https://github.com/ImagoXV/NanoASV and Zenodo archive at https://doi.org/10.5281/zenodo.14730742.

Software

Predicting coarse-grained representations of biogeochemical cycles from metabarcoding data.

MOTIVATION: Taxonomic analysis of environmental microbial communities is now routinely performed thanks to advances in DNA sequencing. Determining the role of these communities in global biogeochemical cycles requires the identification of their metabolic functions, such as hydrogen oxidation, sulfur reduction, and carbon fixation. These functions can be directly inferred from metagenomics data, but in many environmental applications metabarcoding is still the method of choice. The reconstruction of metabolic functions from metabarcoding data and their integration into coarse-grained representations of biogeochemical cycles remains a difficult bioinformatics problem today. RESULTS: We developed a pipeline, called Tabigecy, which exploits taxonomic affiliations to predict metabolic functions constituting biogeochemical cycles. In a first step, Tabigecy uses the tool EsMeCaTa to predict consensus proteomes from input affiliations. To optimize this process, we generated a precomputed database containing information about 2404 taxa from UniProt. The consensus proteomes are searched using bigecyhmm, a newly developed Python package relying on Hidden Markov Models to identify key enzymes involved in metabolic function of biogeochemical cycles. The metabolic functions are then projected on coarse-grained representation of the cycles. We applied Tabigecy to two salt cavern datasets and validated its predictions with microbial activity and hydrochemistry measurements performed on the samples. The results highlight the utility of the approach to investigate the impact of microbial communities on biogeochemical processes. AVAILABILITY AND IMPLEMENTATION: The Tabigecy pipeline is available at https://github.com/ArnaudBelcour/tabigecy. The Python package bigecyhmm and the precomputed EsMeCaTa database are also separately available at https://github.com/ArnaudBelcour/bigecyhmm and https://doi.org/10.5281/zenodo.13354073, respectively.

Metagenomics

The role of stochasticity in fungal community assembly: explaining apparent stochasticity with field experiments.

Stochasticity is a main process in community assembly. However, experimental studies rarely target stochasticity in natural communities, and hence experimental validation of stochasticity estimates in observational studies is lacking. Here, we combine experimental and observational data to unravel the role of stochasticity in the assembly of wood-inhabiting fungi. We carried out a replicated field experiment where the natural colonization of a focal fungal species was simulated through inoculation, and the local fungal communities were monitored through DNA metabarcoding before and after the inoculations. The amount of stochasticity in fungal colonization was less pronounced than expected from the amount of unpredictability in observational data, suggesting that stochasticity may play a smaller role in fungal occurrence than previously anticipated, or that it may be a stronger influence in the dispersal and establishment phases than in colonization per se. Stochasticity was more prominent in the initial phase of community succession, with the earliest successional stage involving a higher level of stochasticity than the later stage after 2 years. We conclude that experimentally measuring the role of stochasticity in community assembly is feasible for species-rich communities under natural conditions and highlight the importance of experimentally testing the accuracy of stochasticity estimates based on observational data.

Stochastic Processes

Multi-step genomics on single cells and live cultures in sub-nanoliter capsules.

Single-cell sequencing methods uncover natural and induced variation between cells. Many functional genomic methods, however, require multiple steps that cannot yet be scaled to high throughput, including assays on living cells. Here we develop capsules with amphiphilic gel envelopes (CAGEs), which selectively retain cells and large analytes while being freely accessible to media, enzymes and reagents. Capsules enable high-throughput multi-step assays combining live-cell culture with genome-wide readouts. We establish methods for barcoding CAGE DNA libraries, and apply them to measure persistence of gene expression programs in cells by capturing the transcriptomes of tens of thousands of expanding clones in CAGEs. The compatibility of CAGEs with diverse enzymatic reactions will facilitate the expansion of the current repertoire of single-cell, high-throughput measurements and extension to live-cell assays.

Journal Article

Getting to the Core of the Matter-Assessing the Role of Replication in Metabarcoding-Based sedaDNA.

Replication is central to most experimental and sampling designs, increasing inferential power and capturing fine-scale data heterogeneity. However, its importance remains poorly evaluated in some ecological and evolutionary settings. This is the case of metabarcoding studies using DNA recovered from sedimentary archives, in which biological signals integrate ecological information through depositional and burial processes, yet are commonly inferred from a single sediment core per site. Here, we evaluated the effect of different types of replication using sedimentary DNA metabarcoding data from two genetic markers (mitochondrial COI and nuclear 18S) using a nested sampling design. The design included three intertidal sites, three spatially separated sediment cores per site (biological replicates), two sediment horizons per core, and eight PCR (technical) replicates per sediment sample. Variance partitioning showed that site identity and sediment age group together explained >&#x2009;70% of the variation in beta diversity, indicating that among-site spatial and stratigraphic differences were the dominant drivers of community composition. PERMANOVA likewise identified non-significant effects of biological replication. Among PCR replicates from the same sediment sample, richness varied substantially, whereas Shannon diversity was more consistent. Despite this variability, differences in community composition among technical replicates remained smaller than those associated with biological replication or site identity, indicating a limited influence on broader ecological patterns. Community composition was highly similar among replicate cores within sites, consistent with stratigraphic coherence. These results indicate limited within-site heterogeneity and suggest that, under stratigraphically coherent conditions, increasing biological replication may provide little additional information, whereas enhancing technical replication and stratigraphic resolution can improve ecological inference from sedimentary DNA metabarcoding datasets.

DNA Barcoding, Taxonomic

Integration of Imaging-based and Sequencing-based Spatial Omics Mapping on the Same Tissue Section via DBiTplus.

Spatially mapping the transcriptome and proteome in the same tissue section can significantly advance our understanding of heterogeneous cellular processes and connect cell type to function. Here, we present Deterministic Barcoding in Tissue sequencing plus (DBiTplus), an integrative multi-modality spatial omics approach that combines sequencing-based spatial transcriptomics and image-based spatial protein profiling on the same tissue section to enable both single-cell resolution cell typing and genome-scale interrogation of biological pathways. DBiTplus begins with in situ reverse transcription for cDNA synthesis, microfluidic delivery of DNA oligos for spatial barcoding, retrieval of barcoded cDNA using RNaseH, an enzyme that selectively degrades RNA in an RNA-DNA hybrid, preserving the intact tissue section for high-plex protein imaging with CODEX. We developed computational pipelines to register data from two distinct modalities. Performing both DBiT-seq and CODEX on the same tissue slide enables accurate cell typing in each spatial transcriptome spot and subsequently image-guided decomposition to generate single-cell resolved spatial transcriptome atlases. DBiTplus was applied to mouse embryos with limited protein markers but still demonstrated excellent integration for single-cell transcriptome decomposition, to normal human lymph nodes with high-plex protein profiling to yield a single-cell spatial transcriptome map, and to human lymphoma FFPE tissue to explore the mechanisms of lymphomagenesis and progression. DBiTplusCODEX is a unified workflow including integrative experimental procedure and computational innovation for spatially resolved single-cell atlasing and exploration of biological pathways cell-by-cell at genome-scale.

Journal Article

Integration of Imaging-based and Sequencing-based Spatial Omics Mapping on the Same Tissue Section via DBiTplus.

Spatially mapping the transcriptome and proteome in the same tissue section can significantly advance our understanding of heterogeneous cellular processes and connect cell type to function. Here, we present Deterministic Barcoding in Tissue sequencing plus (DBiTplus), an integrative multi-modality spatial omics approach that combines sequencing-based spatial transcriptomics and image-based spatial protein profiling on the same tissue section to enable both single-cell resolution cell typing and genome-scale interrogation of biological pathways. DBiTplus begins with in situ reverse transcription for cDNA synthesis, microfluidic delivery of DNA oligos for spatial barcoding, retrieval of barcoded cDNA using RNaseH, an enzyme that selectively degrades RNA in an RNA-DNA hybrid, preserving the intact tissue section for high-plex protein imaging with CODEX. We developed computational pipelines to register data from two distinct modalities. Performing both DBiT-seq and CODEX on the same tissue slide enables accurate cell typing in each spatial transcriptome spot and subsequently image-guided decomposition to generate single-cell resolved spatial transcriptome atlases. DBiTplus was applied to mouse embryos with limited protein markers but still demonstrated excellent integration for single-cell transcriptome decomposition, to normal human lymph nodes with high-plex protein profiling to yield a single-cell spatial transcriptome map, and to human lymphoma FFPE tissue to explore the mechanisms of lymphomagenesis and progression. DBiTplusCODEX is a unified workflow including integrative experimental procedure and computational innovation for spatially resolved single-cell atlasing and exploration of biological pathways cell-by-cell at genome-scale.

Journal Article

Plasticity of extrachromosomal DNA segregation during drug adaptation.

Uneven segregation during mitosis is a striking feature of extrachromosomal DNA (ecDNA). Because ecDNA lacks a centromere, it is thought to segregate stochastically, generating intratumoral heterogeneity in genomic copy number. Drug treatment can readily change ecDNA copy number, enabling cells to acquire drug resistance, yet whether these changes reflect static selection of pre-existing clones or active reconfiguration under stress remains unresolved. To address this, we develop a high-throughput framework combining single-cell DNA sequencing with cellular barcoding for clonal tracking. Single-cell cloning reveals that not all clones exhibit identical segregation modes even under drug-free conditions. Under treatment, resistant populations do not simply arise from pre-existing clones with favorable ecDNA states; instead, some clones actively reconfigure their segregation behavior to generate resistant cells. Thus, although ecDNA generally segregates stochastically, it can undergo nonrandom, actively regulated segregation under drug stress, raising the possibility of therapeutically targeting ecDNA segregation mechanisms to counteract adaptive resistance.

Extrachromosomal DNA

Benchmarking DNA extraction protocols across use cases for culture-independent Nanopore metagenomics.

Oxford Nanopore Technologies (ONT) sequencing offers several advantages for metagenomics, including long reads, rapid turnaround, low upfront cost, scalability and portability. However, for ONT metagenomics, DNA yield, quality and integrity are important considerations when selecting an extraction method. Many metagenomic extraction methods use harsh lysis conditions to extract a wide range of species and provide an accurate community composition, but these conditions can compromise DNA fragment length. Therefore, extraction methods for ONT metagenomics must balance DNA shearing and recovery with representative community lysis. We systematically evaluated DNA extraction methods for ONT metagenomic sequencing using a use case-oriented framework. Among nearly 50 extraction methods screened, 7 were selected for detailed comparison based on suitability for metagenomics, variation in methodology, availability, cost and processing time: Norgen BioTek Corp's Stool DNA Isolation (NG), Zymo Research's ZymoBIOMICS Quick-DNA HMW MagBead (ZMG), Qiagen's DNeasy Blood and Tissue (QBT), Macherey-Nagel's NucleoMag DNA Microbiome (MN), Zymo Research's ZymoBIOMICS DNA Mini Prep (ZMI), Qiagen's DNeasy PowerSoil/QIAamp PowerFecal Pro (PS) and Qiagen's QIAamp Fast DNA Stool Mini (QIA). Methods were tested using Zymo Research's ZymoBIOMICS Microbial Community Standard (MCS), a matrix-free mock community with known composition. DNA extracts were sequenced on an ONT PromethION using the Rapid Barcoding Kit, except QIA due to insufficient DNA yield. Metrics for the method, DNA extracts, sequencing and genomes were evaluated, revealing trade-offs between methods. The two magnetic bead methods, MN and ZMG, produced the highest mean read length N50 values (13.9 and 16.5&#x2009;kb, respectively) but showed apparent community compositions skewed towards Gram-negative bacteria. In contrast, ZMI and PS maintained a community composition close to expected, with reduced mean read length N50 values (4.5 vs. 7.5&#x2009;kb). Performance across various metrics is presented in the context of the following use cases: maximizing genome coverage and assembly completeness, preserving composition accuracy, targeting specific species and limiting required resources (equipment, time or budget). The metrics and use case considerations presented offer practical guidance for informed selection of DNA extraction methods for ONT metagenomics. For accurate community composition, ZMI or PS are recommended, while PS and ZMG perform best at maximizing genome coverage and assembly completeness. NG and QBT may be the most economical options, though performance trade-offs were observed. Finally, PS may be the preferred method for time-sensitive diagnostic or field applications.

Metagenomics

Is There a Fly in My Soup? To What Extent Do Metabarcoding and Individual Barcoding Tell the Same Story?

Metabarcoding has become the method of choice for characterizing complex arthropod communities. The extent to which metabarcoded bulk samples will recover the same community composition as individual sequencing of all individuals in the sample remains poorly quantified. Biases such as unequal extraction of DNA from different taxa, primer mismatches and non-random PCR may cause the selective drop-out of species from metabarcoding data. At the same time, DNA metabarcoding may reveal arthropod taxa present not as individuals, but as DNA residues on the surface or in the gut of insects. To quantify the consistency in sample contents established by different means, we metabarcoded 45 bulk insect samples, then extracted all arthropods and sequenced them individually. Metabarcoding targeted 418&#x2009;bp at the 3' end of the Folmer barcoding region, while individual barcodes captured the entire 658&#x2009;bp Folmer region. The metabarcoding workflow, including PCR amplification, sequencing and bioinformatics, was performed in three replicates from three separate lysate aliquots per sample. For the main analyses, sequences were assigned to Barcode Index Numbers (BINs) as identical taxonomic categories across data types, thereby allowing the detection of even rare but biologically true taxa. Since such reference-based validation will be unavailable to any researcher dealing with metabarcoding data alone, we validated our key findings through an alternative workflow, i.e., de novo clustering of sequences. We found that metabarcoding is replicable, as different replicates of the same sample recover similar species richness and composition. Individual barcoding and metabarcoding provide similar impressions of relative differences in community structure: species-rich vs. species-poor samples rank similarly among data types (Spearman's &#x2374;&#x2009;=&#x2009;0.88-0.99) as do differences in relative dissimilarity between sample pairs (Spearman's &#x2374;&#x2009;=&#x2009;0.55-0.90). Dissimilarity between data types varies with BIN richness in the sample, but this relationship reflects nestedness rather than turnover: metabarcoding recovers the same set of core species as individual barcoding but adds hundreds of species on top. Any BIN recovered as an individual occurred with high probability in the metabarcoding data, and any BIN found in high read abundances by metabarcoding was likely found as an individual (p&#x2009;>&#x2009;0.8). In terms of abundances, the number of individual insects per BIN was well predicted by the number of metabarcoding reads (R2&#x2009;>&#x2009;0.68 for a model including taxonomy as a random effect). Our analysis suggests that metabarcoding data will be informative of the sample contents in terms of arthropod species richness, composition and taxon-specific abundances. Taxa recovered in low copy numbers in metabarcoding sequence data will likely represent DNA left as residues from past biotic interactions. Barring sequencing errors, both types of data yield biologically relevant insights into the taxa present in the source community.

Animals

Directed evolution of engineered virus-like particles with improved production and transduction efficiencies.

Engineered virus-like particles (eVLPs) are promising vehicles for transient delivery of proteins and RNAs, including gene editing agents. We report a system for the laboratory evolution of eVLPs that enables the discovery of eVLP variants with improved properties. The system uses barcoded guide RNAs loaded within DNA-free eVLP-packaged cargos to uniquely label each eVLP variant in a library, enabling the identification of desired variants following selections for desired properties. We applied this system to mutate and select eVLP capsids with improved eVLP production properties or transduction efficiencies in human cells. By combining beneficial capsid mutations, we developed fifth-generation (v5) eVLPs, which exhibit a 2-4-fold increase in cultured mammalian cell delivery potency compared to previous-best v4 eVLPs. Analyses of v5 eVLPs suggest that these capsid mutations optimize packaging and delivery of desired ribonucleoprotein cargos rather than native viral genomes and substantially alter eVLP capsid structure. These findings suggest the potential of barcoded eVLP evolution to support the development of improved eVLPs.

Humans