PubMed HealthSearch

SEARCH · PubMed Health

Results for “taxonomy assignment”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

10 recordsLinked to original sources

Persistent gaps and errors in reference databases impede ecologically meaningful taxonomy assignments in 18S rRNA studies: a case study of terrestrial and marine nematodes.

In metabarcoding studies, Linnaean taxonomy assignments of Operational Taxonomic Units (OTUs) or Amplicon Sequence Variants (ASVs) underpin many downstream bioinformatics analyses and ecological interpretations of environmental DNA (eDNA) datasets. However, public molecular databases (i.e., SILVA, EUKARYOME, BOLD) for most microbial metazoan phyla (nematodes, tardigrades, kinorhynchs, etc.) are sparsely populated, negatively impacting our ability to assign ecologically meaningful taxonomy to these understudied groups. Additionally, the choice of bioinformatics parameters and computational algorithms can further impact the accuracy of eDNA taxonomy assignments. Here, we use two in-silico datasets to show that taxonomy assignments using the 18S rRNA gene can be dramatically improved by curating Linnaean taxonomy strings associated with each reference sequence and closing phylogenetic gaps by improving taxon sampling. Using free-living nematodes as a case study, we applied two commonly used taxonomy assignment algorithms (BLAST+ and the QIIME2 Naïve Bayes classifier) across six iterations of the SILVA 138 reference database to evaluate the precision and accuracy of taxonomy assignments. The BLAST+ top hit with a 90% sequence similarity cutoff often returned the highest percentage of correctly assigned taxonomy at the genus level, and the QIIME2 Naïve Bayes classifier performed similarly well when paired with a reference database containing corrected taxonomy strings. Our results highlight the urgent need for phylogenetically-informed expansions of public reference databases (encompassing both genomes and common gene markers), focused on poorly sampled lineages which are now robustly recovered via eDNA metabarcoding approaches. Additional taxonomy curation efforts should be applied to popular reference databases such as SILVA, and taxon sampling could be rapidly improved by more frequent incorporation of newly published GenBank sequences linked to genus and/or species level identifications.

18S rRNA metabarcoding

A novel transformer model of protein domains for viral taxonomy classification.

MOTIVATION: Viruses with carefully curated taxonomic assignments (such as those in the ICTV taxonomy) still represent only a small fraction of viruses identified through sequencing data from virome or microbiome projects. It is therefore critical to develop methods that can assign viruses at multiple taxonomic ranks, so that a virus deemed novel at a given rank may still be placed into a higher-level taxon. Sequence-similarity-based approaches can classify viruses that share substantial genomic similarity with known viruses (e.g. those belonging to the same species or genus); however, their performance drops significantly when applied to more divergent viruses. Recent deep learning models, such as ViTax, which utilize DNA language models, aim to address these limitations, but their performance also degrades when applied to novel viruses lacking genus-level similarity to known references. Proteins are more conserved than genomic sequences, and the multiple proteins encoded by a virus can be leveraged to reveal evolutionary relationships among viruses. RESULTS: We propose a new tool, D2T (Domain-to-Taxonomy), that leverages recent advances in protein language models to improve viral taxonomic assignment. D2T represents a virus as a sequence of protein domain tokens and learns a transformer-based model for taxonomic classification. Experiments on multiple closed-set and open-set datasets show that D2T excels at assigning higher-level taxonomic labels (family and above). Furthermore, by combining D2T with Kraken2, which performs well at the genus level, the hybrid method (K+D2T) achieves accurate viral taxonomic classification across multiple taxonomic ranks. AVAILABILITY AND IMPLEMENTATION: D2T is available as a GitHub repository at https://github.com/mgtools/D2T.

Viruses

Comparison of whole-genome sequencing-based analysis methods for taxonomic classification of isolates unclassified by MALDI-TOF MS.

Taxonomic identification of clinical isolates is routinely achieved using matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF MS). If the species cannot be reliably identified, whole-genome sequencing can be applied. The aim of this study was to compare the results of approaches for taxonomic assignment for classification of isolates that are difficult to identify. Fifty-seven isolates were included in the study. The isolates were whole-genome sequenced and de novo assembled. Assembly-based classification was performed with the Genome Taxonomy Database Toolkit (GTDB-Tk), BLAST against 16S rRNA gene databases, the Type (Strain) Genome Server (TYGS), and ribosomal MLST (rMLST). Read-based classification was performed with MetaPhlAn4 and Kraken2. Thirty-two isolates were assigned to the same species with all four assembly-based classifiers, while the remaining 25 showed diverging assignments. When evaluating the results for the latter isolates, GTDB-Tk performed better than the other classifiers regarding which assignments were most likely correct. Of the read-based classifiers, MetaPhlAn4 performed better than Kraken2. Our evaluation identified GTDB-Tk to be the strongest tool for taxonomic assignment of isolates that are difficult to identify. Disagreements between classifiers are likely due to database limitations, wrongly assigned taxonomy, or unreliable 16S rRNA gene-based assignments.

Spectrometry, Mass, Matrix-Assisted Laser Desorpti

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100 000 and 1 000 000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

Two Saccharopolyspora isolates from archaeological excavation sites: polyphasic taxonomy, biosynthetic potential, bioactivity profiling and description of Saccharopolyspora antiqui sp. nov.

Archaeological excavation sites represent underexplored microbial habitats with the potential to recover taxonomically and biotechnologically valuable actinomycetes. In this study, two Saccharopolyspora strains, 5N708T and 5N102, were isolated from soil samples collected from the Gaziantep-Doliche-Dülük and Bitlis-Ahlat-Selçuklu Cemetery archaeological excavation sites in Türkiye. A polyphasic taxonomic approach, including 16S rRNA gene sequencing, phylogenetic and phylogenomic analyses, average nucleotide identity, digital DNA-DNA hybridization, phenotypic characterization, and chemotaxonomic analyses, showed that strain 5N708T represents a novel species of the genus Saccharopolyspora, for which the name Saccharopolyspora antiqui sp. nov. is proposed, whereas strain 5N102 was assigned to Saccharopolyspora elongata. Both isolates were further evaluated for their antimicrobial, antioxidant, and cytotoxic activities, and their biosynthetic potential was investigated by genome mining. Both strains showed activity against Staphylococcus aureus, with strain 5N708T producing the larger inhibition zone. Strain 5N102 exhibited markedly stronger antioxidant activity than strain 5N708T in radical scavenging, ferric reducing antioxidant power, and reducing power assays. In contrast, strain 5N708T showed more promising cytotoxic activity, with relative selectivity toward MIA PaCa-2 pancreatic cancer cells compared with HEK293 cells after prolonged incubation. Genome mining revealed multiple biosynthetic gene clusters in both isolates, supporting their capacity to produce secondary metabolites. These findings indicate that archaeological soils are promising reservoirs of taxonomically novel and biologically active Saccharopolyspora strains.

Saccharopolyspora

Microbiome Datahub: an open-access platform integrating environmental metadata, taxonomy, and functional annotation for comprehensive metagenome-assembled genome datasets.

BACKGROUND: Metagenome-assembled genomes (MAGs) provide crucial insights into the genomic diversity of uncultured microbes. However, MAG datasets deposited in public repositories such as INSDC are often difficult to reuse due to heterogeneous quality, inconsistent taxonomic and functional annotations, and insufficiently curated environmental metadata. While secondary MAG databases such as MGnify, IMG/M, and SPIRE provide standardized resources, they reconstruct MAGs de novo from public metagenomic reads and therefore do not represent the original MAGs reported in publications. RESULTS: To address this gap, we developed Microbiome Datahub, an open-access platform that systematically aggregates and re-annotates original MAGs from INSDC. We collected 214,427 MAGs, predicted genes by DFAST, performed quality assessment with CheckM, standardized taxonomic assignments with GTDB-Tk, inferred 27 phenotypic traits using Bac2Feature, assigned proteins to MBGD ortholog clusters and KEGG Orthology IDs using PZLAST, and annotated environmental metadata with the Metagenome and Microbes Environmental Ontology. Across these MAGs, the average completeness was 80.5% and contamination 1.8%; notably, the most frequent values were&#x2009;>95% completeness and&#x2009;<1% contamination, indicating that the majority of MAGs are of high quality. Comparative analyses showed that Microbiome Datahub provides phylogenetically and environmentally diverse MAGs: while the majority originated from vertebrate gut environments, a substantial number were also recovered from other habitats such as groundwater, including nearly 10,000 MAGs from the Patescibacteria. Inference of 27 phenotypic traits, including optimum growth temperature, further revealed ecological differentiation across phyla. Protein clustering revealed 56 million identity 40% clusters, with the majority unique compared with MGnify and GlobDB, and&#x2009;~19% of proteins unassigned to MBGD ortholog clusters, underscoring their novelty. CONCLUSIONS: Microbiome Datahub integrates MAG genome sequences, gene and protein predictions, quality metrics, environmental and taxonomic annotations, ortholog cluster assignments, and phenotype predictions, all accessible via a web interface, API, and bulk downloads. By combining original MAGs with curated metadata and functional annotations, Microbiome Datahub constitutes a comprehensive and reusable resource that will accelerate microbiome and microbial genomics research. Video Abstract.

Metagenome

Molecular identification and diversity assessment of Tyrrhenian Romulea species (Iridaceae).

Taxonomic assignments based only on morphology are often insufficient for delimiting species, particularly in complexes shaped by hybridization and polyploidy, where species boundaries are unclear. This limitation hinders progress in ecological, biogeographic and conservation research. The genus Romulea, distributed across Africa and the Mediterranean Basin, exemplifies this challenge. Despite its remarkable diversity, Mediterranean Romulea has not received much attention from genetic and molecular studies. Here, we present the first multilocus genotype analysis of Mediterranean Romulea taxa, focusing on the Tyrrhenian biogeographic province. Using target-capture sequencing with the universal Angiosperms353 kit, we generated genomic data for 272 individuals representing 18 putative taxa. Our findings reveal genetic groups that align with current taxonomy, the existence of cryptic divergence, and highlight the role of hybridization. Furthermore, analysis of intra-individual genetic diversity suggests one or several allopolyploid origins for Mediterranean Romulea. Four taxa (R. assumptionis, R. revelieri, R. ligustica, R. rollii) are consistently well differentiated across nuclear and plastid datasets, supporting their recognition as distinct species. In contrast, the widespread species R. ramiflora and R. columnae contain well-differentiated groups that may represent cryptic speciation. Several other taxa, including R. x melitensis, R. corsica, and R. bulbocodium, exhibit genomic signatures consistent with hybrid origins. Plastid and nuclear variation patterns are consistent with a hypothesis of rapid radiation in the Tyrrhenian region. These results provide a primary genomic framework for the integrative taxonomy of Romulea.

Genetic Variation

Is There a Fly in My Soup? To What Extent Do Metabarcoding and Individual Barcoding Tell the Same Story?

Metabarcoding has become the method of choice for characterizing complex arthropod communities. The extent to which metabarcoded bulk samples will recover the same community composition as individual sequencing of all individuals in the sample remains poorly quantified. Biases such as unequal extraction of DNA from different taxa, primer mismatches and non-random PCR may cause the selective drop-out of species from metabarcoding data. At the same time, DNA metabarcoding may reveal arthropod taxa present not as individuals, but as DNA residues on the surface or in the gut of insects. To quantify the consistency in sample contents established by different means, we metabarcoded 45 bulk insect samples, then extracted all arthropods and sequenced them individually. Metabarcoding targeted 418&#x2009;bp at the 3' end of the Folmer barcoding region, while individual barcodes captured the entire 658&#x2009;bp Folmer region. The metabarcoding workflow, including PCR amplification, sequencing and bioinformatics, was performed in three replicates from three separate lysate aliquots per sample. For the main analyses, sequences were assigned to Barcode Index Numbers (BINs) as identical taxonomic categories across data types, thereby allowing the detection of even rare but biologically true taxa. Since such reference-based validation will be unavailable to any researcher dealing with metabarcoding data alone, we validated our key findings through an alternative workflow, i.e., de novo clustering of sequences. We found that metabarcoding is replicable, as different replicates of the same sample recover similar species richness and composition. Individual barcoding and metabarcoding provide similar impressions of relative differences in community structure: species-rich vs. species-poor samples rank similarly among data types (Spearman's &#x2374;&#x2009;=&#x2009;0.88-0.99) as do differences in relative dissimilarity between sample pairs (Spearman's &#x2374;&#x2009;=&#x2009;0.55-0.90). Dissimilarity between data types varies with BIN richness in the sample, but this relationship reflects nestedness rather than turnover: metabarcoding recovers the same set of core species as individual barcoding but adds hundreds of species on top. Any BIN recovered as an individual occurred with high probability in the metabarcoding data, and any BIN found in high read abundances by metabarcoding was likely found as an individual (p&#x2009;>&#x2009;0.8). In terms of abundances, the number of individual insects per BIN was well predicted by the number of metabarcoding reads (R2&#x2009;>&#x2009;0.68 for a model including taxonomy as a random effect). Our analysis suggests that metabarcoding data will be informative of the sample contents in terms of arthropod species richness, composition and taxon-specific abundances. Taxa recovered in low copy numbers in metabarcoding sequence data will likely represent DNA left as residues from past biotic interactions. Barring sequencing errors, both types of data yield biologically relevant insights into the taxa present in the source community.

Animals

Phylogenomics of Desulfuromonadia supports reclassification of Geobacter psychrophilus as Irobacter psychrophilus comb. nov. and proposal of Geosyntrophus gen. nov.

Genome-resolved phylogenomics reveals widespread misclassification of metal-reducing bacteria historically assigned to Geobacter based on 16S rRNA gene phylogeny, and highlights species that persist only as 16S rRNA entries without genomes for robust taxonomic resolution. Here, we resolve two such lineages by integrating whole-genome phylogeny with average amino acid identity (AAI) and percentage of conserved proteins (POCP) across 418 dereplicated genomes of Desulfuromonadia. We report a draft genome of the psychrophilic iron-reducing bacterium Geobacter psychrophilus (100% completeness). Phylogenomic analyses place both Geobacter psychrophilus and the GTDB placeholder genus g__JACRCG01 within the family 'Pseudopelobacteraceae', outside Geobacteraceae sensu stricto. Within this framework, G. psychrophilus forms a distinct, well-supported lineage separated from neighbouring genera by discontinuities in AAI and POCP, supporting its reclassification as Irobacter psychrophilus comb. nov. Additionally, we show that Geosyntrophus acetoxidans, a non-axenic syntrophic bacterium, forms a coherent genus with 51 other environmental genomes (placeholder genus g__JACRCG01), for which we propose the replacement name Geosyntrophus gen. nov. Comparative genome analysis revealed conserved family-level metabolic traits together with genus-specific differences in respiratory metabolism, while ANI-based clustering identified substantial species-level diversity within both proposed genera. Metagenome and 16S rRNA-gene survey data further show that Geosyntrophus and Irobacter occur in broadly similar aquatic and subsurface habitats spanning from the Arctic to the Antarctic. Together, these results resolve the taxonomy of two previously ambiguous Desulfuromonadales lineages and shed light on their environmental distribution.

AAI

Genomic and polyphasic characterization of six novel Hymenobacter species isolated from soil in Korea.

Six bacterial strains (BT523T, BT559T, 15J16-1T3BT, BT730T, DG25AT, and DG25BT) were isolated from soil samples in Korea and assigned to the family Hymenobacteraceae (order Cytophagales, class Cytophagia). Phylogenetic analysis based on 16S rRNA gene sequences showed that the strains formed distinct lineages within the genus Hymenobacter. Strains BT523T, BT559T, and 15J16-1T3BT exhibited highest sequence similarities to Hymenobacter armeniacus BT189T (97.6%), Hymenobacter polaris RP-2-7&#xa0;T (97.8%), and Hymenobacter paludis KBP-30&#xa0;T (98.3%), respectively, while strains BT730T, DG25AT, and DG25BT were most closely related to Hymenobacter tibetensis XTM003T, with similarities of 96.3-96.6%. All strains were Gram-negative, aerobic, rod-shaped, and formed red to pink pigmented colonies. Whole-genome analysis revealed genome sizes ranging from 3.78 to 6.33&#xa0;Mb with G&#x2009;+&#x2009;C contents of 55.5-65.0%. Average nucleotide identity (ANI) and digital DNA-DNA hybridization (dDDH) values between the strains and their closest relatives were below the accepted thresholds for species delineation, supporting their classification as novel species. Functional annotation indicated the presence of genes associated with core metabolism, stress response, and pigment biosynthesis, reflecting adaptation to soil environments. Secondary metabolite analysis further revealed the presence of biosynthetic gene clusters, including terpene and siderophore pathways. Based on polyphasic taxonomic evidence, the six strains are proposed to represent six novel species of the genus Hymenobacter, for which the names Hymenobacter miniatus sp. nov., Hymenobacter madidus sp. nov., Hymenobacter convexus sp. nov., Hymenobacter rubellus sp. nov., Hymenobacter erythromyxa sp. nov., and Hymenobacter radioresistens sp. nov. are proposed. The type strains are BT523T (=&#x2009;KCTC 72341&#xa0;T&#x2009;=&#x2009;NBRC 114851&#xa0;T), BT559T (=&#x2009;KACC 21821&#xa0;T&#x2009;=&#x2009;NBRC 114852&#xa0;T), 15J16-1T3BT (=&#x2009;KCTC 42995&#xa0;T&#x2009;=&#x2009;NBRC 112818&#xa0;T), BT730T (=&#x2009;KACC 22457&#xa0;T&#x2009;=&#x2009;NBRC 116482&#xa0;T), DG25AT (=&#x2009;KCTC 32451&#xa0;T&#x2009;=&#x2009;TBRC 19856&#xa0;T), and DG25BT (=&#x2009;KCTC 32452&#xa0;T&#x2009;=&#x2009;JCM 19444&#xa0;T).

Soil Microbiology