PubMed HealthSearch

SEARCH · PubMed Health

Results for “Taxonomy”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Persistent gaps and errors in reference databases impede ecologically meaningful taxonomy assignments in 18S rRNA studies: a case study of terrestrial and marine nematodes.

In metabarcoding studies, Linnaean taxonomy assignments of Operational Taxonomic Units (OTUs) or Amplicon Sequence Variants (ASVs) underpin many downstream bioinformatics analyses and ecological interpretations of environmental DNA (eDNA) datasets. However, public molecular databases (i.e., SILVA, EUKARYOME, BOLD) for most microbial metazoan phyla (nematodes, tardigrades, kinorhynchs, etc.) are sparsely populated, negatively impacting our ability to assign ecologically meaningful taxonomy to these understudied groups. Additionally, the choice of bioinformatics parameters and computational algorithms can further impact the accuracy of eDNA taxonomy assignments. Here, we use two in-silico datasets to show that taxonomy assignments using the 18S rRNA gene can be dramatically improved by curating Linnaean taxonomy strings associated with each reference sequence and closing phylogenetic gaps by improving taxon sampling. Using free-living nematodes as a case study, we applied two commonly used taxonomy assignment algorithms (BLAST+ and the QIIME2 Naïve Bayes classifier) across six iterations of the SILVA 138 reference database to evaluate the precision and accuracy of taxonomy assignments. The BLAST+ top hit with a 90% sequence similarity cutoff often returned the highest percentage of correctly assigned taxonomy at the genus level, and the QIIME2 Naïve Bayes classifier performed similarly well when paired with a reference database containing corrected taxonomy strings. Our results highlight the urgent need for phylogenetically-informed expansions of public reference databases (encompassing both genomes and common gene markers), focused on poorly sampled lineages which are now robustly recovered via eDNA metabarcoding approaches. Additional taxonomy curation efforts should be applied to popular reference databases such as SILVA, and taxon sampling could be rapidly improved by more frequent incorporation of newly published GenBank sequences linked to genus and/or species level identifications.

18S rRNA metabarcoding

Programmatic access to ICTV virus taxonomy through a public ontology API.

BACKGROUND: The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. FINDINGS: To address this, we developed a public and sustainable solution leveraging ontology-based APIs. All available ICTV Master Species List (MSL) releases, from MSL1 to MSL41, were transformed into a unified, semantically structured ontology comprising more than 195,000 current and historical entities and deployed through the Ontology Lookup Service (OLS). The ontology is automatically rebuilt and republished whenever a new MSL release becomes available. Complementary ICTV-NCBI mappings and helper libraries support integration into downstream systems. CONCLUSIONS: Together, these resources enable, for the first time, public programmatic retrieval of current and historical ICTV taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints, including resolution of former taxonomic terms to their current accepted taxon or taxa and retrieval of taxon histories across releases. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API

Programmatic access to ICTV virus taxonomy through a public ontology API.

The International Committee on Taxonomy of Viruses (ICTV) is responsible for developing and maintaining a universal virus taxonomy. As the reference framework for organising the viral world, it is essential for virology and related fields. Despite its widespread use in research and public health, programmatic access to ICTV taxonomy has remained limited, posing challenges for integration, versioning, and interoperability across databases and bioinformatics resources requiring up-to-date virus taxonomy. To address this, we developed a public and sustainable solution leveraging ontology-based APIs. Successive ICTV Master Species List (MSL) releases were transformed into a structured ontology and deployed as a unified representation through the Ontology Lookup Service (OLS). The framework also provides ICTV-NCBI mappings and helper libraries for integration into downstream systems. This enables, for the first time, public programmatic retrieval of current and historical virological taxon names, taxonomic relationships, metadata, and persistent identifiers through stable endpoints. More broadly, this work illustrates a general strategy for transforming structured biological datasets into semantically enriched graph resources exposed through scalable public APIs. These developments enhance interoperability, reduce manual curation, and support FAIR-aligned taxonomic data management in virology and pandemic preparedness.

API

A novel transformer model of protein domains for viral taxonomy classification.

MOTIVATION: Viruses with carefully curated taxonomic assignments (such as those in the ICTV taxonomy) still represent only a small fraction of viruses identified through sequencing data from virome or microbiome projects. It is therefore critical to develop methods that can assign viruses at multiple taxonomic ranks, so that a virus deemed novel at a given rank may still be placed into a higher-level taxon. Sequence-similarity-based approaches can classify viruses that share substantial genomic similarity with known viruses (e.g. those belonging to the same species or genus); however, their performance drops significantly when applied to more divergent viruses. Recent deep learning models, such as ViTax, which utilize DNA language models, aim to address these limitations, but their performance also degrades when applied to novel viruses lacking genus-level similarity to known references. Proteins are more conserved than genomic sequences, and the multiple proteins encoded by a virus can be leveraged to reveal evolutionary relationships among viruses. RESULTS: We propose a new tool, D2T (Domain-to-Taxonomy), that leverages recent advances in protein language models to improve viral taxonomic assignment. D2T represents a virus as a sequence of protein domain tokens and learns a transformer-based model for taxonomic classification. Experiments on multiple closed-set and open-set datasets show that D2T excels at assigning higher-level taxonomic labels (family and above). Furthermore, by combining D2T with Kraken2, which performs well at the genus level, the hybrid method (K+D2T) achieves accurate viral taxonomic classification across multiple taxonomic ranks. AVAILABILITY AND IMPLEMENTATION: D2T is available as a GitHub repository at https://github.com/mgtools/D2T.

Viruses

Stage-Independent Real-Time Subtype Classification and Comprehensive Biopsy Profiling of Urothelial Carcinomas by the Lund Taxonomy System.

Bladder cancer is a heterogeneous malignancy with diverse clinical outcomes, and conventional pathological assessment alone is insufficient to capture its underlying biology. Gene expression profiling can stratify tumors into molecular subtypes with prognostic and predictive potential, but the reliability of transcriptomic classification and its clinical utility remains to be established. The translational/observational UROSCANSEQ study (ISRCTN15459149) prospectively evaluates RNA-based Lund Taxonomy (LundTax) molecular subtype classification in a clinical setting. Among 784 consecutive biopsies collected between 2018 and 2022, RNA sequencing was successful for 90% of all biopsies, encompassing 662 bladder cancer patients with a stage distribution of 48% Ta, 27% T1, 24% ≥T2, and 1% CIS. We demonstrate that the LundTax subtype classification algorithm, applied to individual samples, accurately identifies cancer cell phenotypes with characteristic gene and protein expression patterns in a manner robust to RNA quality, data preprocessing strategies, and batch effects, supporting its clinical feasibility across both non-muscle-invasive and muscle-invasive disease. We further extend the LundTax framework by incorporating single-sample molecular risk scores reflecting tumor grade, proliferation, and progression risk, as well as tumor microenvironment signatures. Both risk scores and overall immune and stromal content in biopsies were significantly associated with an increased risk of clinical progression in noninvasive disease. In a separate analysis of the relative cellular composition of the tumor microenvironment, however, only the fraction of natural killer cells remained significant. Together, the expanded LundTax system provides a comprehensive molecular portrait of individual tumor biopsies. By explicitly separating cancer cell-intrinsic phenotypes, prognostic indexes, and microenvironmental signals, the framework minimizes biological confounding and establishes a strong foundation for future studies evaluating clinical outcomes and treatment responses.

Humans

The landscape of pruning for large language models: A systematic review and unified taxonomy.

Confronting the inherent tension between the exceptional capabilities and the immense computational costs of Large Language Models (LLMs), pruning has become a crucial technique for achieving efficient deployment. However, a systematic analytical framework dedicated specifically to LLM pruning remains absent. In this paper, we aim to bridge this gap. We first elucidate the theoretical foundations that underpin the effectiveness of pruning, namely overparameterization and redundancy, and then propose a multidimensional taxonomy that organizes existing approaches along the axes of granularity, timing, and criteria. Building upon this unified perspective, we further analyze performance recovery mechanisms and the broader evaluation ecosystem, while also exploring forward-looking challenges such as interpretability, automation, and hardware-algorithm co-design. Through this comprehensive synthesis, we seek to provide an integrated and coherent analytical lens for advancing both research and practice in LLM pruning.

Large Language Models

Up-to-date, and taxonomy-curated mcrA reference databases for methanogen community profiling.

The methyl-coenzyme M reductase subunit alpha gene (mcrA) is an important phylogenetic marker for high throughput ecological profiling of methanogenic archaea, central to industrial biological methane production and greenhouse gas emissions. Yet, dedicated reference databases predate current relevant NCBI sequence accumulation and archaeal taxonomic revision. We present three updated mcrA reference databases: (i) one derived from NCBI-catalogued methanogen genomes (1572 sequences); (ii) a database built by expansion of a previously published reference dataset, leveraging the NCBI nucleotide collection (27,942 sequences); (iii) a curated-taxonomy version of the latter. The updated amplicon databases provide a ∼ 3.5-fold sequence richness expansion, extend genus-level richness from 31 to 83 taxa, more than 4-fold species-level richness, and incorporate novel lineages compared with the previous reference dataset (e.g. Thermoplasmatota-encompassed). All databases were formatted to support analysis with relevant contemporary software pipelines and packages. Overall, the generated databases facilitate a highly improved characterization of methanogen diversity and ecology.

Archaea

Chromosome-Level Genome Assembly of Eden's Whale Clarifies the Taxonomy and Speciation of Bryde's Whale Complex.

Eden's whale (Balaenoptera edeni), a poorly understood baleen cetacean, has long been shrouded in taxonomic ambiguity due to limited genomic resources, obscuring its distinction from closely related species and its position within the cetacean Tree of Life. In this paper, we present a high-quality chromosomal-level genome of B. edeni and conduct comparative genomic analyses to address long-standing taxonomic confusion and elucidate speciation of balaenopterids. Our phylogenomic analysis and demographic reconstruction reveal that B. edeni is a distinct sister to Bryde's whale (Balaenoptera brydei), sharing a common ancestor that diverged approximately 7.84 million years ago during the late Miocene. Their genetic divergence exceeds typical intraspecific variation in whales, supporting the reinstatement of B. brydei as a valid species. Chromosomal syntenic analyses suggest that macro-fragment inversions contributed to speciation in balaenopterid whales and uncover unexpected large-scale complex genome rearrangements in Bryde's whale, offering novel insights into cetacean genome evolution. Functional enrichment analysis of inverted regions between B. edeni and Balaenoptera musculus indicates their predominant association with metabolism and biosynthesis, as well as responses to various substances, stress, and stimuli. These genomic resources for B. edeni not only lay a critical foundation for comparative genetic and evolutionary research of cetaceans but also advance our understanding of the taxonomy and evolutionary dynamics of the Bryde's whale complex, with broader implications for baleen whale conservation and biodiversity.

Animals

Stenotrophomonas maltophilia in the Antimicrobial Resistance Era: Species-Complex Taxonomy, Pathogenesis, Evolving Therapeutic Priorities, and Genomic Surveillance.

Stenotrophomonas maltophilia is a globally distributed, aerobic, non-fermenting Gram-negative bacillus increasingly recognized as an opportunistic pathogen in hospitalized and immunocompromised patients. Clinical interpretation is challenging because respiratory and device-associated isolates may represent colonization, polymicrobial infection, or true invasive disease. Recent genomic studies further suggest that organisms historically identified as S. maltophilia comprise a genetically diverse species complex, with implications for epidemiology, virulence, resistance surveillance, and susceptibility testing. Treatment is difficult because of biofilm formation, persistence in water-associated healthcare reservoirs, and intrinsic or acquired resistance mediated by L1 and L2 β-lactamases, multidrug efflux pumps, reduced permeability, mobile resistance determinants, and biofilm-associated tolerance. Current IDSA guidance identifies cefiderocol monotherapy as the preferred treatment for invasive S. maltophilia infection, whereas aztreonam-avibactam and agents such as trimethoprim-sulfamethoxazole, levofloxacin, and minocycline occupy alternative or combination-based roles. Nevertheless, the therapeutic evidence base remains uneven, and clinical decisions should integrate infection severity, source control, susceptibility findings, pharmacokinetic/pharmacodynamic (PK/PD) exposure, toxicity, infection site, and host-related factors. This review summarizes advances in taxonomy, epidemiology, pathogenesis, diagnostics, resistance, treatment, infection prevention, and genomic surveillance, and highlights the need for standardized identification, validated breakpoints, prospective comparative-effectiveness studies, and pragmatic or adaptive trial designs.

L1 β-lactamase

[Immunological characteristics of the protein antigens of the family Neisseriaceae. II. The importance of an immunochemical analysis of the protein complexes for a study of taxonomy problems].

The study of antigenic interrelations in the family Neisseriaceae resulted in the isolation of 2 main immunologically separated variants of protein complexes: the first variant was characteristic of nonpathogenic and pathogenic species of the genus Neisseria as well as of 7 taxonomically undefined Neisseria species (N. lactamicus, N. cuniculi, N. ellongata, N. ovis, N. animalis, N. cinerea, N. canis) and Gemella haemolysans; the second variant was represented by the genera Branhamella and Acinetobacter. N. caviae and 3 out of 14 Neisseria strains of undefined species had no common antigens with the genera Neisseria and Branhamella. The importance of the immunotyping of protein complexes for studying the problems connected with the taxonomy of the family Neisseriaceae was considered.

Antigens, Bacterial

Arenavirus taxonomy: a review.

Despite a late beginning, the construction of the arenavirus taxon and its placement in the scheme of the International Committee on Taxonomy of Viruses has now been completed. The bringing together of the member viruses has already provided valuable indications of promising laboratory and field study approaches; in the future this classification will contribute further to our understanding of the natural history and disease processes of the human pathogens of the group.

Arboviruses

Two Saccharopolyspora isolates from archaeological excavation sites: polyphasic taxonomy, biosynthetic potential, bioactivity profiling and description of Saccharopolyspora antiqui sp. nov.

Archaeological excavation sites represent underexplored microbial habitats with the potential to recover taxonomically and biotechnologically valuable actinomycetes. In this study, two Saccharopolyspora strains, 5N708T and 5N102, were isolated from soil samples collected from the Gaziantep-Doliche-Dülük and Bitlis-Ahlat-Selçuklu Cemetery archaeological excavation sites in Türkiye. A polyphasic taxonomic approach, including 16S rRNA gene sequencing, phylogenetic and phylogenomic analyses, average nucleotide identity, digital DNA-DNA hybridization, phenotypic characterization, and chemotaxonomic analyses, showed that strain 5N708T represents a novel species of the genus Saccharopolyspora, for which the name Saccharopolyspora antiqui sp. nov. is proposed, whereas strain 5N102 was assigned to Saccharopolyspora elongata. Both isolates were further evaluated for their antimicrobial, antioxidant, and cytotoxic activities, and their biosynthetic potential was investigated by genome mining. Both strains showed activity against Staphylococcus aureus, with strain 5N708T producing the larger inhibition zone. Strain 5N102 exhibited markedly stronger antioxidant activity than strain 5N708T in radical scavenging, ferric reducing antioxidant power, and reducing power assays. In contrast, strain 5N708T showed more promising cytotoxic activity, with relative selectivity toward MIA PaCa-2 pancreatic cancer cells compared with HEK293 cells after prolonged incubation. Genome mining revealed multiple biosynthetic gene clusters in both isolates, supporting their capacity to produce secondary metabolites. These findings indicate that archaeological soils are promising reservoirs of taxonomically novel and biologically active Saccharopolyspora strains.

Saccharopolyspora

Revisiting Papillomavirus Taxonomy: A Proposal for Updating the Current Classification in Line with Evolutionary Evidence.

Papillomaviruses infect a wide array of animal hosts and are responsible for roughly 5% of all human cancers. Comparative genomics between different virus types belonging to specific taxonomic groupings (e.g., species, and genera) has the potential to illuminate physiological differences between viruses with different biological outcomes. Likewise, extrapolation of features between related viruses can be very powerful but requires a solid foundation supporting the evolutionary relationships between viruses. The current papillomavirus classification system is based on pairwise sequence identity. However, with the advent of metagenomics as facilitated by high-throughput sequencing and molecular tools of enriching circular DNA molecules using rolling circle amplification, there has been a dramatic increase in the described diversity of this viral family. Not surprisingly, this resulted in a dramatic increase in absolute number of viral types (i.e., sequences sharing <90% L1 gene pairwise identity). Many of these novel viruses are the sole member of a novel species within a novel genus (i.e., singletons), highlighting that we have only scratched the surface of papillomavirus diversity. I will discuss how this increase in observed sequence diversity complicates papillomavirus classification. I will propose a potential solution to these issues by explicitly basing the species and genera classification on the evolutionary history of these viruses based on the core viral proteins (E1, E2, and L1) of papillomaviruses. This strategy means that it is possible that a virus identified as the closest neighbor based on the E1, E2, L1 phylogenetic tree, is not the closest neighbor based on L1 nucleotide identity. In this case, I propose that a virus would be considered a novel type if it shares less than 90% identity with its closest neighbors in the E1, E2, L1 phylogenetic tree.

Animals

Taxonomy of the Clostridia: ribosomal ribonucleic acid homologies among the species.

rRNA homologies have been determined on reference strains representing 56 species of Clostridium. Competition experiments using tritium-labelled 23S rRNA were employed. The majority of the species had DNA with 27 to 28% guanine plus cytosine (%GC). These fell into rRNA homology groups I and II, which were well defined, and a third group which consisted of species which did not belong in groups I and II. Species whose DNA was 41 to 45% GC comprised a fourth group. Thirty species were placed into rRNA homology group I on the basis of having 50% or greater homology with Clostridium butyricum, C. perfringens, C. carnis, C. sporogenes, C. novyi or C. pasteurianum. Ten subgroups were delineated in homology group I. Species in each subgroup either had high homology with a particular reference species or a similar pattern of homologies to all of the reference organisms. The eleven species in rRNA homology group II had 69% or greater homology to C. lituseburense. Species in groups I and II had intergroup homologies of 20 to 40%. The six species in group II had very low homologies with groups I and II. Negligible homology also resulted when five of the species were tested against the sixth, C. ramosum. The five species having DNA with 41 to 45% GC were C. innocuum, C. sphenoides, C. indolis, C. barkeri and C. orotic um. Little rRNA homology was apparent between C. innocuum and the other high % GC species or with several Bacillus species having similar %GC DNA. Correlations between homology results and phenotypic characteristics are discussed.

Base Sequence

Characterization of microbial dark matter at scale with MetaSBT and taxonomy-aware Sequence Bloom Trees.

Metagenomics has become a powerful tool for studying microbial communities, allowing researchers to investigate microbial diversity within complex environmental samples. Recent advances in sequencing technology have enabled the recovery of near-complete microbial genomes directly from metagenomic samples, also known as metagenome-assembled genomes (MAGs). However, accurately characterizing these genomes remains a significant challenge due to the presence of sequencing errors, incomplete assembly, and contamination. Here we present MetaSBT, a new tool for organizing, indexing, and characterizing microbial reference genomes and MAGs. It is able to identify clusters of genomes at all seven taxonomic levels, from the kingdom all the way down to the species level, using the Sequence Bloom Tree (SBT) data structure that relies on Bloom Filters (BFs) to index massive amounts of genomes based on their k-mers composition. We have built an initial set of databases composed of over 190 thousand viral genomes from NCBI GenBank and public sources grouped into sequence consistent clusters at different taxonomic levels, making it the first software solution for the classification of viruses at different ranks, including still unknown ones. This results in the definition of over 40 thousand species clusters where ~80% do not match with any known viral species in reference databases to date. Furthermore, we show how our databases can be used as a new basis for existing quantitative metagenomic profilers to unlock the detection of unknown microbes and the estimation of their abundance in metagenomic samples. Finally, the framework is released open-source and, along with its public databases, is fully integrated into the Galaxy Platform enabling broad accessibility.

metagenome-assembled genomes

Suid evolution and correlation of African hominid localities: an alternative taxonomy.

New phylogenies were recently proposed by White and Harris, who recognized 7 genera and 16 species of fossil and extant suids from sub-Saharan Africa. This scheme is regarded here as oversimplified and an alternative is suggested, in which 9 genera and 21 species are recognized. The taxonomic and phylogenetic differences do not have any significant effect on the stratigraphic interpretations offered by White and Harris.

Africa