PubMed HealthSearch

SEARCH · PubMed Health

Results for “reference databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Persistent gaps and errors in reference databases impede ecologically meaningful taxonomy assignments in 18S rRNA studies: a case study of terrestrial and marine nematodes.

In metabarcoding studies, Linnaean taxonomy assignments of Operational Taxonomic Units (OTUs) or Amplicon Sequence Variants (ASVs) underpin many downstream bioinformatics analyses and ecological interpretations of environmental DNA (eDNA) datasets. However, public molecular databases (i.e., SILVA, EUKARYOME, BOLD) for most microbial metazoan phyla (nematodes, tardigrades, kinorhynchs, etc.) are sparsely populated, negatively impacting our ability to assign ecologically meaningful taxonomy to these understudied groups. Additionally, the choice of bioinformatics parameters and computational algorithms can further impact the accuracy of eDNA taxonomy assignments. Here, we use two in-silico datasets to show that taxonomy assignments using the 18S rRNA gene can be dramatically improved by curating Linnaean taxonomy strings associated with each reference sequence and closing phylogenetic gaps by improving taxon sampling. Using free-living nematodes as a case study, we applied two commonly used taxonomy assignment algorithms (BLAST+ and the QIIME2 Naïve Bayes classifier) across six iterations of the SILVA 138 reference database to evaluate the precision and accuracy of taxonomy assignments. The BLAST+ top hit with a 90% sequence similarity cutoff often returned the highest percentage of correctly assigned taxonomy at the genus level, and the QIIME2 Naïve Bayes classifier performed similarly well when paired with a reference database containing corrected taxonomy strings. Our results highlight the urgent need for phylogenetically-informed expansions of public reference databases (encompassing both genomes and common gene markers), focused on poorly sampled lineages which are now robustly recovered via eDNA metabarcoding approaches. Additional taxonomy curation efforts should be applied to popular reference databases such as SILVA, and taxon sampling could be rapidly improved by more frequent incorporation of newly published GenBank sequences linked to genus and/or species level identifications.

18S rRNA metabarcoding

Up-to-date, and taxonomy-curated mcrA reference databases for methanogen community profiling.

The methyl-coenzyme M reductase subunit alpha gene (mcrA) is an important phylogenetic marker for high throughput ecological profiling of methanogenic archaea, central to industrial biological methane production and greenhouse gas emissions. Yet, dedicated reference databases predate current relevant NCBI sequence accumulation and archaeal taxonomic revision. We present three updated mcrA reference databases: (i) one derived from NCBI-catalogued methanogen genomes (1572 sequences); (ii) a database built by expansion of a previously published reference dataset, leveraging the NCBI nucleotide collection (27,942 sequences); (iii) a curated-taxonomy version of the latter. The updated amplicon databases provide a ∼ 3.5-fold sequence richness expansion, extend genus-level richness from 31 to 83 taxa, more than 4-fold species-level richness, and incorporate novel lineages compared with the previous reference dataset (e.g. Thermoplasmatota-encompassed). All databases were formatted to support analysis with relevant contemporary software pipelines and packages. Overall, the generated databases facilitate a highly improved characterization of methanogen diversity and ecology.

Archaea

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software

Human liver protein map: a reference database established by microsequencing and gel comparison.

This publication establishes a reference human liver protein map obtained with immobilized pH gradients. By microsequencing, 57 spots or 42 polypeptide chains were identified. By protein map comparison and matching (liver, red blood cell and plasma sample maps), 8 additional proteins were identified. The new polypeptides and previously known proteins are listed in a table and/or labeled on the protein map, thus providing a human liver two-dimensional gel database. This reference map can be used to identify protein spots on other samples such as rectal cancer biopsies.

Amino Acid Sequence

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100 000 and 1 000 000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

The development of a multicenter database for reference values in clinical neurophysiology--principles and examples.

This paper describes the work undertaken to establish principles for the development of multicenter databases for reference values in clinical neurophysiology. The study was initiated because of interest of the involved laboratories in knowledge-based systems in electromyographic diagnosis, for which it was necessary to formalize the key concepts in the diagnostic process: diseases, pathophysiology and test results. The paper deals specifically with the structuring of results of motor and sensory nerve conduction studies.

Action Potentials

MegaPX: fast and space-efficient peptide assignment method using IBF-based multi-indexing.

MOTIVATION: A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the analyzed microbiomes. Using a database search, the standard approach requires prior knowledge of which proteins and taxa to include in the protein reference database or to use tailored metagenome-derived databases, which are expensive and error-prone in their generation. A possible strategy to circumvent this database search issue is de novo sequencing, where peptide sequences are directly identified from mass spectra. However, these sequences must still be mapped back to potentially extensive databases. Here, alignment-based approaches enable robust and precise results, with the potential drawback of high memory usage and long run times. RESULTS: We present MegaPX, a software for rapidly classifying de novo peptide sequences against large protein databases. MegaPX implemented as a C++-based tool, uses an alignment-free, k-mer approach as a taxonomic classification method with the possibility of generating mutated reference databases for error-tolerant searching. It uses various algorithms, including interleaved Bloom filters, to efficiently compute approximate membership queries, ensuring fast processing times while querying and indexing large databases in a multi-indexing fashion. We demonstrate the potential of MegaPX by analyzing different samples, including metaproteomics, against extensive reference databases, highlighting its use as a fast screening tool.

Software

Medical Facts File: a self service database of reference information.

The Dahlgren Memorial Library, Georgetown University Medical Center, will demonstrate Medical Facts File, a newly developed in-house database of general medical information. The file content emerged from the library's experience with commonly asked reference questions and the need to develop a database as an online source for users seeking quick answers to medical queries. Medical Facts File joins a growing family of over 18 databases which comprise Georgetown's IAIMS Knowledge Network. Use scenarios will demonstrate how an online search is initiated, either directly or as a prompt from one of the other online databases. The design of Medical Facts File at the Dahlgren Memorial Library began in late 1989 with a publishing section on instructions for authors planning to submit manuscripts to a variety of prominent medical journals. Since then, seven sections have been identified for the database. Three sections are highlighted for presentation, although work on the project is on-going. Medical Facts File is an easy-to-use, time saving system that facilitates tedious searching through a multitude of library sources. It provides users with a self-service, information look-up system.

Databases, Factual

Extensive Analysis of Genetic Diversity in HLA-DMA, HLA-DMB, HLA-DOA and HLA-DOB: Characterisation of 236 Novel Alleles.

HLA-DMA, -DMB, -DOA and -DOB are non-classical HLA Class II genes that play a crucial role in the selection of highly stable HLA Class II/peptide complexes on antigen-presenting cells. Although the genes were initially thought to have a limited diversity with less than 13 alleles per gene documented in the IPD-IMGT/HLA Database in 2022, recent studies suggest a potential impact of certain alleles on the outcome of hematopoietic cell transplantation. To gain a deeper understanding of allelic diversity, we sequenced HLA-DMA, -DMB, -DOA and -DOB of 1880 potential stem cell donors from Germany, Poland, Great Britain and Chile, achieving full-gene resolution. Remarkably, we identified 3968 previously undescribed sequences, including 28 distinct novel proteins. The observed allele frequencies were consistent across all studied populations with one dominating protein for each gene: HLA-DMA*01:01 (> 77%), HLA-DMB*01:01 (> 63%), HLA-DOA*01:01 (> 97%) and HLA-DOB*01:01 (> 77%). Notably, a much higher diversity was observed in full-genomic resolution. Finally, we submitted 51 distinct novel sequences for HLA-DMA, 58 for HLA-DMB, 80 for HLA-DOA and 47 for HLA-DOB to the IPD-IMGT/HLA Database. This comprehensive reference database update will not only simplify future genotyping of HLA-DMA, -DMB, -DOA and -DOB but will hopefully also enhance our understanding of the complex process of peptide selection and loading to the HLA Class II proteins.

Humans

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

Effectiveness of mass spectrometry and genomic analysis in the surveillance of nontuberculous Mycobacterium in Taiwan.

Nontuberculous mycobacteria (NTM) are diverse, and species-level identification remains challenging in routine diagnostics. We analyzed NTM isolates collected at three regional centers of the National Taiwan University Hospital (NTUH) from 2019 to 2024 to assess geographic variation and identification performance after implementation of matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS). Among 3,188 cases meeting the microbiological criteria for probable pulmonary NTM disease, the species distribution differed by region: Mycobacterium avium complex predominated in central Taiwan (Yunlin, 47.3%), whereas M. abscessus complex (Taipei, 26.5%) and M. kansasii (Hsinchu, 12.4%) were more common in northern Taiwan. In 2019, 14.5% of isolates were reported to be unidentified by MALDI-TOF MS; with workflow optimization and database updates, this percentage decreased but plateaued at 4.5-4.8%. Whole-genome sequencing (WGS) of 61 randomly selected persistently unidentified isolates revealed eight average nucleotide identity (ANI)-defined clusters; 55 isolates (90.2%) could not be assigned to known species using current reference databases. Two clusters detected only in Hsinchu were phylogenetically closest to M. kyorinense, with ANI values below the species demarcation threshold. Overall, we observed marked regional heterogeneity of NTM in Taiwan and a persistent identification gap that remained after MALDI-TOF MS optimization and follow-up WGS.IMPORTANCEThis study characterized regional differences in the NTM species distribution across Taiwan, and the results highlight the limitations of current identification approaches. MALDI-TOF MS identifies most isolates, but locally circulating lineages represent a persistent gap in global reference libraries. Even with whole-genome sequencing (WGS), 90.2% (55/61) of persistently unresolved isolates could not be assigned to known species in the current reference databases despite the formation of clear ANI- and phylogeny-defined clusters. These findings show that both proteomic and genomic reference resources for clinical NTM remain incomplete. Expanding regionally representative databases and performing WGS for isolates that remain unresolved by MALDI-TOF MS will be necessary to improve species-level resolution for surveillance and clinical interpretation.

Taiwan