PubMed HealthSearch

SEARCH · PubMed Health

Results for “Reference database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Persistent gaps and errors in reference databases impede ecologically meaningful taxonomy assignments in 18S rRNA studies: a case study of terrestrial and marine nematodes.

In metabarcoding studies, Linnaean taxonomy assignments of Operational Taxonomic Units (OTUs) or Amplicon Sequence Variants (ASVs) underpin many downstream bioinformatics analyses and ecological interpretations of environmental DNA (eDNA) datasets. However, public molecular databases (i.e., SILVA, EUKARYOME, BOLD) for most microbial metazoan phyla (nematodes, tardigrades, kinorhynchs, etc.) are sparsely populated, negatively impacting our ability to assign ecologically meaningful taxonomy to these understudied groups. Additionally, the choice of bioinformatics parameters and computational algorithms can further impact the accuracy of eDNA taxonomy assignments. Here, we use two in-silico datasets to show that taxonomy assignments using the 18S rRNA gene can be dramatically improved by curating Linnaean taxonomy strings associated with each reference sequence and closing phylogenetic gaps by improving taxon sampling. Using free-living nematodes as a case study, we applied two commonly used taxonomy assignment algorithms (BLAST+ and the QIIME2 Naïve Bayes classifier) across six iterations of the SILVA 138 reference database to evaluate the precision and accuracy of taxonomy assignments. The BLAST+ top hit with a 90% sequence similarity cutoff often returned the highest percentage of correctly assigned taxonomy at the genus level, and the QIIME2 Naïve Bayes classifier performed similarly well when paired with a reference database containing corrected taxonomy strings. Our results highlight the urgent need for phylogenetically-informed expansions of public reference databases (encompassing both genomes and common gene markers), focused on poorly sampled lineages which are now robustly recovered via eDNA metabarcoding approaches. Additional taxonomy curation efforts should be applied to popular reference databases such as SILVA, and taxon sampling could be rapidly improved by more frequent incorporation of newly published GenBank sequences linked to genus and/or species level identifications.

18S rRNA metabarcoding

Up-to-date, and taxonomy-curated mcrA reference databases for methanogen community profiling.

The methyl-coenzyme M reductase subunit alpha gene (mcrA) is an important phylogenetic marker for high throughput ecological profiling of methanogenic archaea, central to industrial biological methane production and greenhouse gas emissions. Yet, dedicated reference databases predate current relevant NCBI sequence accumulation and archaeal taxonomic revision. We present three updated mcrA reference databases: (i) one derived from NCBI-catalogued methanogen genomes (1572 sequences); (ii) a database built by expansion of a previously published reference dataset, leveraging the NCBI nucleotide collection (27,942 sequences); (iii) a curated-taxonomy version of the latter. The updated amplicon databases provide a ∼ 3.5-fold sequence richness expansion, extend genus-level richness from 31 to 83 taxa, more than 4-fold species-level richness, and incorporate novel lineages compared with the previous reference dataset (e.g. Thermoplasmatota-encompassed). All databases were formatted to support analysis with relevant contemporary software pipelines and packages. Overall, the generated databases facilitate a highly improved characterization of methanogen diversity and ecology.

Archaea

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software

Human liver protein map: a reference database established by microsequencing and gel comparison.

This publication establishes a reference human liver protein map obtained with immobilized pH gradients. By microsequencing, 57 spots or 42 polypeptide chains were identified. By protein map comparison and matching (liver, red blood cell and plasma sample maps), 8 additional proteins were identified. The new polypeptides and previously known proteins are listed in a table and/or labeled on the protein map, thus providing a human liver two-dimensional gel database. This reference map can be used to identify protein spots on other samples such as rectal cancer biopsies.

Amino Acid Sequence

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100 000 and 1 000 000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

MegaPX: fast and space-efficient peptide assignment method using IBF-based multi-indexing.

MOTIVATION: A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the analyzed microbiomes. Using a database search, the standard approach requires prior knowledge of which proteins and taxa to include in the protein reference database or to use tailored metagenome-derived databases, which are expensive and error-prone in their generation. A possible strategy to circumvent this database search issue is de novo sequencing, where peptide sequences are directly identified from mass spectra. However, these sequences must still be mapped back to potentially extensive databases. Here, alignment-based approaches enable robust and precise results, with the potential drawback of high memory usage and long run times. RESULTS: We present MegaPX, a software for rapidly classifying de novo peptide sequences against large protein databases. MegaPX implemented as a C++-based tool, uses an alignment-free, k-mer approach as a taxonomic classification method with the possibility of generating mutated reference databases for error-tolerant searching. It uses various algorithms, including interleaved Bloom filters, to efficiently compute approximate membership queries, ensuring fast processing times while querying and indexing large databases in a multi-indexing fashion. We demonstrate the potential of MegaPX by analyzing different samples, including metaproteomics, against extensive reference databases, highlighting its use as a fast screening tool.

Software

Extensive Analysis of Genetic Diversity in HLA-DMA, HLA-DMB, HLA-DOA and HLA-DOB: Characterisation of 236 Novel Alleles.

HLA-DMA, -DMB, -DOA and -DOB are non-classical HLA Class II genes that play a crucial role in the selection of highly stable HLA Class II/peptide complexes on antigen-presenting cells. Although the genes were initially thought to have a limited diversity with less than 13 alleles per gene documented in the IPD-IMGT/HLA Database in 2022, recent studies suggest a potential impact of certain alleles on the outcome of hematopoietic cell transplantation. To gain a deeper understanding of allelic diversity, we sequenced HLA-DMA, -DMB, -DOA and -DOB of 1880 potential stem cell donors from Germany, Poland, Great Britain and Chile, achieving full-gene resolution. Remarkably, we identified 3968 previously undescribed sequences, including 28 distinct novel proteins. The observed allele frequencies were consistent across all studied populations with one dominating protein for each gene: HLA-DMA*01:01 (> 77%), HLA-DMB*01:01 (> 63%), HLA-DOA*01:01 (> 97%) and HLA-DOB*01:01 (> 77%). Notably, a much higher diversity was observed in full-genomic resolution. Finally, we submitted 51 distinct novel sequences for HLA-DMA, 58 for HLA-DMB, 80 for HLA-DOA and 47 for HLA-DOB to the IPD-IMGT/HLA Database. This comprehensive reference database update will not only simplify future genotyping of HLA-DMA, -DMB, -DOA and -DOB but will hopefully also enhance our understanding of the complex process of peptide selection and loading to the HLA Class II proteins.

Humans

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

Effectiveness of mass spectrometry and genomic analysis in the surveillance of nontuberculous Mycobacterium in Taiwan.

Nontuberculous mycobacteria (NTM) are diverse, and species-level identification remains challenging in routine diagnostics. We analyzed NTM isolates collected at three regional centers of the National Taiwan University Hospital (NTUH) from 2019 to 2024 to assess geographic variation and identification performance after implementation of matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS). Among 3,188 cases meeting the microbiological criteria for probable pulmonary NTM disease, the species distribution differed by region: Mycobacterium avium complex predominated in central Taiwan (Yunlin, 47.3%), whereas M. abscessus complex (Taipei, 26.5%) and M. kansasii (Hsinchu, 12.4%) were more common in northern Taiwan. In 2019, 14.5% of isolates were reported to be unidentified by MALDI-TOF MS; with workflow optimization and database updates, this percentage decreased but plateaued at 4.5-4.8%. Whole-genome sequencing (WGS) of 61 randomly selected persistently unidentified isolates revealed eight average nucleotide identity (ANI)-defined clusters; 55 isolates (90.2%) could not be assigned to known species using current reference databases. Two clusters detected only in Hsinchu were phylogenetically closest to M. kyorinense, with ANI values below the species demarcation threshold. Overall, we observed marked regional heterogeneity of NTM in Taiwan and a persistent identification gap that remained after MALDI-TOF MS optimization and follow-up WGS.IMPORTANCEThis study characterized regional differences in the NTM species distribution across Taiwan, and the results highlight the limitations of current identification approaches. MALDI-TOF MS identifies most isolates, but locally circulating lineages represent a persistent gap in global reference libraries. Even with whole-genome sequencing (WGS), 90.2% (55/61) of persistently unresolved isolates could not be assigned to known species in the current reference databases despite the formation of clear ANI- and phylogeny-defined clusters. These findings show that both proteomic and genomic reference resources for clinical NTM remain incomplete. Expanding regionally representative databases and performing WGS for isolates that remain unresolved by MALDI-TOF MS will be necessary to improve species-level resolution for surveillance and clinical interpretation.

Taiwan

Comparison of the effect of different reference data on Lunar DPX and Hologic QDR-1000 dual-energy X-ray absorptiometers.

We have investigated whether the Lunar DPX (software 3.4) and Hologic QDR-1000 dual-energy X-ray absorptiometers have comparable normal reference databases for the spine and femur of white UK and USA subjects. After conversion for systematic differences in absolute bone density values between the two systems, the reference databases were very similar for the spine in young subjects, but there were clear differences in the femur databases of young females and males of all ages. These differences were confirmed by comparing the percent age-matched and young values determined by the two systems for subjects scanned on both systems. Thus the diagnosis and management of a patient could differ, depending on the system used for the bone density measurements.

Absorptiometry, Photon

CamK-DB: A k-mer MinHash fingerprint database for reference-free genotyping of Camellia accessions.

Tea (Camellia sinensis L.), a major global economic crop in Asia, poses challenges for genetic identification because its highly heterozygous, repetitive genome reduces the efficacy of conventional single-nucleotide polymorphism (SNP) and microsatellite markers, and interspecific hybridization further complicates the situation. To address these issues, CamK-DB was developed as a reference-free Camellia fingerprinting database built on MIKE MinHash sketches. We curated 418 candidate resequencing datasets, and built a database using standardized 5× genome-coverage fingerprints. Each accession is stored as a MIKE. jac fingerprint generated with k = 21 and recommended sketch/pre_cnt = 2000. CamK-DB provides a command-line interface for data management and a custom C++ query engine that computes top-10 matches using Jaccard similarity, complemented by a QT-based graphical interface for interactive analysis. This resource offers a robust and scalable framework for precise and routine germplasm identification, genomic phylogenetic inference, and strategic breeding program design. CamK-DB (database and code) is publicly available at https://github.com/sc-zhang/CamK-DB. CamK-DB binaries are provided for Windows 10/11 and Linux (x86_64, glibc ≥ 2.27).

Databases, Genetic

Meta2DB: curated shotgun metagenomic feature sets and metadata for health state prediction.

SUMMARY: Meta2DB is a curated metagenomic and metadata database that provides structurally consistent microbiome taxonomy feature count tables for 13 897 samples across 84 studies, 23 disease states, and 34 geographical locations. All samples were uniformly processed using a streamlined metagenomic classification pipeline that employs a unique and comprehensive reference database indexed to contain all sequences across all kingdoms of life that were present in the NCBI Nucleotide (nt) database retrieved on 4 January 2023. This pipeline leverages high-performance computing (HPC) resources at Lawrence Livermore National Laboratory and was used to process 50TB of publicly available raw metagenomic sequence data. Extensive metadata curation was carried out through a combination of manual curation and automated parsing, producing a consistent inter-study metadata table specifically structured to facilitate training of ML models for prediction of human health. AVAILABILITY: Data is available at https://gdo-meta2db.llnl.gov/ and https://zenodo.org/records/17315984.

Metadata

resLens: genomic language models to enhance antibiotic resistance gene detection.

The rise of antibiotic resistance necessitates advanced tools to detect and analyze antibiotic resistance genes (ARGs). We present resLens, a family of genomic language models that leverage latent genomic representations to enhance ARG detection and analysis. Unlike alignment-based methods constrained by reference databases, resLens fine-tunes a pre-trained DNA language model on curated ARG datasets, achieving competitive or superior performance in classifying resistance genes across multiple evaluation scenarios, including when ARGs exhibit sequences and mechanisms of resistance dissimilar to those in reference datasets.

Journal Article

nf-core/magmap: Map metatranscriptomes to large collections of genomes.

SUMMARY: The lack of publicly available reference genomes has forced annotation of metatranscriptomes to either use direct alignment of sequence reads to reference databases or de novo assembly. As more and more natural environments are covered by metagenomic surveys, this is rapidly changing. This opens up the possibility of genome-resolved studies of prokaryotic metatranscriptomes by mapping to genomes from public repositories or metagenome-assembled genomes derived from the same environment. Here, we present the nf-core/magmap pipeline that provides a reproducible, easy-to-access, and well-documented workflow for selecting reference genomes, mapping to them, and quantifying features. Genomes can be drawn from public sources or originate from private collections. The pipeline is primarily aimed at prokaryotic communities but can, together with collections of reference mature gene sequences, also be applied to eukaryotes. AVAILABILITY AND IMPLEMENTATION: The nf-core/magmap pipeline is implemented in Nextflow and part of the nf-core collaboration. The pipeline is available at the nf-core website (https://nf-co.re/magmap) and GitHub (https://github.com/nf-core/magmap).

Software

Genome-based predictions of metabolic preferences and substrate phenotypes in psychrotrophic bacteria from permafrost environments.

Genomes reveal vast functional potential, but harbor genomic noise that obscures prediction of metabolic and environmental preferences. Genomic databases are skewed towards clinically relevant and easily cultivated bacteria, limiting predictions for diverse and underrepresented environmental taxa. Psychrotrophic bacteria, which can survive and grow in cold, nutrient-limited, dry, and saline environments, are especially underrepresented despite their relevance for understanding microbial responses to changing cold environments and potential biotechnological value given growth at low temperatures. Assembling complete genomes of 48 isolates from Alaskan permafrost, seasonally frozen active layer soils, and terrestrial ice, we used Kyoto Encyclopedia of Genes and Genomes (KEGG) ortholog annotations to evaluate the predictability of metabolic resource-use traits observed using phenotypic tests. Genome-predicted values for glycolytic versus gluconeogenic catabolic preference index, or sugar-acid preference (SAP), explained over 50% of the variance in empirically observed SAP. SAP was inversely correlated to genomic GC content, which follows phylum-level trends, indicating that coarse metabolic preference covaries with phylogeny. Regularized elastic net models offered a more granular view, linking KEGG genes to specific substrate utilization and sensitivity phenotypes and yielding moderate but reproducible accuracy (AUC 0.70-0.79) for 11 substrates, demonstrating that specific substrate responses may be predictable from relatively small subsets of KO genes. These results extend recent advances, such as the SAP metric, and highlight associations among genomic GC content, phylum, and broad metabolic strategy. Linking genomic content to phenotype using isolates is a necessary step toward predictive models of microbial function in environmental communities, and this work can be used for hypothesis generation, with applications towards more expansive data sets.IMPORTANCECold region soils and ice host psychrotrophic bacteria with metabolic traits and adaptations that enable persistence in harsh, resource-limited environments. However, these taxa are underrepresented in genomic reference databases dominated by well-studied, mesophilic organisms. This gap limits inference of ecological strategies and our ability to predict how these microbes may influence the large, thaw-vulnerable carbon reservoirs in permafrost. Here, we show that genomic GC content is associated with the sugar-versus-acid catabolic preference (SAP) of isolates across major phyla, suggesting that broad genomic features may provide a coarse signal of metabolic strategy. We demonstrate that a modified SAP metric, using binary (positive/negative) substrate utilization rather than detailed growth rate measurements, is moderately predictive, thus extending its application to slow-growing or difficult-to-culture taxa. Together, these advances broaden the toolkit for linking genome content to resource-use traits (phenotype) in poorly characterized, cold-adapted bacteria and offer a tractable entry point to broad prediction and hypothesis generation.

Genome, Bacterial

Metax enables accurate cross-domain taxonomic profiling of metagenomes.

Taxonomic profiling is fundamental to microbiome research, yet achieving high species-level accuracy remains challenging for complex communities that span bacteria, viruses, eukaryotes, and archaea, and these limitations are exacerbated in low-biomass, host-dominated samples. We introduce Metax, a cross-domain taxonomic profiler that integrates coverage-based probabilistic modeling with an expectation-maximization framework to distinguish true microbial signals from artifacts. Across >600 samples from host-associated, environmental, wastewater, and low-biomass clinical settings, including benchmarks with limited reference representation, Metax improved profiling accuracy, achieving on average 55% higher F1 scores and 45% lower Bray-Curtis dissimilarity than other methods. Moreover, this broad evaluation demonstrated that Metax resolved bacterial and viral signatures of peri-implantitis in oral microbiomes and revealed signals suggestive of reagent-borne contaminants and reference misassemblies in plasma-cell-free DNA. By leveraging genome-wide coverage evidence, Metax enables robust cross-domain profiling across diverse sample types and sequencing depths, including settings where reference databases are highly incomplete.

abundance estimation