PubMed HealthSearch

SEARCH · PubMed Health

Results for “Reference database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

The Role of Community Science in DNA-Based Biodiversity Monitoring.

The mutual interest in nature by the general public and scientists has led to many collaborations, past and present. Community science shows great potential for monitoring species occurrences and distributions, especially in combination with scalable and (semi)-automated methods such as DNA-based monitoring, helping to obtain data from a broader geographic and temporal range than would be possible by the scientific community alone. Here, we present an overview of the complementarity between community science and DNA-based biomonitoring through examples from ongoing projects. The involvement of hobby experts is particularly crucial for building up the necessary species reference databases that enable DNA-based monitoring. Based on this overview, we identify some key points related to learning opportunities and participant recognition to maximise the success, impact and benefit of community participants in DNA-based monitoring.

Biodiversity

Population genetics in forensic DNA typing.

Variable number of tandem repeat (VNTR) sequences are used to link defendants with crimes by matching DNA patterns. The probative value of a match is often calculated by multiplying together the estimated frequencies with which each particular VNTR pattern occurs in a reference database. However, this method is liable to potentially serious errors because ethnic subgroups within major racial categories exhibit genetic differences that are maintained by endogamy. The multiplication procedure currently in use can be made scientifically valid only by extensive sampling of VNTR frequency distributions in a variety of ethnic groups, similar to the ethnic studies of various blood groups done in the past. Alternative approaches for dealing with subpopulation heterogeneity are discussed.

Alleles

Proposal to validate Listeria swaminathanii sp. nov. and reassign the type strain to UTK S2-0008.

Listeria swaminathanii UTK S2-0008, isolated from soil collected in the Nantahala National Forest in North Carolina, USA, is the only L. swaminathanii strain eligible to serve as the type strain, which is needed to achieve valid status. The previously effectively published type, L. swaminathanii FSL L7-0020T, and previously described strains UTK C1-0015 and UTK C1-0024 do not conform to the International Code of Nomenclature of Prokaryotes' rules for type strains. Additionally, the currently designated type strain (FSL L7-0020T = ATCC TSD-239T) is an atypical representative of L. swaminathanii as it is the only strain lacking catalase activity. Therefore, it is proposed to reassign the type to L. swaminathanii (UTK S2-0008T = CCUG 77280T = LMG 33255T). Whole-genome sequence-based average nucleotide identity (ANI) showed that this strain clustered with the three previously described L. swaminathanii strains (FSL L7-0020 = ATCC TSD-239, UTK C1-0015, and UTK C1-0024; pairwise ANI ranged from 98.71% to 98.83%). All four strains, including the one described here, could not be classified as any validly published Listeria species and showed the highest similarity to Listeria marthii (maximum ANI of 93.92%, in silico DNA-DNA hybridization of 56.2%). L. swaminathanii exhibits the phenotypic characteristics that are currently expected of the Listeria sensu stricto species. This species lacks phenotypic characteristics associated with Listeria pathogenicity (non-hemolytic and negative for phosphatidylinositol-specific phospholipase C activity); the genomes lack genes associated with virulence (all genes found on the Listeria pathogenicity island 1 [LIPI-1], as well as the internalin genes inlA and inlB), which support L. swaminathanii is nonpathogenic.IMPORTANCEThe genus Listeria includes species of significant relevance to food safety, environmental microbiology, and public health. Accurate species identification is critical because misidentification of nonpathogenic species as pathogenic ones can lead to unnecessary recalls and regulatory complications. The validation of Listeria swaminathanii sp. nov. will ensure that this species is formally recognized and has a type strain (UTK S2-0008T) that is representative of the species. This work strengthens diagnostic accuracy by enabling the inclusion of this species in reference databases and inclusivity studies, reducing the risk of false identification. Furthermore, the identification and characterization of Listeria swaminathanii sp. nov. expands our understanding of the genetic and ecological diversity within the genus Listeria, particularly among soil-dwelling strains.

Listeria

Genetic Research on Cardiac Channelopathies in African and African-Descent Populations: A Scoping Review.

Cardiac channelopathies are inherited arrhythmias that can lead to sudden cardiac death. Despite Africa's extensive genomic diversity, African and African-descent populations remain underrepresented in genetic research, creating gaps in variant interpretation and clinical care. This scoping review aims to map the extent, range, and nature of genetic research on cardiac channelopathies in these populations and to identify key geographic, thematic, and methodological gaps. Using the Joanna Briggs Institute scoping review methodology and the Population-Concept-Context framework, systematic searches in PubMed, Embase, and Web of Science identified original human studies on cardiac channelopathies with genetic data. Extracted variables included study characteristics, populations, types of channelopathies, and reported genes and variants. Forty-four studies met the inclusion criteria. Most studies originated from the United States and South Africa, while West, Central, and East Africa were largely underrepresented. US Black individuals and South African individuals of continental African or African-descended ancestry (excluding populations of European descent such as Cape Afrikaner people) were the most studied groups, with other continental African groups rarely included. Long QT syndrome was the predominant focus, and SCN5A, KCNQ1, and KCNH2 were the most frequently analyzed genes. Many of the genetic variants discussed remained of uncertain significance due to limited functional validation and the underrepresentation of African genomes in reference databases. Genetic research on cardiac channelopathies in populations of African ancestry is limited, restricting variant interpretation, counseling, and risk prediction. Broader African inclusion, expanded gene screening, and functional studies are essential to improve diagnostics and promote equity in genomic medicine.

Humans

Using Mapping-Profiles to Refine Strain-Level Metagenomic Classification.

Metagenomic classification at the strain level remains challenging due to high sequence similarity among closely related genomes, which leads to ambiguous read mappings and frequent false-positive strain detections. Reducing such errors improves the reliability of strain-level analyses, which is critical for applications such as pathogen detection. We introduce StrainRefine, a post-mapping refinement method that analyzes read-reference mapping profiles to resolve ambiguous assignments among highly similar genomes. The method represents candidate reference genomes using binary profiles that capture read-support patterns and measures similarity between references based on profile overlap. The method clusters references based on similar mapping profiles, filters weakly supported genomes, and reassigns reads to representative references, reducing redundant reporting of near-identical strains. StrainRefine substantially reduces false-positive strain detections while preserving recall and improving agreement between predicted and true abundance profiles. On large-scale metagenomic datasets, it achieves a substantially improved precision-recall balance compared with existing mapping-based approaches, with the standalone method obtaining the highest read-level classification accuracy on the most complex evaluated dataset. Unlike many strain-level tools designed for individual species, StrainRefine operates without prior assumptions about sample composition or curated species-specific reference collections, while still achieving comparable performance in single-species settings on species-specific reference databases. These results highlight mapping-profile similarity as an effective signal for improving strain-level metagenomic classification.

false-positive reduction

Establishing the ELIXIR Microbiome Community.

Microbiome research has grown substantially over the past decade in terms of the range of biomes sampled, identified taxa, and the volume of data derived from the samples. In particular, experimental approaches such as metagenomics, metabarcoding, metatranscriptomics and metaproteomics have provided profound insights into the vast, hitherto unknown, microbial biodiversity. The ELIXIR Marine Metagenomics Community, initiated amongst researchers focusing on marine microbiomes, has concentrated on promoting standards around microbiome-derived sequence analysis, as well as understanding the gaps in methods and reference databases, and identifying solutions to the computational overheads of performing such analyses. Nevertheless, the methods used and the challenges faced are not confined to marine microbiome studies, but are broadly applicable to other biomes. Thus, expanding this Marine Metagenomics Community to a more inclusive ELIXIR Microbiome Community will enable it to encompass a broader range of biomes and link expertise across 'omics technologies. Furthermore, engaging with a large number of researchers will improve the efficiency and sustainability of bioinformatics infrastructure and resources for microbiome research (standards, data, tools, workflows, training), which will enable a deeper understanding of the function and taxonomic composition of the different microbial communities.

Computational Biology

Expanding kinetoplastid genome annotation through protein structure comparison.

Kinetoplastids belong to the Discoba supergroup, an early divergent eukaryotic clade. Although the amount of genomic information on these parasites has grown substantially, assigning gene functions through traditional sequence-based homology methods remains challenging. Recently, significant advancements have been made in in-silico protein structure prediction and algorithms for rapid and precise large-scale protein structure comparisons. In this work, we developed a protein structure-based homology search pipeline (ASC, Annotation by Structural Comparisons) and applied it to transfer biological information to all kinetoplastid proteins available in TriTrypDB, the reference database for this lineage. Our pipeline enabled the assignment of structural similarity to a substantial portion of kinetoplastid proteins, improving current knowledge through annotation transfer. Additionally, we identified structural homologs for representatives of 6,700 uncharacterized proteins across 33 kinetoplastid species, proteins that could not be annotated using existing sequence-based tools and databases. As a result, this approach allowed us to infer potential biological information for a considerable number of kinetoplastid proteins. Among these, we identified structural homologs to ubiquitous eukaryotic proteins that are challenging to detect in kinetoplastid genomes through standard genome annotation pipelines. The results (KASC, Kinetoplastid Annotation by Structural Comparison) are openly accessible to the community at kasc.fcien.edu.uy through a user-friendly, gene-by-gene interface that enables visual inspection of the data.

Kinetoplastida

Oral bacteriome in pediatric patients with malignancies prior to chemotherapy: a pilot study using full-length 16S rRNA sequencing.

OBJECTIVE: To characterize the composition, diversity, and ecological features of the oral bacteriome in pediatric patients with malignancies prior to chemotherapy initiation. METHODS: In this prospective pilot study,supragingival plaque samples were collected from 10 pediatric cancer patients prior to the initiation of chemotherapy. Bacterial genomic DNA was extracted from each sample, and the full-length 16S rRNA gene was amplified and sequenced on the PacBio Sequel II platform using circular consensus sequencing (CCS). Raw CCS reads were quality-filtered and denoised into amplicon sequence variants (ASVs) using DADA2, and taxonomic assignment was performed against the SILVA 138 reference database. Alpha diversity was assessed using the Chao1, Shannon, Simpson, and Faith's phylogenetic diversity (PD whole tree) indices, while beta diversity was evaluated through principal coordinate analysis (PCoA), and non-metric multidimensional scaling (NMDS). Microbial co-occurrence networks were constructed to characterize bacterial interactions, and functional potential was predicted using PICRUSt2, and BugBase. RESULTS: A total of 614,473 high-quality CCS reads were generated, yielding 1,697 ASVs. Alpha diversity analysis revealed substantial inter-individual variation in microbial richness and diversity among the pediatric cancer patients. The bacterial community was dominated by the phyla Firmicutes, Proteobacteria, Bacteroidota, Actinobacteriota. At the genus level, Streptococcus, Prevotella, Neisseria, and Haemophilus were the most abundant taxa. Beta diversity analysis revealed distinct clustering patterns, indicating highly individualized microbial profiles. Co-occurrence network analysis identified several keystone taxa and potential pathogenic associations within the supragingival plaque community. Functional prediction indicated that the dominant metabolic pathways were related to amino acid metabolism, carbohydrate metabolism, and membrane transport. CONCLUSION: These preliminary findings reveal a taxonomically diverse, highly individualized pre-chemotherapy oral bacteriome, providing foundational baseline profiles to guide future longitudinal investigations of chemotherapy-induced dysbiosis and personalized interventions.

Humans

Structural variant discovery and diagnostic impact in rare diseases from short-read and long-read sequencing.

Rare diseases collectively affect 1 in 10 individuals, yet current genetic testing fails to identify a causal variant for most cases. At present, cytogenetic methods and/or sequencing approaches such as exome (ES) or short-read genome sequencing (srGS) represent the state-of-the-art for comprehensive clinical discovery of sequence and structural variants (SVs), including copy number variants, balanced SVs, complex SVs, and tandem repeats (TRs). Recently, long-read genome sequencing (lrGS), coupled with multiomics data, has presented great promise to resolve variation in genomic regions recalcitrant to characterization by srGS such as highly repetitive simple repeat sequences and segmental duplications. However, there are few guidelines to enable clinical interpretation of genetic variation in these highly repetitive genomic regions, and the enthusiasm of the field in adopting lrGS has made it difficult to assess the true added diagnostic yield of this technology due to widely variable and inconsistently applied analytic pipelines and variable degrees of pre-screening by ES or srGS. Here, we investigated the contribution of SVs to rare diseases using srGS as a front-line strategy when paired with highly sensitive SV discovery and evaluate the added diagnostic yield of incorporating lrGS for a subset of cases. Our srGS analysis encompassed 1,462 families (3,450 individuals) recruited through the Broad Institute Center for Mendelian Genetics and the Genomics Research to Elucidate the Genetics of Rare Diseases (GREGoR) programs. Diagnostic SVs were identified in 5.4% of cases (79/1,462), of which 80% were uniquely detectable by srGS compared to standard cytogenetic techniques. For 96 families (including 10 families with a heterozygous variant observed in a known recessive gene of clinical relevance), we performed lrGS with methylation profiling, as well as long-read transcriptomic analyses in a subset of 20 trios. Analyses with lrGS yielded over 25,000 SVs per genome, 63% of which were not captured by srGS, along with an additional ~200 rare SNV/indels per genome not previously captured and 12 differentially methylated regions per genome. Among these, we identified only one diagnostic variant not interpreted by srGS, an apparently mosaic de novo SNV in CASK that was absent in the srGS callset due to allelic imbalance. No new diagnoses were supported by long-read transcriptomics or episignatures. In this well characterized rare disease cohort, the added diagnostic yield was thus 1.04% (1/96 families). Following a systematic literature review of prior lrGS studies, we find that most reported diagnoses were detectable by srGS and that our added diagnostic yield is consistent with those prior studies. These studies emphasize the significant impact of comprehensive SV discovery in rare disease cases and further demonstrate the power for increased discovery of novel genomic variation and episignatures from lrGS. Nonetheless, they also serve to temper expectations of dramatic diagnostic advances in rare disease patients until there is more extensive annotation of the functional and clinical impact of all coding and noncoding variation uniquely accessible to lrGS with extensive reference databases spanning highly repetitive genomic sequencing that could be enabled by this transformative technology.

Journal Article

PC CLIN-SIM: a toolbook based clinical simulation environment.

The Departments of Computer Medicine, Health Care Sciences, Medicine, and Electrical Engineering & Computer Science at the George Washington University have joined forces to create a clinical simulation program. The purpose of this program is to provide experience in the management of complex patient populations (eg geriatrics). A number of simulation programs are available commercially, however none provide adequate geriatric content, or were deemed to lack functionality important to the developers. The immediate goal of this effort was to create a computer-based, core curriculum in geriatric medicine for medical and allied health students. The curriculum includes case simulations linked to a comprehensive reference database. The development objectives were to create an intuitive, friendly, consistent user interface which could serve as a shell for additional content areas. In order to increase fidelity, free text entry and time simulation were included.

Computer Graphics

MetaFX: feature extraction from whole-genome metagenomic sequencing data.

MOTIVATION: Microbial communities consist of thousands of microorganisms and viruses and have a tight connection with an environment, such as gut microbiota modulation of host body metabolism. However, the direct relationship between the presence of certain microorganism and the host state often remains unknown. Toolkits using reference-based approaches are limited to microbes present in databases. Reference-free methods often require enormous resources for metagenomic assembly or results in many poorly interpretable features based on k-mers. RESULTS: Here we present MetaFX-an open-source library for feature extraction from whole-genome metagenomic sequencing data and classification of groups of samples. Using a large volume of metagenomic samples deposited in databases, MetaFX compares samples grouped by metadata criteria (e.g. disease, treatment, etc.) and constructs genomic features distinct for certain types of communities. Features constructed based on statistical k-mer analysis and de Bruijn graphs partition. Those features are used in machine learning models for classification of novel samples. Extracted features can be visualized on de Bruijn graphs and annotated for providing biological insights. We demonstrate the utility of MetaFX by building classification models for 590 human gut samples with inflammatory bowel disease. Our results outperform the previous research disease prediction accuracy up to 17%, and improves classification results compared to taxonomic analysis by 9±10% on average. AVAILABILITY AND IMPLEMENTATION: MetaFX is a feature extraction toolkit applicable for metagenomic datasets analysis and samples classification. The source code, test data, and relevant information for MetaFX are freely accessible at https://github.com/ctlab/metafx under the MIT License. Alternatively, MetaFX can be obtained via http://doi.org/10.5281/zenodo.16949369.

Metagenomics

Genome-related datasets within the E. coli Genetic Stock Center database.

The contents of the E. coli Genetic Stock Center database and the availability in electronic form of the subset of information most relevant to sequence databases are described. The database uses the long-standing Stock Center records (developed and curated by Dr B.J.Bachmann) in describing genotypes of mutant derivatives of E.coli K-12 in terms of alleles, structural mutations, mating type, and plasmids as well as the derivation, names and originators of the strain, and references. The database includes descriptions of mutations, mutation properties, genes, gene properties, and gene products, with EC number identifiers for enzymes. Sequence information is not included, but entries refer to sequence database accession numbers for sequenced regions. A gene is described as a subtype of a more general category of chromosome interval called Site. Since sites are used to describe any chromosomal interval, mapping information is associated with sites. Alleles are described as mutations of those sites and they are not primary map objects, but inherit map position information from the corresponding site description. The database design is intended to preserve richness of detail where it is known and uncertainty of measurements or information as it occurs in order to represent the stock center records as accurately as possible.

Bacterial Proteins

Comparison and evaluation of nine bibliographic databases concerning adverse drug reactions.

Few evaluations and statistical comparisons of bibliographic databases have been published. As a drug information center, we were particularly interested in databases providing references on adverse drug reactions (ADRs). Ten drugs were randomly chosen from the 2000 files at our center. Nine databases were selected according to the high frequency of references concerning ADRs: eight online systems (MEDLINE, BIOSIS, TOXLINE, Iowa Drug Information System, PASCAL, EMBASE, PHARMLINE, and International Pharmaceutical Abstracts [IPA]), and one Compact Disk Read Only Memory (CD-ROM) system (Core MEDLINE). The total number of references, the number of references from 1987 to 1989, and the number of relevant references from 1987 to 1989 were analyzed using the Friedman two-way ANOVA by ranks. The overlap between databases for only one drug, carboplatin, and the quality:cost ratio were also studied. Considering the total number of references, TOXLINE and EMBASE were significantly superior to IPA, PHARMLINE, PASCAL, and Core MEDLINE. For the period 1987-1989, EMBASE was significantly superior to PASCAL, IPA, PHARMLINE, and Core MEDLINE with regard to total number of references, and significantly superior to PASCAL, Core MEDLINE, and IPA with regard to relevance. MEDLINE, TOXLINE, and EMBASE had the best quality:cost ratio. EMBASE had the slightest overlap of references, with 53 percent of the unique references on carboplatin. This comparative evaluation showed that the ability of bibliographic databases to provide information on ADRs is dependent on both the size and the quality of each database.

Databases, Bibliographic

Algorithm for point-to-point correlation of geometrically nearly similar microscopic objects.

An algorithm is presented that compares two quasi similar images by correlating selected points on them--assuming their coordinates are available. The procedure involves translational, magnificational and rotational operations to find corresponding point pairs on the pictures. The algorithm automatically compensates for slight dissimilarities between images and constructs a reference point database for correlation during the evaluation process. Establishment of the reference point networks on the images prior to the examination is avoided.

Algorithms

theBIGbam: compression and interactive exploration of large-scale sequencing alignments with circular mapping support.

SUMMARY: theBIGbam (github.com/bhagavadgitadu22/theBIGbam) is a genome browser and alignment viewer designed for massive metagenomic and metatranscriptomic datasets. The tool takes BAM files containing read alignments, together with genome assemblies in FASTA format or annotated genome sequences in GenBank format. Alternatively, it can start from raw FASTQ reads and generate alignments using a modified mapper that supports circular genomes, enabling seamless read mapping across genome ends. theBIGbam can compress hundreds of gigabytes of input files 10- to 100-fold into dedicated databases while retaining key per-position information, including coverage depth and recurrent mismatches, insertions, and deletions between reads and the reference. These databases can be served to a local web browser, enabling interactive exploration of any contig in any sample using DNAFeaturesViewer for genome maps and Bokeh for mapping-derived features. Contig-sample pairs available for visualization can be filtered using a range of summary metrics calculated per contig, per sample, and per contig-sample pair to guide users toward the most relevant signals. Through its interactive visualization, theBIGbam facilitates the exploration of complex datasets, while its integrated database-combining assembly features, annotated features, and mapping-derived features-provides the information needed to investigate biological hypotheses systematically. Designed to complement existing browsing tools like IGV and Anvi'o, theBIGbam is particularly suited for examining misassemblies, subpopulations, microdiversity, and contig topology in large-scale datasets. AVAILABILITY AND IMPLEMENTATION: theBIGbam is an open-source Rust/Python package that can be installed from Bioconda or PyPI. The source code and documentation are available on GitHub (github.com/bhagavadgitadu22/theBIGbam).

Software

Mapping the oral microbiome opens links to periodontitis.

Many microbiome analysis techniques can only detect the microbes present in the reference genome database used. In this issue of Cell Host & Microbe, Cha et al. establish an improved genome database of the human oral microbiome, which they use to discover a connection between periodontitis and an enigmatic bacterial phylum.

Humans

Medical reference filing systems.

Low priced shareware packages may be adequate for those whose main need is to index a collection of articles read in journals. Those who search the literature electronically will need a package able to import mass data. Pro-Cite is set up for bibliographic functions only, but Reference File and Notebook II can be used to index any kind of collection, with Reference File having the advantage of being memory-resident. Other reference filing databases being considered for purchase could be measured against those reviewed here.

Abstracting and Indexing

Wheat gliadin: digital imaging and database construction using a 4-band reference system of agarose isoelectric focusing patterns.

An isoelectric focusing method using thin-layer agarose gel has been developed for wheat gliadin. Using flat-bed units with a third electrode, up to 72 samples per gel may be analyzed. Advantages over traditional acid polyacrylamide gel electrophoresis methodology include: faster run times, nontoxic media, and greater sample capacity. The method is suitable for fingerprinting or purity testing of wheat varieties. Using digital images captured by a flat-bed scanner, a 4-band reference system using isoelectric points was devised. Software enables separated bands to be assigned pI values based upon reference tracks. Precision of assigned isoelectric points is shown to be on the order of 0.02 pH units. Captured images may be stored in a computer database and compared to unknown patterns to enable an identification. Parameters for a match with a stored pattern may be adjusted for pI interval required for a match, and number of best matches.

Databases, Factual