PubMed HealthSearch

SEARCH · PubMed Health

Results for “reference databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

ProteoParc: A Reference Protein Database Builder for Ancient and Nonmodel Organisms.

Over the past few years, the increasing interest in analyzing the proteome of extinct and nonmodel organisms has generated a new field of research expanding the scope of proteomics. The lack of curated databases and/or molecular data from these organisms forces researchers to manually search in different public repositories for related protein sequences, either for MS/MS peptide identification or ZooMS marker annotation. This can lead to format incongruences and hinder reproducibility between studies. To address this issue, we introduce ProteoParc, a user-friendly software that builds reference databases by systematically downloading and processing protein sequences from the most widely used public repositories. The pipeline's output is a nonredundant protein database, formatted in a way to be interpreted by typical peptide identification software. Moreover, the user can adjust the database dimension and composition by applying different criteria to include only a certain number of genes or species. Thus, ProteoParc is an easy and fast, custom-made bioinformatic tool useful for future paleoproteomics analysis in ancient samples related to understudied organisms.

Databases, Protein

raxtax: a k-mer-based non-Bayesian taxonomic classifier.

MOTIVATION: Taxonomic classification in biodiversity studies is the process of assigning the anonymous sequences of a marker gene (barcode) or whole genomes (metagenomics) to a specific lineage using a reference database that contains named sequences in a known taxonomy. This classification is important for assessing the diversity of biological systems. Taxonomic classification faces two main challenges: first, accuracy is critical as errors can propagate to downstream analysis results; and second, the classification time requirements can limit study size and study design, in particular when considering the constantly growing reference databases. To address these two challenges, we introduce raxtax, an efficient, novel taxonomic classification tool for barcodes that uses common k-mers between all pairs of query and reference sequences. We also introduce two novel uncertainty scores which take into account the fundamental biases of reference databases. RESULTS: We validate raxtax on three widely-used empirical reference databases and show that it is 2.7-100 times faster than competing state-of-the-art tools on the largest database while being equally accurate. In particular, raxtax exhibits increasing speedups with growing query and reference sequence numbers compared to existing tools (for 100 000 and 1 000 000 query and reference sequences overall, it is 1.3 and 2.9 times faster, respectively), and therefore alleviates the taxonomic classification scalability challenge. AVAILABILITY AND IMPLEMENTATION: raxtax is available at https://github.com/noahares/raxtax under a CC-NC-BY-SA license. The scripts and summary metrics used in our analyses are available at https://github.com/noahares/raxtax_paper_scripts. The source code, sequence data, and summarized results of the analyses are available at https://doi.org/10.5281/zenodo.15057027.

Software

The development of a multicenter database for reference values in clinical neurophysiology--principles and examples.

This paper describes the work undertaken to establish principles for the development of multicenter databases for reference values in clinical neurophysiology. The study was initiated because of interest of the involved laboratories in knowledge-based systems in electromyographic diagnosis, for which it was necessary to formalize the key concepts in the diagnostic process: diseases, pathophysiology and test results. The paper deals specifically with the structuring of results of motor and sensory nerve conduction studies.

Action Potentials

MegaPX: fast and space-efficient peptide assignment method using IBF-based multi-indexing.

MOTIVATION: A central problem for metaproteomic analysis is the often-unknown taxonomic composition of the analyzed microbiomes. Using a database search, the standard approach requires prior knowledge of which proteins and taxa to include in the protein reference database or to use tailored metagenome-derived databases, which are expensive and error-prone in their generation. A possible strategy to circumvent this database search issue is de novo sequencing, where peptide sequences are directly identified from mass spectra. However, these sequences must still be mapped back to potentially extensive databases. Here, alignment-based approaches enable robust and precise results, with the potential drawback of high memory usage and long run times. RESULTS: We present MegaPX, a software for rapidly classifying de novo peptide sequences against large protein databases. MegaPX implemented as a C++-based tool, uses an alignment-free, k-mer approach as a taxonomic classification method with the possibility of generating mutated reference databases for error-tolerant searching. It uses various algorithms, including interleaved Bloom filters, to efficiently compute approximate membership queries, ensuring fast processing times while querying and indexing large databases in a multi-indexing fashion. We demonstrate the potential of MegaPX by analyzing different samples, including metaproteomics, against extensive reference databases, highlighting its use as a fast screening tool.

Software

Referring physician database: the cornerstone of a successful marketing program.

In the language of authors Thomas Dunlap and Hedy Rogers, increased competition, "turf wars" and shifting demographics have made developing a keen competitive edge vital to growing and maintaining group practice income. They present the referring physician database as one method for satisfying the needs of a group's customer base.

Databases, Factual

Medical Facts File: a self service database of reference information.

The Dahlgren Memorial Library, Georgetown University Medical Center, will demonstrate Medical Facts File, a newly developed in-house database of general medical information. The file content emerged from the library's experience with commonly asked reference questions and the need to develop a database as an online source for users seeking quick answers to medical queries. Medical Facts File joins a growing family of over 18 databases which comprise Georgetown's IAIMS Knowledge Network. Use scenarios will demonstrate how an online search is initiated, either directly or as a prompt from one of the other online databases. The design of Medical Facts File at the Dahlgren Memorial Library began in late 1989 with a publishing section on instructions for authors planning to submit manuscripts to a variety of prominent medical journals. Since then, seven sections have been identified for the database. Three sections are highlighted for presentation, although work on the project is on-going. Medical Facts File is an easy-to-use, time saving system that facilitates tedious searching through a multitude of library sources. It provides users with a self-service, information look-up system.

Databases, Factual

Extensive Analysis of Genetic Diversity in HLA-DMA, HLA-DMB, HLA-DOA and HLA-DOB: Characterisation of 236 Novel Alleles.

HLA-DMA, -DMB, -DOA and -DOB are non-classical HLA Class II genes that play a crucial role in the selection of highly stable HLA Class II/peptide complexes on antigen-presenting cells. Although the genes were initially thought to have a limited diversity with less than 13 alleles per gene documented in the IPD-IMGT/HLA Database in 2022, recent studies suggest a potential impact of certain alleles on the outcome of hematopoietic cell transplantation. To gain a deeper understanding of allelic diversity, we sequenced HLA-DMA, -DMB, -DOA and -DOB of 1880 potential stem cell donors from Germany, Poland, Great Britain and Chile, achieving full-gene resolution. Remarkably, we identified 3968 previously undescribed sequences, including 28 distinct novel proteins. The observed allele frequencies were consistent across all studied populations with one dominating protein for each gene: HLA-DMA*01:01 (> 77%), HLA-DMB*01:01 (> 63%), HLA-DOA*01:01 (> 97%) and HLA-DOB*01:01 (> 77%). Notably, a much higher diversity was observed in full-genomic resolution. Finally, we submitted 51 distinct novel sequences for HLA-DMA, 58 for HLA-DMB, 80 for HLA-DOA and 47 for HLA-DOB to the IPD-IMGT/HLA Database. This comprehensive reference database update will not only simplify future genotyping of HLA-DMA, -DMB, -DOA and -DOB but will hopefully also enhance our understanding of the complex process of peptide selection and loading to the HLA Class II proteins.

Humans

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus

Demixer: a probabilistic generative model to delineate different strains of a microbial species in a mixed infection sample.

MOTIVATION: Multi-drug resistant or hetero-resistant tuberculosis (TB) hinders the successful treatment of TB. Hetero-resistant TB occurs when multiple strains of the TB-causing bacterium with varying degrees of drug susceptibility are present in an individual. Existing studies predicting the proportion and identity of strains in a mixed infection sample rely on a reference database of known strains. A main challenge then is to identify de novo strains not present in the reference database, while quantifying the proportion of known strains. RESULTS: We present Demixer, a probabilistic generative model that uses a combination of reference-based and reference-free techniques to delineate mixed infection strains in whole genome sequencing (WGS) data. Demixer extends a topic model widely used in text mining to represent known mutations and discover novel ones. Parallelization and other heuristics enabled Demixer to process large datasets like CRyPTIC (Comprehensive Resistance Prediction for Tuberculosis: an International Consortium). In both synthetic and experimental benchmark datasets, our proposed method precisely detected the identity (e.g. 91.67% accuracy on the experimental in vitro dataset) as well as the proportions of the mixed strains. In real-world applications, Demixer revealed novel high confidence mixed infections (101 out of 1963 Malawi samples analysed), and new insights into the global frequency of mixed infection (2% at the most stringent threshold in the CRyPTIC dataset) and its significant association to drug resistance. Our approach is generalizable and hence applicable to any bacterial and viral WGS data. AVAILABILITY AND IMPLEMENTATION: All code relevant to Demixer is available at https://github.com/BIRDSgroup/Demixer.

Mycobacterium tuberculosis

Evaluation of Hinf I-generated VNTR profile frequencies determined using various ethnic databases.

Concerns have been raised about hypothetical problems arising from the use of statistics for determining the likelihood of occurrence of DNA profiles for forensic purposes. A major contention is that reference databases based on subgroups of a major population category rather than on general (or major) population groups, might yield large differences in the estimated likelihood of occurrence of DNA profiles. This hypothetical issue is based on the assertion by some people that the differences among subgroups within a race would be greater than between races (at least for forensic purposes). To evaluate the effects of the above concern the likelihood of occurrence of 615 Hinf I-generated target DNA profiles was estimated using fixed bin frequencies from various ethnic databases and the multiplication rule. Based on the data in this study, differences in allele frequencies at a particular locus do not have substantial effects on VNTR profile frequency estimates when subgroup reference databases from within a major population group are compared. In contrast, the greatest variation in statistical estimates occurs across-major population groups. Therefore, the assertion, by some critics that the differences among subgroups within a race would be greater than between races (at least for forensic purposes), is unfounded. The data in the study support that comparisons across major population groups provide valid estimates of DNA profile frequencies without forensically significant consequences. The data do not support the need for alternate procedures, such as the ceiling principle approach, for deriving statistical estimates of DNA profile frequencies.

Black People

Effectiveness of mass spectrometry and genomic analysis in the surveillance of nontuberculous Mycobacterium in Taiwan.

Nontuberculous mycobacteria (NTM) are diverse, and species-level identification remains challenging in routine diagnostics. We analyzed NTM isolates collected at three regional centers of the National Taiwan University Hospital (NTUH) from 2019 to 2024 to assess geographic variation and identification performance after implementation of matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS). Among 3,188 cases meeting the microbiological criteria for probable pulmonary NTM disease, the species distribution differed by region: Mycobacterium avium complex predominated in central Taiwan (Yunlin, 47.3%), whereas M. abscessus complex (Taipei, 26.5%) and M. kansasii (Hsinchu, 12.4%) were more common in northern Taiwan. In 2019, 14.5% of isolates were reported to be unidentified by MALDI-TOF MS; with workflow optimization and database updates, this percentage decreased but plateaued at 4.5-4.8%. Whole-genome sequencing (WGS) of 61 randomly selected persistently unidentified isolates revealed eight average nucleotide identity (ANI)-defined clusters; 55 isolates (90.2%) could not be assigned to known species using current reference databases. Two clusters detected only in Hsinchu were phylogenetically closest to M. kyorinense, with ANI values below the species demarcation threshold. Overall, we observed marked regional heterogeneity of NTM in Taiwan and a persistent identification gap that remained after MALDI-TOF MS optimization and follow-up WGS.IMPORTANCEThis study characterized regional differences in the NTM species distribution across Taiwan, and the results highlight the limitations of current identification approaches. MALDI-TOF MS identifies most isolates, but locally circulating lineages represent a persistent gap in global reference libraries. Even with whole-genome sequencing (WGS), 90.2% (55/61) of persistently unresolved isolates could not be assigned to known species in the current reference databases despite the formation of clear ANI- and phylogeny-defined clusters. These findings show that both proteomic and genomic reference resources for clinical NTM remain incomplete. Expanding regionally representative databases and performing WGS for isolates that remain unresolved by MALDI-TOF MS will be necessary to improve species-level resolution for surveillance and clinical interpretation.

Taiwan

High cost factors for leukaemia and lymphoma patients: a new analysis of costs within these diagnosis related groups.

STUDY OBJECTIVE: To determine high cost factors to help managers and clinicians to analyse the reasons of adverse costs and provide indications for financial negotiation. DESIGN: To locate high cost or long stay patients, the analysis was designed on the basis of a mixture of Weibull distributions. In this new model, the proportion of high cost patients was expressed according to the multinomial logistic regression, permitting the determination of high cost factors. SETTING: The 1993 French reference database, constituted in the framework of the national study of DRG costs, conducted by the French Ministry of Health. The database of discharge abstracts recorded in 1993 in the Dijon public teaching hospital. PARTICIPANTS: The analyses were based on 1352 abstracts from the French reference database and 368 from the Dijon database concerning patients, aged 18 and over, suffering from leukaemia and lymphoma. MAIN RESULTS: High cost and long stay factors were the same: number of stays, death, transfer, acute leukaemia, neutropenia, septicaemia, high dose aplastic chemotherapy, central venous catheterisation, parenteral nutrition, protected or laminar airflow room, blood transfusion, and intravenous antibiotherapy. CONCLUSIONS: Taking into account high cost predictive factors, as shown in the case of leukaemia and lymphoma patients, would help to reduce the adverse effects of a prospective payment system.

Adolescent

Comparison of the effect of different reference data on Lunar DPX and Hologic QDR-1000 dual-energy X-ray absorptiometers.

We have investigated whether the Lunar DPX (software 3.4) and Hologic QDR-1000 dual-energy X-ray absorptiometers have comparable normal reference databases for the spine and femur of white UK and USA subjects. After conversion for systematic differences in absolute bone density values between the two systems, the reference databases were very similar for the spine in young subjects, but there were clear differences in the femur databases of young females and males of all ages. These differences were confirmed by comparing the percent age-matched and young values determined by the two systems for subjects scanned on both systems. Thus the diagnosis and management of a patient could differ, depending on the system used for the bone density measurements.

Absorptiometry, Photon

Comparison of Hae-III-generated VNTR profiles of Chinese in Hong Kong and Singapore.

Some statistical analysis and experiments have been carried out on the single-locus VNTR reference databases collected separately from the Chinese in Hong Kong and Singapore to assess the issues of sample size and representativeness in the context of forensic science. We employed the bootstrap method and the Pearson chi-squared test for independence, and found that the different reference databases have very little effect in assessing the frequencies of occurrence of variable number of tandem repeats (VNTR) profiles. The results of random match probability assessment also indicate that the sample sizes studied appeared adequate to provide representative data. It has been found that a battery of four loci was enough to distinguish the 458 individuals under study.

Asian People

How the new Hologic hip normal reference values affect the densitometric diagnosis of osteoporosis.

In February 1997, Hologic supplied new software to all QDR dual-energy X-ray absorptiometry (DXA) machines replacing the previous femoral normative reference database with the NHANES III normative data. In addition to changing the normative database (and therefore T-scores) for all regions of the hip, the new software has changed the primary region of interest from the femoral neck to the total hip. In the present study we examined how these changes influence the densitometric diagnosis of osteoporosis in a large clinical referral population (n = 2311, mean age 62.7 years). The patients had spine and hip DXA performed at either of two centers using a Hologic QDR-2000 over a 4-year period. T-scores were derived for each patient using both previous and current young normal reference databases. Intraindividual differences in T-scores were calculated. The prevalence of osteoporosis based on the two normative databases and the difference between the prevalence was calculated for each skeletal site. The average paired difference between current and previous T-scores at femoral neck is 0.64, the difference increasing with age. Using the new normative database, the percentage of osteoporotic patients decreases from 49% of all patients at the femoral neck to 28% at the femoral neck and 20% at the total hip. In conclusion, the densitometric diagnosis of osteoporosis will be affected in a significant proportion of women as a result of the implementation of the new hip normative database supplied by Hologic. Whether this will translate into fewer patients being treated remains to be seen.

Absorptiometry, Photon

Bone mineral density measurements using the Hologic QD2000 in 175 Singaporean women aged 20-80.

AIMS: The first aim is to obtain a reference database of bone mineral density (BMD) measurements for Singaporean women across different age groups and to compare this with an American database using the same machine. The second is to study the lifestyle factors that may influence bone mass in these women. METHODS: Subjects were recruited from hospital staff and their relatives. They needed to fulfil inclusion criteria and those with confounding factors such as being on hormone replacement therapy (HRT) were excluded. A lifestyle questionnaire was administered. RESULTS: Across every five-year age band, the mean BMD measurements of the Singaporean women were 3%-8% lower than their American counterparts at the AP spine and 6%-11% lower at the femoral neck. 40.6% either drank milk, ate cheese or took calcium supplements everyday. 16.6% did some form of weight bearing exercise at least once a week. CONCLUSION: A local reference database is needed for the accurate interpretation of BMD measurements as there is significant variation compared to Caucasian values.

Absorptiometry, Photon

18F-FDG PET studies in patients with extratemporal and temporal epilepsy: evaluation of an observer-independent analysis.

UNLABELLED: The aim of this study was to evaluate an observer-independent analysis of 18F-fluorodeoxyglucose (FDG) PET studies in patients with temporal or extratemporal epilepsy. METHODS: Twenty-seven patients with temporal epilepsy and 22 patients with extratemporal epilepsy were included in the study. All patients with temporal epilepsy and 7 patients with extratemporal epilepsy underwent surgical treatment. In patients who showed significant postoperative improvement (temporal, n = 23; extratemporal, n = 6), the epileptogenic focus was assumed to be located in the area of surgical resection. In extratemporal epilepsy patients who did not undergo surgery, the focus localization was determined using a combination of semiology, ictal and interictal electroencephalography, [99mTc]ethyl cysteinate dimer SPECT, MRI and [11C]flumazenil PET. Visual analysis was performed by two experienced and two less experienced blinded observers using sagittal, axial and coronal images. In the automated analysis after anatomic standardization and generation of three-dimensional stereotactic surface projections (SSPs), a pixelwise comparison of 18F-FDG uptake with an age-matched reference database (n = 20) was performed, resulting in z score images. Pixels with the maximum deviation were detected, summarized and attached to one of 20 predefined surface regions of interest. For comparison with 18F-FDG PET and MR images, three-dimensional overlay images were generated. RESULTS: In patients with temporal epilepsy, the sensitivity was comparable for visual and observer-independent analysis (three-dimensional SSP 86%, experienced observers 86%-90%, less experienced observers 77%-86%). In patients with extratemporal epilepsy, three-dimensional SSP showed a significantly higher sensitivity in detecting the epileptogenic focus (67%) than did visual analysis (experienced 33%-38%, each less experienced 19%). In temporal lobe epilepsy, there was moderate to good agreement between the localization found with three-dimensional SSP and the different observers. In patients with extratemporal epilepsy, there was a high interobserver variability and only a weak agreement between the localization found with three-dimensional SSP and the different observers. Although three-dimensional SSP detected multiple lesions more often than visual analysis, the determination of the highest deviation from the reference database allowed the identification of the epileptogenic focus with a higher accuracy than subjective criteria, especially in extratemporal epilepsy. CONCLUSION: Three-dimensional SSP increases sensitivity and reduces observer variability of the analysis of 18F-FDG PET images in patients with extratemporal epilepsy and is, therefore, a useful tool in the evaluation of this patient group. The benefit of this analytical approach in patients with temporal epilepsy is less apparent.

Adolescent