PubMed HealthSearch

SEARCH · PubMed Health

Results for “database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2Linked to original sources

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software

Maintenance of a nutrient database for clinical trials.

Maintenance of a nutrient database for use in dietary analysis for clinical trials and other medical research studies is described. The database, maintained at the University of Minnesota's Nutrition Coordinating Center (NCC), has been used to calculate dietary intake data for a wide range of diet-disease related investigations including studies on cardiovascular disease, hypertension, cancer, gastroenterology, and osteoporosis. Potential sources of error associated with nutrient databases are identified. Criteria are provided for the selection of a nutrient database to meet study objectives and to minimize the potential for errors and inconsistencies. NCC database maintenance procedures, designed to provide updated and verified nutrient calculations for clinical research, involve adherence to standardized procedures for all aspects of database maintenance including data selection, imputations, quality control, recipe calculations, and documentation. By maintaining multiple versions of the database, the NCC is able to update and expand a working version of the database while providing database stability for individual research studies.

Clinical Trials as Topic

Construction of validated, non-redundant composite protein sequence databases.

A strategy has been developed for the construction of a validated, comprehensive composite protein sequence database. Entries are amalgamated from primary source data bases by a largely automated set of processes in which redundant and trivially different entries are eliminated. A modular approach has been adopted to allow scientific judgement to be used at each stage of database processing and amalgamation. Source databases are assigned a priority depending on the quality of sequence validation and commenting. Rejection of entries from the lower priority database, in each pairwise comparison of databases, is carried out according to optionally defined redundancy criteria based on sequence segment mismatches. Efficient algorithms for this methodology are embodied in the COMPO software system. COMPO has been applied for over 2 years in construction and regular updating of the OWL composite protein sequence database from the source databases NBRF-PIR, SWISS-PROT, a GenBank translation retrieved from the feature tables, NBRF-NEW, NEWAT86, PSD-KYOTO and the sequences contained in the Brookhaven protein structure databank. OWL is part of the ISIS integrated data resource of protein sequence and structure [Akrigg et al. (1988) Nature, 335, 745-746]. The modular nature of the integration process greatly facilitates the frequent updating of OWL following releases of the source databases. The extent of redundancy in these sources is revealed by the comparison process. The advantages of a robust composite database for sequence similarity searching and information retrieval are discussed.

Amino Acid Sequence

An object-based architecture for biomedical expert database systems.

Objects play a major role in both database and artificial intelligence research. In this paper, we present a novel architecture for expert database systems that introduces an object-based interface between relational databases and expert systems. We exploit a semantic model of the database structure to map relations automatically into object templates, where each template can be a complex combination of join and projection operations. Moreover, we arrange the templates into object networks that represent different views of the same database. Separate processes instantiate those templates using data from the base relations, cache the resulting instances in main memory, navigate through a given network's objects, and update the database according to changes made at the object layer. In the context of an immunologic-research application, we demonstrate the capabilities of a prototype implementation of the architecture. The resulting model provides enhanced tools for database structuring and manipulation. In addition, this architecture supports efficient bidirectional communication between database and expert systems through the shared object layer.

Database Management Systems

A model for database design.

Computerized databases can facilitate several types of occupational therapy research. The value and usefulness of any database, however, is dependent on how well it has been designed. In this paper, a systematic, sequential-process model for the development of a computerized database is introduced. Each component of the model is illustrated by examples of its application to the actual design of a database for a community agency that provides occupational therapy services. The model focuses on issues related to the development of the contents of a database rather than on computer hardware and software. The issues addressed by the model include decisions about the purpose of the database, selection of the variables, and identification of the most appropriate measures with which to operationalize these variables. Content-related development issues have been given little attention in the literature, yet their neglect typically results in important limitations on the usefulness of a database. Therefore, this paper provides a set of guidelines for occupational therapists planning to establish a database for facilitating research.

Databases, Factual

BACOMP--database of bioactive compounds for structure-activity relationship.

BACOMP database is presented for structure-activity relationship (SAR) investigations; it was realized on a BK-1300 general purpose microcomputer using the MICRO-SETOR network database management system. Some general considerations of database design are given and the models and facilities for a development of a microcomputer-based SAR oriented database are described. The database contains the following information for the bioactive compounds: chemical structures, biological activities, trade names, reference numbers and information sources. For computer representation of chemical structures SAR oriented language is used. The database software includes: system software, data capture and data editing software, information retrieval and data processing applications. The software development is done in FORTRAN IV and MACRO assembly language. The programs are written in a completely interactive mode. The information retrieval software includes 12 functions giving an information for the database as a whole and for a single compound as well. The data processing software includes 8 functions for finding common structural fragments among compounds with similar biological activity and for estimating a structural similarity between different compounds. The functions are selected from corresponding screen MENUs. The function realization results are framed as appropriate screen formats and the receipt of hardcopies is available. The database can be used to develop predictive methods in respect of the investigated biological activity.

Drug Information Services

Quantitative exploration of the REF52 protein database: cluster analysis reveals the major protein expression profiles in responses to growth regulation, serum stimulation, and viral transformation.

Quantitative protein databases reveal the response of cells to experimental variables, such as exposure to growth factors or transfection with a transforming gene. The nature of the response depends on the type of cell and its internal state at the time of the stimulus. By constructing a protein database to study a given cell line, we can better understand the differentiated state of the cell, the growth regulatory mechanisms it employs, the particular mechanisms it uses to cope with its environment, and the ways these mechanisms may have been compromised through mutation or transformation. The REF52 database is a quantitative database designed to study growth control and transformation in a well-defined family of normal and transformed rat cell lines. The database, which has been described and analyzed elsewhere (J. I. Garrels and B. R. Franza, J. Biol. Chem. 1989, 264, 5283-5298 and J. I. Garrels and B. R. Franza, J. Biol. Chem. 1989, 264, 5299-5312) is further explored here using cluster analysis. This method reveals the most common protein expression profiles for each series of two-dimensional gels without requiring any prior hypothesis or queries on the part of the investigator. This study reveals, for each experiment, large and small clusters of protein expression profiles, most of which have readily apparent biological meaning. For example, large clusters of proteins induced or repressed during growth to confluence have been revealed, and several clusters of transformation-sensitive proteins reveal differential effects of transformation by DNA- and RNA-tumor viruses. This analysis extends our earlier quantitative explorations of the REF52 protein database and helps to show how such a database can be used to provide context and guidance for molecular studies of regulation in a given cell system.

Algorithms

Database of homology-derived protein structures and the structural meaning of sequence alignment.

The database of known protein three-dimensional structures can be significantly increased by the use of sequence homology, based on the following observations. (1) The database of known sequences, currently at more than 12,000 proteins, is two orders of magnitude larger than the database of known structures. (2) The currently most powerful method of predicting protein structures is model building by homology. (3) Structural homology can be inferred from the level of sequence similarity. (4) The threshold of sequence similarity sufficient for structural homology depends strongly on the length of the alignment. Here, we first quantify the relation between sequence similarity, structure similarity, and alignment length by an exhaustive survey of alignments between proteins of known structure and report a homology threshold curve as a function of alignment length. We then produce a database of homology-derived secondary structure of proteins (HSSP) by aligning to each protein of known structure all sequences deemed homologous on the basis of the threshold curve. For each known protein structure, the derived database contains the aligned sequences, secondary structure, sequence variability, and sequence profile. Tertiary structures of the aligned sequences are implied, but not modeled explicitly. The database effectively increases the number of known protein structures by a factor of five to more than 1800. The results may be useful in assessing the structural significance of matches in sequence database searches, in deriving preferences and patterns for structure prediction, in elucidating the structural role of conserved residues, and in modeling three-dimensional detail by homology.

Amino Acid Sequence

The effect of a multiple literature database search--a numerical evaluation in the domain of Japanese life science.

In literature database searching, we show that it is necessary to use plural databases for a more improved search. We also compare the results of a single database search with that of multiple database search in the domain of Japanese life sciences. We searched the MEDLINE and EMBASE using the same search terms. There were some differences in the results, owing to differences in the journals and recording methods. We herein show some of the differences in the journals contained in both databases. Furthermore, we show the differences in the number of papers derived from the same journal. Next, as an example of a practical search, we selected some universities in Japan, searched both databases regarding papers published from these universities and then merged the results by hand. According to our results, only 63% of all papers were common to both databases.

Biological Science Disciplines

The free form database program as a research tool.

Two types of database programs for the IBM compatible personal computer (PC) are described: the fixed form database and the free form database. The use of the latter in compiling bibliographic databases and in content analysis of interview transcripts is described. Other uses for the free form database program are also discussed. It is suggested that the free form database has advantages over some other custom-made analysis programs in terms of its simplicity and ease of use. It is also suggested that the free form database can be a useful tool for nurse educators and students.

Databases, Bibliographic

An improved FORTRAN 77 recombinant DNA database management system with graphic extensions in GKS.

We have improved an existing clone database management system written in FORTRAN 77 and adapted it to our software environment. Improvements are that the database can be interrogated for any type of information, not just keywords. Also, recombinant DNA constructions can be represented in a simplified 'shorthand', whereafter a program assembles the full nucleotide sequence from the contributing fragments, which may be obtained from nucleotide sequence databases. Another improvement is the replacement of the database manager by programs, running in batch to maintain the databank and verify its consistency automatically. Finally, graphic extensions are written in Graphical Kernel System, to draw linear and circular restriction maps of recombinants. Besides restriction sites, recombinant features can be presented from the feature lines of recombinant database entries, or from the feature tables of nucleotide databases. The clone database management system is fully integrated into the sequence analysis software package from the Pasteur Institute, Paris, and is made accessible through the same menu. As a result, recombinant DNA sequences can directly be analysed by the sequence analysis programs.

Algorithms

Automated assembly of protein blocks for database searching.

A system is described for finding and assembling the most highly conserved regions of related proteins for database searching. First, an automated version of Smith's algorithm for finding motifs is used for sensitive detection of multiple local alignments. Next, the local alignments are converted to blocks and the best set of non-overlapping blocks is determined. When the automated system was applied successively to all 437 groups of related proteins in the PROSITE catalog, 1764 blocks resulted; these could be used for very sensitive searches of sequence databases. Each block was calibrated by searching the SWISS-PROT database to obtain a measure of the chance distribution of matches, and the calibrated blocks were concatenated into a database that could itself be searched. Examples are provided in which distant relationships are detected either using a set of blocks to search a sequence database or using sequences to search the database of blocks. The practical use of the blocks database is demonstrated by detecting previously unknown relationships between oxidoreductases and by evaluating a proposed relationship between HIV Vif protein and thiol proteases.

Algorithms

Expanding vaginal microbiome pangenomes via a custom MIDAS database reveals Lactobacillus crispatus accessory genes associated with cervical dysplasia.

The vaginal microbiome plays a central role in reproductive health. Vaginal microbiome dysbiosis is associated with many adverse reproductive health outcomes, but most studies have focused on associations at the species level. The potential contribution of intraspecies microbial variation, especially gene content differences across bacterial strains, remains underexplored in reproductive health contexts. The Metagenomic Intra-Species Diversity Analysis (MIDAS) framework enables such analyses, but depends on comprehensive reference databases. We constructed a MIDAS-compatible pangenome database from over 18,000 genomes in the Vaginal Microbiome Genome Collection (VMGC). Compared to the Genome Taxonomy Database (GTDB)-derived reference, the VMGC-derived database expanded the pangenomes of prevalent vaginal species, better capturing vaginal-specific intraspecies diversity. Applying this database to vaginal samples from a cervical dysplasia cohort, we identified 13 Lactobacillus crispatus accessory genes significantly associated with cervical dysplasia, including a HicAB toxin-antitoxin system, three transcriptional regulators, and three phage-derived genes. These findings highlight the utility of body site-specific reference resources and shotgun metagenomic sequencing for uncovering intraspecies microbial variation relevant to reproductive health.IMPORTANCEThe vaginal microbiome plays a critical role in reproductive health, and different bacteria from the same species can carry different genes that influence how the strains interact with the host and other microbes. These strain-level differences are often overlooked when microbiomes are analyzed only at the species level. Existing genomic reference databases are heavily biased toward gut and environmental bacteria, leaving the genetic diversity of vaginal microbes understudied. We built a specialized reference database from over 18,000 vaginal bacterial genomes that better reflects this diversity. We then applied this resource to quantify gene-level variation in vaginal samples from a cervical dysplasia cohort. Focusing on Lactobacillus crispatus, a prevalent and often beneficial vaginal species, we identified 13 genes that were more common in women with cervical dysplasia than in controls. This work demonstrates that body site-specific genomic resources are essential for uncovering strain-level bacterial differences relevant to reproductive health.

Lactobacillus crispatus

Carcinogenicity evaluations and ongoing studies: the IARC databases.

Many thousands of chemicals are produced industrially and many more occur naturally. Information on the toxicology of these chemicals is often minimal or absent. The International Agency for Research on Cancer (IARC) has published evaluations of the carcinogenic risk to humans of over 700 chemicals, groups of chemicals, and complex mixtures as a regular series of monographs. A database has been created containing summaries of all the relevant epidemiological, animal carcinogenicity, and other relevant biological data for each chemical or mixture evaluated. Additional databases have been created for ongoing epidemiological studies of cancer in humans and for long-term carcinogenicity studies in rodents, as well as a database containing information on genotoxic and related effects of chemicals. Some of these databases have been published in print form. IARC now plans to publish them electronically, together with other databases, in the form of a CDROM (compact disk, read-only memory). The objective will be to make the entire IARC database of cancer information as widely available as possible in an integrated format conducive to efficient and combined exploitation of all the component databases.

Animals

[A new database system for radiological reports].

We have designed and developed a new database system to facilitate automatic feedback of the content of radiology reports to radiologists. The prototype of this database system has been implemented in the RGSS-IDJ, a developmental computer system that applies artificial intelligence methods to a reporting system. This prototype system was constructed to test the feasibility of overcoming the limitations of conventional database systems. The new database system is based on our semantic model for radiology reports and is able to treat data with unnormalized relations. Operations specific to our database system include the ability to acquire information about a set of reports that contains any semantic expression included in the lexicon and the ability to obtain the expressions that belong to a set of several semantic expressions in the reports. Thus, our new database system will offer a more powerful tool for analyzing the content of reports than conventional database systems.

Databases, Bibliographic

Up-to-date, and taxonomy-curated mcrA reference databases for methanogen community profiling.

The methyl-coenzyme M reductase subunit alpha gene (mcrA) is an important phylogenetic marker for high throughput ecological profiling of methanogenic archaea, central to industrial biological methane production and greenhouse gas emissions. Yet, dedicated reference databases predate current relevant NCBI sequence accumulation and archaeal taxonomic revision. We present three updated mcrA reference databases: (i) one derived from NCBI-catalogued methanogen genomes (1572 sequences); (ii) a database built by expansion of a previously published reference dataset, leveraging the NCBI nucleotide collection (27,942 sequences); (iii) a curated-taxonomy version of the latter. The updated amplicon databases provide a ∼ 3.5-fold sequence richness expansion, extend genus-level richness from 31 to 83 taxa, more than 4-fold species-level richness, and incorporate novel lineages compared with the previous reference dataset (e.g. Thermoplasmatota-encompassed). All databases were formatted to support analysis with relevant contemporary software pipelines and packages. Overall, the generated databases facilitate a highly improved characterization of methanogen diversity and ecology.

Archaea

Characteristics of the U.S. EPA's Office of Pesticide Programs' toxicity information databases.

The United States Environmental Protection Agency's Office of Pesticide Programs (OPP) requires that data from toxicity testing be submitted to the OPP to support the registration of pesticide chemicals. Once the toxicity data are submitted, they are entered into various toxicity databases. The studies are listed in an archival database to catalog and allow retrieval of the study for review. Reviews of toxicity studies are then placed into a separate database that can be retrieved to support a regulatory position. Toxicity information for health effects other than cancer and gene mutations from chronic exposure is reviewed through a reference dose (RfD) approach, and these decisions and supporting data are entered into an RfD database. Carcinogenicity data are reviewed by a peer review process, and these decisions are entered into a newly developed database to show the regulatory decision with supporting data. The mutagenicity data are reviewed and acceptable data are entered into the Genetic Activity Profile system to catalog and display the submitted information. These databases contain the information used for hazard evaluations as part of the OPP review of pesticide chemicals.

Animals

Non-sequence databases for biological activity and physicochemical properties.

A biological activity database and a physicochemical property database are described. They are intended to complement the protein sequence database of PIR-International. The Biological Activity Database and the Physicochemical Property Database contain information regarding the biological activity and the physicochemical properties of proteins, respectively. In addition they also provide information about wild-type molecules with which information concerning variant molecules may be compared. Data on artificial variant molecules are stored in the Artificial Variant Database which is described separately.

Amino Acid Sequence