PubMed HealthSearch

SEARCH · PubMed Health

Results for “database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software

Databases and software for the analysis of mutations in the human p53 gene, the human hprt gene and the lacZ gene in transgenic rodents.

We have created databases and software applications for the analysis of DNA mutations in the human p53 gene, the human hprt gene and the rodent transgenic lacZ locus. The databases themselves are stand-alone dBase files and the software for analysis of the databases runs on IBM- compatible computers. The software created for these databases permits filtering, ordering, report generation and display of information in the database. In addition, a significant number of routines have been developed for the analysis of single base substitutions. One method of obtaining the databases and software is via the World Wide Web (WWW). Open home page http://sunsite.unc.edu/dnam/mainpage.ht ml with a WWW browser. Alternatively, the databases and programs are available via public ftp from anonymous@sunsite.unc.edu. There is no password required to enter the system. The databases and software are found in subdirectory pub/academic/biology/dna-mutations. Two other programs are available at the WWW site, a program for comparison of mutational spectra and a program for entry of mutational data into a relational database.

Animals

Identifying the active general practice workforce in one division of general practice: the utility of public domain databases.

OBJECTIVE: To identify the non-specialist medical practitioner workforce engaged in active general practice in the region served by the Division of General Practice-Northern Tasmania and to determine the usefulness of public domain databases for enumeration of individual non-specialists providing general practice services. METHODS: A masterlist of the active general practice workforce was compiled by obtaining the names and addresses/postcodes of all non-specialist medical practitioners who were listed in at least one of nine public domain databases and who were confirmed by selected local medical practitioners to be in active general practice in the three months prior to 30 June 1994. This masterlist was used in calculating the sensitivity and positive predictive value (PPV) of each of the nine databases for enumerating non-specialist practitioners in active general practice. RESULTS: Combining the databases resulted in a list of 475 practitioners, which was refined to 139 practitioners who, by our criteria, were in active general practice. Databases had a range of sensitivities and PPVs, but those with high sensitivity tended to have low PPVs, and vice versa. The most useful database for enumerating these practitioners was the mailing list for Australian Family Physician (sensitivity, 94%; PPV, 0.79). CONCLUSIONS: When used alone, no single database had both high sensitivity and high positive predictive value for identifying the active general practice workforce. Combining multiple databases may improve precision. Developing methods to identify recent departures from local active practice has the potential to improve the PPV of existing highly sensitive databases.

Australia

Performance of online biomedical databases in rheumatology.

OBJECTIVE: To compare the performance of MEDLINE, EMBASE, and BIOSIS in selected rheumatology topics. METHODS: Online literature searches were conducted with regard to the epidemiology of rheumatoid arthritis (RA), systemic lupus erythematosus (SLE), and ankylosing spondylitis (AS), as well as for 3 specific questions representing clinical, clinical/laboratory, and therapeutic topics in rheumatology. Total number of citations retrieved, type and language of publication, percentage of contribution from rheumatology journals, and degree of overlap among the databases were recorded. Publications retrieved for the 3 specific questions were also graded for relevance. RESULTS: For 1991, each online biomedical database (OBD) retrieved more than 1,100 citations for RA, over 600 for SLE, and over 110 for AS. For the epidemiology subtopic, fewer than 25% of the citations were retrieved by more than one of the databases. About 3/4 of the citations obtained for the specific search questions were retrieved by a single database. No major differences were observed among databases in relation to number of relevance of citations retrieved. Over 60% of the papers assessed had low relevance in relation to the topic of the search. Efficiency was estimated as the percentage of all relevant citations retrieved by each OBD. Results varied according to the topic, but in most cases each database retrieved at least 50% of the relevant citations. About 45% of the citations retrieved for the 3 search questions were published in nonrheumatology journals. CONCLUSION: No database was superior in all respects. The majority of the citations were retrieved by a single database. A high percentage of the articles retrieved were not relevant, implying low specificity. If a comprehensive online search in rheumatology is required, 2 or more databases should be utilized.

Arthritis, Rheumatoid

Maintaining patient confidentiality in the public domain Internet Autopsy Database (IAD).

The Internet provides the opportunity of permitting public access to large databases containing patient information that can be shared and utilized by epidemiologists, health planners, and medical researchers. Until now, large databases containing patient information have been held in strict confidence, with database access available only to approved researchers or to researchers with access limited to only specific portions of the database. The Internet Autopsy Database (IAD) consists of demographic and pathologic data from over 49,000 autopsies contributed by over a dozen academic medical institutions. Each autopsy record in the public database consists of a uniform set of demographics and SNOMED-compatible terms. To make the database publicly available, a strategy had to be devised that assured the privacy of every person included in the database. A key step involved translating the autopsy facesheets into a listing of SNOMED-compatible terms that effectively eliminated identifying terminology, replacing free text with a generic nomenclature that preserves diagnostic information. The entire database is available on the Internet at: http:@www.med.jhu.edu/pathology/iad.html

Autopsy

Use of large databases for resolving critical care problems.

Large databases allow for rapid access to large volumes of data. To convert raw data to information, large numbers of data points must be correlated into a descriptive pattern that can be interpreted by the user. Databases must be constructed so as to allow reliable extraction of the raw data into a format that supports analysis of events in a meaningful, objective, and reproducible manner. Databases must be responsive to a variety of users. They must not demand unrealistic amounts of effort on those responsible for data entry. Standard protocols in various stages of development will make databases easier to use and more reliable. Database management tools such as the Internet and the National Library of Medicine will become more integrated into the practice of critical care medicine at all levels, including administration, clinical care, and research. This article provides an overview of the capabilities and difficulties associated with large databases. The major areas of use of large databases in the hospital setting are administration, bibliographic, patient care, research, and education. Each of these areas has different requirements and is supported by different types of databases. The advantages and disadvantages of linear, relational, and object-oriented databases are discussed. Issues relating to methods of data entry and the accuracy and reliability of data are discussed. The challenges involving integration of various sources of data and the interfacing of devices are reviewed.

Critical Care

Maintenance of a nutrient database for clinical trials.

Maintenance of a nutrient database for use in dietary analysis for clinical trials and other medical research studies is described. The database, maintained at the University of Minnesota's Nutrition Coordinating Center (NCC), has been used to calculate dietary intake data for a wide range of diet-disease related investigations including studies on cardiovascular disease, hypertension, cancer, gastroenterology, and osteoporosis. Potential sources of error associated with nutrient databases are identified. Criteria are provided for the selection of a nutrient database to meet study objectives and to minimize the potential for errors and inconsistencies. NCC database maintenance procedures, designed to provide updated and verified nutrient calculations for clinical research, involve adherence to standardized procedures for all aspects of database maintenance including data selection, imputations, quality control, recipe calculations, and documentation. By maintaining multiple versions of the database, the NCC is able to update and expand a working version of the database while providing database stability for individual research studies.

Clinical Trials as Topic

An assessment of data quality in the Vermont-Oxford Trials Network database.

The Vermont-Oxford Trials Network is a voluntary collaborative research group of neonatologists that maintains a database for very low birthweight infants (501-1500 g). The database (1) provides core data for randomized trials, (2) serves as a resource for outcomes research in neonatology, and (3) generates quality management reports for participating sites. To assess the reliability of this database and to determine the sources of error, we reviewed 635 medical records chosen at random from among the 4341 eligible infants born at 40 participating data generating sites during an 18-month period beginning January 1, 1990. The estimated frequencies of disagreement between the medical record and database for each of the 10 data items studied and the standard errors of the estimates (in parentheses) were: date of birth 1.3% (0.4), date of admission 2.5% (0.6), date of discharge 8.8% (1.0), birthweight (difference > 50 g) 2.9% (0.6), location of birth (inborn or outborn) 2.1% (0.5), multiple birth 2.2% (0.5), cesarean section 2.5% (0.6), gender 2.1% (0.5), status 28 days after birth 3.4% (0.6), final status 2.9% (0.6). The overall proportions and mean values for items in the database were close to the estimated values based on the random sample of records. There were a total of 247 disagreements between the database and the medical records in the sample. Twenty-three were due to data keying errors. Two hundred twenty-four were due to errors in transcription or interpretation. The rate of data keying errors decreased from over 50 errors per 10,000 fields to less than 15 errors per 10,000 fields when specific quality control procedures, including visual inspection, were instituted. Data keying errors accounted for 13.7% of all disagreements between the database and medical record before improved data entry methods were introduced, and only 3.7% of all errors after they were introduced. We concluded that the Vermont-Oxford Trials Network Database is reliable. Data keying errors have been reduced by the introduction of additional quality control measures. Further reductions in database errors will require measures aimed at minimizing transcription or interpretation errors by individuals completing the data forms.

Computer Communication Networks

Construction of validated, non-redundant composite protein sequence databases.

A strategy has been developed for the construction of a validated, comprehensive composite protein sequence database. Entries are amalgamated from primary source data bases by a largely automated set of processes in which redundant and trivially different entries are eliminated. A modular approach has been adopted to allow scientific judgement to be used at each stage of database processing and amalgamation. Source databases are assigned a priority depending on the quality of sequence validation and commenting. Rejection of entries from the lower priority database, in each pairwise comparison of databases, is carried out according to optionally defined redundancy criteria based on sequence segment mismatches. Efficient algorithms for this methodology are embodied in the COMPO software system. COMPO has been applied for over 2 years in construction and regular updating of the OWL composite protein sequence database from the source databases NBRF-PIR, SWISS-PROT, a GenBank translation retrieved from the feature tables, NBRF-NEW, NEWAT86, PSD-KYOTO and the sequences contained in the Brookhaven protein structure databank. OWL is part of the ISIS integrated data resource of protein sequence and structure [Akrigg et al. (1988) Nature, 335, 745-746]. The modular nature of the integration process greatly facilitates the frequent updating of OWL following releases of the source databases. The extent of redundancy in these sources is revealed by the comparison process. The advantages of a robust composite database for sequence similarity searching and information retrieval are discussed.

Amino Acid Sequence

Uses of clinical databases.

Clinical databases consist of observational data collected on patients who meet specific criteria. The uses of these databases depend on whether the observations are drawn from a single institution, multiple clinical centers, or are population-based. Single institution databases frequently are used to profile patient accrual. In cases of rare diseases or unusual procedures, multicenter databases are required to amass sufficient information for study. Multicenter databases can be used to address issues related to intercenter variation and to develop statistical models to predict outcome based on prognostic factors. Population-based databases are required to assess incidence, prevalence, and mortality rates of disease. Although the inferences that can be drawn from observational data are limited by selection bias, clinical databases are valuable tools in planning clinical research. Clearly, however, resources are required to develop and maintain clinical databases. In an era when much research funding is directed at hypothesis-driven research, the importance of these clinical databases in developing clinical research hypotheses should not be overlooked.

Information Systems

A strategy for database interoperation.

To realize the full potential of biological databases (DBs) requires more than the interactive, hypertext flavor of database interoperation that is now so popular in the bioinformatics community. Interoperation based on declarative queries to multiple network-accessible databases will support analyses and investigations that are orders of magnitude faster and more powerful than what can be accomplished through interactive navigation. I present a vision of the capabilities that a query-based interoperation infrastructure should provide, and identify assumptions underlying, and requirements of, this vision. I then propose an architecture for query-based interoperation that includes a number of novel components of an information infrastructure for molecular biology. These components include a knowledge base that describes relationships among the conceptualizations used in different biological databases, a module that can determine the DBs that are relevant to a particular query, a module that can translate a query and its results from one conceptualization to another, a collection of DB drivers that provide uniform physical access to different database management systems, a suite of translators that can interconvert among different database schema languages, and a database that describes the network location and access methods for biological databases. A number of the components are translators that bridge the heterogeneities that exist between biological DBs at several different levels, including the conceptual level, the data model, the query language, and data formats.

Artificial Intelligence

The Protein Disease Database of human body fluids: II. Computer methods and data issues.

The Protein Disease Database (PDD) is a relational database of proteins and diseases. With this database it is possible to screen for quantitative protein abnormalities associated with disease states. These quantitative relationships use data drawn from the peer-reviewed biomedical literature. Assays may also include those observed in high-resolution electrophoretic gels that offer the potential to quantitate many proteins in a single test as well as data gathered by enzymatic or immunologic assays. We are using the Internet World Wide Web (WWW) and the Web browser paradigm as an access method for wide distribution and querying of the Protein Disease Database. The WWW hypertext transfer protocol and its Common Gateway Interface make it possible to build powerful graphical user interfaces that can support easy-to-use data retrieval using query specification forms or images. The details of these interactions are totally transparent to the users of these forms. Using a client-server SQL relational database, user query access, initial data entry and database maintenance are all performed over the Internet with a Web browser. We discuss the underlying design issues, mapping mechanisms and assumptions that we used in constructing the system, data entry, access to the database server, security, and synthesis of derived two-dimensional gel image maps and hypertext documents resulting from SQL database searches.

Body Fluids

Derivation of rules for comparative protein modeling from a database of protein structure alignments.

We describe a database of protein structure alignments as well as methods and tools that use this database to improve comparative protein modeling. The current version of the database contains 105 alignments of similar proteins or protein segments. The database comprises 416 entries, 78,495 residues, 1,233 equivalent entry pairs, and 230,396 pairs of equivalent alignment positions. At present, the main application of the database is to improve comparative modeling by satisfaction of spatial restraints implemented in the program MODELLER (Sali A, Blundell TL, 1993, J Mol Biol 234:779-815). To illustrate the usefulness of the database, the restraints on the conformation of a disulfide bridge provided by an equivalent disulfide bridge in a related structure are derived from the alignments; the prediction success of the disulfide dihedral angle classes is increased to approximately 80%, compared to approximately 55% for modeling that relies on the stereochemistry of disulfide bridges alone. The second example of the use of the database is the derivation of the probability density function for comparative modeling of the cis/trans isomerism of the proline residues; the prediction success is increased from 0% to 82.9% for cis-proline and from 93.3% to 96.2% for trans-proline. The database is available via electronic mail.

Amino Acid Sequence

An object-based architecture for biomedical expert database systems.

Objects play a major role in both database and artificial intelligence research. In this paper, we present a novel architecture for expert database systems that introduces an object-based interface between relational databases and expert systems. We exploit a semantic model of the database structure to map relations automatically into object templates, where each template can be a complex combination of join and projection operations. Moreover, we arrange the templates into object networks that represent different views of the same database. Separate processes instantiate those templates using data from the base relations, cache the resulting instances in main memory, navigate through a given network's objects, and update the database according to changes made at the object layer. In the context of an immunologic-research application, we demonstrate the capabilities of a prototype implementation of the architecture. The resulting model provides enhanced tools for database structuring and manipulation. In addition, this architecture supports efficient bidirectional communication between database and expert systems through the shared object layer.

Database Management Systems

Issues in designing a student database.

Many larger colleges of nursing have centralized student databases in which all elements of students' records are collated. The smaller department or college may leave the design of such databases to course leaders. This paper highlights stages and considerations in the development of student databases. It offers definitions of flat-file and relational databases; it identifies issues in prior planning and then of direct plannings of the database. It offers a discussion of a number of issues related to what sort of information is stored in a student database. It closes with a discussion of the training needs of those staff who are going to use such a database. The principles described in this paper can also be adapted for use in the designing of any health care related database system. The paper also addresses the question of why all colleges are not, already, using such systems.

Computer User Training

Evaluating the options for developing databases to support research-based medicine at the NHS Centre for Reviews and Dissemination.

The paper presents the results of a questionnaire survey and user interface review group experiments to determine the value and ease of use of the Database of Abstracts of Reviews of Effectiveness (DARE) and the NHS Economic Evaluations Database (NEED). The results are interpreted with other recent studies of the use of electronic databases, including the NHS Research Register User Requirements Specification. The study found that most frequent users of the DARE database tend to use the CD-ROM version. Librarians were found to have the greatest awareness of the databases, with relatively low levels of use by operational NHS Staff. Untrained users found the online databases difficult to access and had erroneous perceptions of the database content which were only realised when queries returned unexpected answers. Experienced users of online information systems tended to want more sophisticated search facilities than inexperienced users. Nearly all users in the review groups wanted to access the databases in conjunction with other information sources, such as the Cochrane Library, Medline and the ACP Journal Club, highlighting the need for cross organisational strategies for the dissemination of research-based information.

CD-ROM

The PIR-International databases.

PIR-International is an association of macromolecular sequence data collection centers dedicated to fostering international cooperation as an essential element in the development of scientific databases. PIR-International is most noted for the Protein Sequence Database. This database originated in the early 1960's with the pioneering work of the late Margaret Dayhoff as a research tool for the study of protein evolution and intersequence relationships; it is maintained as a scientific resource, organized by biological concepts, using sequence homology as a guiding principle. PIR-International also maintains a number of other genomic, protein sequence, and sequence-related databases. The databases of PIR-International are made widely available. This paper briefly describes the architecture of the Protein Sequence Database, a number of other PIR-International databases, and mechanisms for providing access to and for distribution of these databases.

Amino Acid Sequence

Comparing checklists and databases with physicians' ratings as measures of students' history and physical-examination skills.

PURPOSE: To compare two methods of rating students' performances on history and physical examination: (1) by using checklists completed by standardized patients (SPs) and databases completed by students, and (2) by using ratings of students by three physicians for each SP-student encounter. METHOD: Four cases were chosen for the study, and 30 students were examined per case. The students were all in their fourth year at the Southern Illinois University School of Medicine in the spring of 1991. Two of the cases had both checklists and databases, and the remaining two had databases only. Each SP-student encounter was videotaped and was viewed independently by three physicians unfamiliar with the contents of the checklists and databases. The physicians' pooled ratings were then compared with the checklist and database scores. Uncorrected and corrected correlations were obtained, with the generalizability coefficient used as the index of reliability. RESULTS: Interrater generalizability of physicians' ratings was very good, ranging from .65 to .93 for overall ratings. Generalizability of physicians' ratings pooled across the four cases was .85. Checklist scores tended to correlate higher with physicians' ratings than did database scores: across the cases, correlation coefficients between physicians' ratings and checklist scores and database scores were .65 and .39, respectively. CONCLUSION: The checklist scores correlated strongly with the physicians' ratings of history and physical-examination skills, providing some evidence of validity for their use. The checklist scores correlated much better with the physicians' ratings than did the database scores. Possible explanations for this finding are discussed.

Clinical Clerkship