PubMed HealthSearch

SEARCH · PubMed Health

Results for “database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Federated two-dimensional electrophoresis database: a simple means of publishing two-dimensional electrophoresis data.

While a two-dimensional electrophoresis (2-DE) database is a relatively old concept, in recent years it generated renewed interest within the 2-DE community due to two main factors: (i) The high reproducibility of the current 2-DE method allows 2-DE images to be exchanged and compared between laboratories. (ii) The recent development of faster and more powerful techniques for protein identification such as microsequencing, matrix-assisted laser desorption ionization-mass spectrometry (MALDI-MS) and amino acid composition makes the production of reference protein maps and 2-DE databases cost- and time-effective. Additionally, the Internet network's current increase in popularity, combined with the rapid growth of Internet-connected laboratories, provides a straightforward means of publishing and sharing 2-DE data. While a small number of laboratories have already successfully published their data over the net, the increasing number of 2-DE database servers that are currently being set up will sooner or later require some kind of standardization. Unfortunately, standardization can be a long and cumbersome process inevitably leading to undesirable compromises. A federated database offers a simple and efficient way to publish and share 2-DE data without the need for standardization. Taking advantage of Internet protocols such as World Wide Web, they allow each laboratory to maintain their own database and to interconnect it with other similar databases through the use of active cross-references. This paper first presents guidelines for building a federated 2-DE database that may easily be followed by most laboratories. It then briefly reviews the state-of-the-art in networked 2-DE databases, and finally describes the SWISS-2DPAGE database which fully implements the concept of a federated 2-DE database.

Computer Communication Networks

Databases and software for the analysis of mutations in the human p53 gene, the human hprt gene and both the lacI and lacZ gene in transgenic rodents.

We have created databases and software applications for the analysis of DNA mutations at the humanp53gene, the humanhprtgene and both the rodent transgeniclacIandlacZlocus. The databases themselves are stand-alone dBASE files and the software for analysis of the databases runs on IBM-compatible computers. Each database has a separate software analysis program. The software created for these databases permit the filtering, ordering, report generation and display of information in the database. In addition, a significant number of routines have been developed for the analysis of single base substitutions. One method of obtaining the databases and software is via the World Wide Web (WWW). Open the following home page with a Web Browser: http://sunsite.unc.edu/dnam/mainpage.ht ml . Alternatively, the databases and programs are available via public FTP from: anonymous@sunsite.unc.edu . There is no password required to enter the system. The databases and software are found beneath the subdirectory: pub/academic/biology/dna-mutations. Two other programs are available at the site-a program for comparison of mutational spectra and a program for entry of mutational data into a relational database.

Animals

The Jepson Lecture 1991. Clinical databases and surgical research.

Modern science is preoccupied with basic mechanisms at the level of the gene, molecule, atom and fundamental particle. This preoccupation has affected attitudes to medical science. Paradoxically, politicians and Departments of Health are largely concerned with clinical epiphenomena--outcomes, interventions, effectiveness, efficiency--that are often disregarded by medical scientists and viewed without favour by granting bodies. In the surgical sciences, there is renewed interest in clinical research. The small computer and the readily available database with its compatible statistical package have added a new dimension to clinical research. Database design and analysis are as much a part of the surgical investigator's skills as are laboratory techniques. There are simple rules that govern database design. The database should be simple and flexible, but provide an adequate patient profile. Entry should be standardized by precise inclusion criteria and precise definitions. Data input should be numerical whenever possible. The database program should convert simply to a comprehensive statistical program. The clinical database is useful for both observational and investigational studies. Case series, case control studies and cohort studies can all be developed from well maintained databases. In the Department of Surgery at Westmead Hospital, databases have been maintained for 10 years. The hepatobiliary and pancreatic group of databases have led to 34 publications. The liver tumour database has produced 12 studies and nearly 20 papers. The process of development of one study is outlined in detail. The chance observation of an excess incidence of gallstones in patients having regular ultrasound examinations after major abdominal surgery has been confirmed in a case control study.(ABSTRACT TRUNCATED AT 250 WORDS)

Clinical Protocols

Database use in neonatal intensive care units: success or failure.

The purpose of this national survey was to define the extent and features of database use by 445 tertiary level neonatal intensive care nurseries in the United States. Of the 305 centers responding to our survey, 78% had a database in use in 1989 and 15% planned to develop one in the future. Nurseries varied remarkably in the volume of data collected, the amount of time devoted to completing data collection forms, and the personnel involved in data collection. Although data were used primarily for statistical reports (93% of nurseries), quality assurance (73%) and research activities (61%) were also enhanced by database information. Neonatal databases were used to generate reports for the permanent medical record in 38% of centers. Satisfaction with the database was dependent on how useful the database information was to centers which collected and actually used a large volume of information. Overall, nurseries expressed a high degree of confidence in the data they collected, and 65% felt their neonatal database information could be used directly in publication of research. It was disturbing that accuracy of data was not monitored formally by the majority of nurseries. Only 27% of centers followed a routine schedule of data quality assurance, and only 53% had built in error messages for data entry. We caution all who receive database information in the form of morbidity and mortality statistics, clinical reports on patients cared for in neonatal units, and published manuscripts to be attentive to the quality of the data they consume. We feel that future database design efforts need to better address data quality control. Our findings stress the importance and need for immediate efforts to better address database quality control.

Data Collection

A protein secondary structure database (PSS).

A protein secondary structure database (PSS) has been designed to correlate the Protein Sequence Database of the PIR-International with the atomic coordinates and bond connectivities database of the Protein Data Bank in the Brookhaven National Laboratory. The present database includes secondary structures determined by X-ray diffraction analysis, but not predicted structures. The database currently contains data from both the Protein Sequence Database and the Protein Data Bank Database, and will encompass the NMR database in the future. The main characteristics of the database are as follows: (1) the secondary structures, sites, regions and domains of structural interest are displayed together with protein primary structures; and (2) the secondary structure of a desired length of peptide fragment is displayed upon request, as are the peptide fragment(s) that correspond to a defined secondary structure. This database also has software to indicate amino acid pairs having hydrogen bonds and to count the occurrence frequency of each pair as well as the conformational parameters widely used in semi-empirical methods of secondary structure prediction.

Amino Acid Sequence

A profile for molecular biology databases and information resources.

This paper examines the requirements for building database management systems and multi-database information resources to support molecular biology research. The paper profiles the most important features of 16 integrated resources and 102 databases related to molecular biology research. The aspects surveyed in this paper include the nature of information in these databases, their sizes, update properties, cross-references, database management system heterogeneity, geographical distribution, data quality, use of temporal information and level of interpretation. The paper also comments on the access patterns to these databases. Since not all these aspects were available for all databases, specific comparisons sometimes compare fewer than the full 102 databases. Consequently, the same set of databases is not necessarily always being compared with respect to every aspect. The paper is organized primarily according to these comparison aspects and ends with some concluding remarks.

Databases, Bibliographic

The rat liver epithelial (RLE) cell protein database.

Computer databases of rat liver epithelial (RLE) cellular polypeptides have been established using high resolution two-dimensional gel electrophoresis and computer-assisted analysis. Databases have been constructed utilizing both [35S]methionine- and [32P]orthophosphate-labeled as well as silver-stained polypeptides from normal RLE cells. The RLE database, which contains both qualitative and quantitative annotations, includes experiments with normal, chemically and oncogene transformed as well as spontaneously transformed cell lines. A total of 2537 [35S]methionine-labeled polypeptides from whole cell lysates (1920 acidic and 617 basic, separated in the first dimension using isoelectric focusing and nonequilibrium pH gradient electrophoresis, respectively) were analyzed and databases constructed using the Elsie 5 gel analysis system. To increase the "viewing window" and hence the usefulness of the RLE database, subcellular fractionation of whole cell preparations was performed and high resolution two-dimensional maps of the individual subcellular components were constructed. Databases representing 1229 cytosolic, 1539 acidic and 674 basic nuclear, 1746 membrane-associated, 415 mitochondrial, 773 in vitro translated and 350 phosphoproteins were established from these maps. The RLE databases contain the Elsie 5 identification number, protein name (if known), molecular weight and pI information, quantitative and spot shape data, and specific information regarding transformation-sensitive, growth-related (exponentially proliferating versus confluent) cell populations as well as those polypeptides modulated by specific growth factors. The RLE databases represent initial efforts toward the establishment of comprehensive databases of rat liver proteins and serve as a vital resource for on-going as well as future studies regarding the regulation of growth and differentiation as well as transformation of RLE cells.

Animals

An efficient disk based data structure for rapid searching of quantitative two-dimensional gel databases.

Fast access of two-dimensional (2-D) gel quantitative databases is important for rapid searching for protein differences between sets of 2-D gels from an experiment. The GELLAB-II system organizes corresponding spots from the gels in the database into reference or "Rspot" sets. These Rspot numeric names index fixed regions in the paged composite gel database file. This is adequate for an existing database, but has several problems. (i) Building the initial database requires guessing how much disk space to pre-allocate for each corresponding spot (i.e. spots from different gels). If it ever runs out of pre-allocated space during this process, it must expand the size of each corresponding set of spots copying the old database data into the new in-place on the disk. (ii) When adding new gels or editing the database, if a new spot is created, the system may also go into this expansion mode. The time spent and wasted disk space can be appreciable--depending on the size of the database (order of 100 gel database). (iii) Because each set of corresponding spots is the same size, we waste space in most spot sets since they do not require the additional space a few spot sets require which contain additional fragmented spots. We present a new low-level disk object-based structure and algorithm, paged indexed buckets (PIB), which optimizes disk space usage while having similar retrieval speed to the original method.

Algorithms

Construction of HSC-2DPAGE: a two-dimensional gel electrophoresis database of heart proteins.

The dissemination of information relating to the characterisation of proteins from two-dimensional electrophoresis (2-DE) gel databases is essential for their effective utilisation in the study of protein expression in cell biology. Since the inception of the World Wide Web and the pioneering development of SWISS-2DPAGE as a tool for retrieving information on proteins separated by 2-DE, the Internet has become the method of choice for disseminating and accessing information on 2-DE protein databases. At Harefield we have established HSC-2DPAGE which is an advanced interface for accessing protein database relating to heart disease. The Web site currently includes databases of proteins from human, dog and rat ventricular tissue and a human endothelial cell line. The databases are searchable individually or as a whole by remote keyword searches. Each database is represented by both synthetic (computer generated) and real (scanned gel) clickable images upon which characterised protein spots are highlighted by hyperlinked symbols. The database conforms to all the rules proposed for federated 2-DE protein databases and individual protein entries are linked to other protein databases such as SWISS-PROT by active cross-references. This paper describes the construction of HSC-2DPAGE, its maintenance, and access via the Internet.

Animals

PICRUSt2-SC: an update to the reference database used for functional prediction within PICRUSt2.

SUMMARY: PICRUSt2 is a bioinformatic tool that predicts microbial functions in amplicon sequencing data using a database of annotated reference genomes. We have constructed an updated database for PICRUSt2 that has substantially increased the number of bacterial (19,493 to 26,868) and archaeal (406 to 1,002) genomes as well as the number of functional annotations present. The previous PICRUSt2 database relied on many timely and computationally intensive manual processes that made it difficult to update. We constructed a new streamlined process to allow regular upgrades to the PICRUSt2 database on an ongoing basis, and used this process to create a new database, PICRUSt2-SC (Sugar-Coated). Additionally, we have shown that this updated database contains genomes that more closely match study sequences from a range of different environments. The genomes contained in the database therefore better represent these environments and this leads to an improvement in the predicted functional annotations obtained from PICRUSt2. AVAILABILITY AND IMPLEMENTATION: PICRUSt2 source code is freely available at https://github.com/picrust/picrust2 and at https://anaconda.org/bioconda/picrust2. The latest version of PICRUSt2 at the time of writing is also archived: https://doi.org/10.5281/zenodo.15119781. The PICRUSt2-SC database comes pre-installed with PICRUSt2 from version 2.6.0 onwards. Step-by-step instructions for making the updated database are at https://github.com/picrust/picrust2/wiki/Updating-the-PICRUSt2-database. All code used for the analyses and figures in this manuscript is at https://github.com/R-Wright-1/PICRUSt2-SC_application_note and https://doi.org/10.5281/zenodo.15119770.

Software

Databases and software for the analysis of mutations in the human p53 gene, the human hprt gene and the lacZ gene in transgenic rodents.

We have created databases and software applications for the analysis of DNA mutations in the human p53 gene, the human hprt gene and the rodent transgenic lacZ locus. The databases themselves are stand-alone dBase files and the software for analysis of the databases runs on IBM- compatible computers. The software created for these databases permits filtering, ordering, report generation and display of information in the database. In addition, a significant number of routines have been developed for the analysis of single base substitutions. One method of obtaining the databases and software is via the World Wide Web (WWW). Open home page http://sunsite.unc.edu/dnam/mainpage.ht ml with a WWW browser. Alternatively, the databases and programs are available via public ftp from anonymous@sunsite.unc.edu. There is no password required to enter the system. The databases and software are found in subdirectory pub/academic/biology/dna-mutations. Two other programs are available at the WWW site, a program for comparison of mutational spectra and a program for entry of mutational data into a relational database.

Animals

Identifying the active general practice workforce in one division of general practice: the utility of public domain databases.

OBJECTIVE: To identify the non-specialist medical practitioner workforce engaged in active general practice in the region served by the Division of General Practice-Northern Tasmania and to determine the usefulness of public domain databases for enumeration of individual non-specialists providing general practice services. METHODS: A masterlist of the active general practice workforce was compiled by obtaining the names and addresses/postcodes of all non-specialist medical practitioners who were listed in at least one of nine public domain databases and who were confirmed by selected local medical practitioners to be in active general practice in the three months prior to 30 June 1994. This masterlist was used in calculating the sensitivity and positive predictive value (PPV) of each of the nine databases for enumerating non-specialist practitioners in active general practice. RESULTS: Combining the databases resulted in a list of 475 practitioners, which was refined to 139 practitioners who, by our criteria, were in active general practice. Databases had a range of sensitivities and PPVs, but those with high sensitivity tended to have low PPVs, and vice versa. The most useful database for enumerating these practitioners was the mailing list for Australian Family Physician (sensitivity, 94%; PPV, 0.79). CONCLUSIONS: When used alone, no single database had both high sensitivity and high positive predictive value for identifying the active general practice workforce. Combining multiple databases may improve precision. Developing methods to identify recent departures from local active practice has the potential to improve the PPV of existing highly sensitive databases.

Australia

Performance of online biomedical databases in rheumatology.

OBJECTIVE: To compare the performance of MEDLINE, EMBASE, and BIOSIS in selected rheumatology topics. METHODS: Online literature searches were conducted with regard to the epidemiology of rheumatoid arthritis (RA), systemic lupus erythematosus (SLE), and ankylosing spondylitis (AS), as well as for 3 specific questions representing clinical, clinical/laboratory, and therapeutic topics in rheumatology. Total number of citations retrieved, type and language of publication, percentage of contribution from rheumatology journals, and degree of overlap among the databases were recorded. Publications retrieved for the 3 specific questions were also graded for relevance. RESULTS: For 1991, each online biomedical database (OBD) retrieved more than 1,100 citations for RA, over 600 for SLE, and over 110 for AS. For the epidemiology subtopic, fewer than 25% of the citations were retrieved by more than one of the databases. About 3/4 of the citations obtained for the specific search questions were retrieved by a single database. No major differences were observed among databases in relation to number of relevance of citations retrieved. Over 60% of the papers assessed had low relevance in relation to the topic of the search. Efficiency was estimated as the percentage of all relevant citations retrieved by each OBD. Results varied according to the topic, but in most cases each database retrieved at least 50% of the relevant citations. About 45% of the citations retrieved for the 3 search questions were published in nonrheumatology journals. CONCLUSION: No database was superior in all respects. The majority of the citations were retrieved by a single database. A high percentage of the articles retrieved were not relevant, implying low specificity. If a comprehensive online search in rheumatology is required, 2 or more databases should be utilized.

Arthritis, Rheumatoid

Maintaining patient confidentiality in the public domain Internet Autopsy Database (IAD).

The Internet provides the opportunity of permitting public access to large databases containing patient information that can be shared and utilized by epidemiologists, health planners, and medical researchers. Until now, large databases containing patient information have been held in strict confidence, with database access available only to approved researchers or to researchers with access limited to only specific portions of the database. The Internet Autopsy Database (IAD) consists of demographic and pathologic data from over 49,000 autopsies contributed by over a dozen academic medical institutions. Each autopsy record in the public database consists of a uniform set of demographics and SNOMED-compatible terms. To make the database publicly available, a strategy had to be devised that assured the privacy of every person included in the database. A key step involved translating the autopsy facesheets into a listing of SNOMED-compatible terms that effectively eliminated identifying terminology, replacing free text with a generic nomenclature that preserves diagnostic information. The entire database is available on the Internet at: http:@www.med.jhu.edu/pathology/iad.html

Autopsy

Use of large databases for resolving critical care problems.

Large databases allow for rapid access to large volumes of data. To convert raw data to information, large numbers of data points must be correlated into a descriptive pattern that can be interpreted by the user. Databases must be constructed so as to allow reliable extraction of the raw data into a format that supports analysis of events in a meaningful, objective, and reproducible manner. Databases must be responsive to a variety of users. They must not demand unrealistic amounts of effort on those responsible for data entry. Standard protocols in various stages of development will make databases easier to use and more reliable. Database management tools such as the Internet and the National Library of Medicine will become more integrated into the practice of critical care medicine at all levels, including administration, clinical care, and research. This article provides an overview of the capabilities and difficulties associated with large databases. The major areas of use of large databases in the hospital setting are administration, bibliographic, patient care, research, and education. Each of these areas has different requirements and is supported by different types of databases. The advantages and disadvantages of linear, relational, and object-oriented databases are discussed. Issues relating to methods of data entry and the accuracy and reliability of data are discussed. The challenges involving integration of various sources of data and the interfacing of devices are reviewed.

Critical Care

Maintenance of a nutrient database for clinical trials.

Maintenance of a nutrient database for use in dietary analysis for clinical trials and other medical research studies is described. The database, maintained at the University of Minnesota's Nutrition Coordinating Center (NCC), has been used to calculate dietary intake data for a wide range of diet-disease related investigations including studies on cardiovascular disease, hypertension, cancer, gastroenterology, and osteoporosis. Potential sources of error associated with nutrient databases are identified. Criteria are provided for the selection of a nutrient database to meet study objectives and to minimize the potential for errors and inconsistencies. NCC database maintenance procedures, designed to provide updated and verified nutrient calculations for clinical research, involve adherence to standardized procedures for all aspects of database maintenance including data selection, imputations, quality control, recipe calculations, and documentation. By maintaining multiple versions of the database, the NCC is able to update and expand a working version of the database while providing database stability for individual research studies.

Clinical Trials as Topic

An assessment of data quality in the Vermont-Oxford Trials Network database.

The Vermont-Oxford Trials Network is a voluntary collaborative research group of neonatologists that maintains a database for very low birthweight infants (501-1500 g). The database (1) provides core data for randomized trials, (2) serves as a resource for outcomes research in neonatology, and (3) generates quality management reports for participating sites. To assess the reliability of this database and to determine the sources of error, we reviewed 635 medical records chosen at random from among the 4341 eligible infants born at 40 participating data generating sites during an 18-month period beginning January 1, 1990. The estimated frequencies of disagreement between the medical record and database for each of the 10 data items studied and the standard errors of the estimates (in parentheses) were: date of birth 1.3% (0.4), date of admission 2.5% (0.6), date of discharge 8.8% (1.0), birthweight (difference > 50 g) 2.9% (0.6), location of birth (inborn or outborn) 2.1% (0.5), multiple birth 2.2% (0.5), cesarean section 2.5% (0.6), gender 2.1% (0.5), status 28 days after birth 3.4% (0.6), final status 2.9% (0.6). The overall proportions and mean values for items in the database were close to the estimated values based on the random sample of records. There were a total of 247 disagreements between the database and the medical records in the sample. Twenty-three were due to data keying errors. Two hundred twenty-four were due to errors in transcription or interpretation. The rate of data keying errors decreased from over 50 errors per 10,000 fields to less than 15 errors per 10,000 fields when specific quality control procedures, including visual inspection, were instituted. Data keying errors accounted for 13.7% of all disagreements between the database and medical record before improved data entry methods were introduced, and only 3.7% of all errors after they were introduced. We concluded that the Vermont-Oxford Trials Network Database is reliable. Data keying errors have been reduced by the introduction of additional quality control measures. Further reductions in database errors will require measures aimed at minimizing transcription or interpretation errors by individuals completing the data forms.

Computer Communication Networks

Construction of validated, non-redundant composite protein sequence databases.

A strategy has been developed for the construction of a validated, comprehensive composite protein sequence database. Entries are amalgamated from primary source data bases by a largely automated set of processes in which redundant and trivially different entries are eliminated. A modular approach has been adopted to allow scientific judgement to be used at each stage of database processing and amalgamation. Source databases are assigned a priority depending on the quality of sequence validation and commenting. Rejection of entries from the lower priority database, in each pairwise comparison of databases, is carried out according to optionally defined redundancy criteria based on sequence segment mismatches. Efficient algorithms for this methodology are embodied in the COMPO software system. COMPO has been applied for over 2 years in construction and regular updating of the OWL composite protein sequence database from the source databases NBRF-PIR, SWISS-PROT, a GenBank translation retrieved from the feature tables, NBRF-NEW, NEWAT86, PSD-KYOTO and the sequences contained in the Brookhaven protein structure databank. OWL is part of the ISIS integrated data resource of protein sequence and structure [Akrigg et al. (1988) Nature, 335, 745-746]. The modular nature of the integration process greatly facilitates the frequent updating of OWL following releases of the source databases. The extent of redundancy in these sources is revealed by the comparison process. The advantages of a robust composite database for sequence similarity searching and information retrieval are discussed.

Amino Acid Sequence