PubMed HealthSearch

SEARCH · PubMed Health

Results for “Database”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Extensions to the time-oriented database model to support temporal reasoning in medical expert systems.

Physicians faced with diagnostic and therapeutic decisions must reason about clinical features that change over time. Database-management systems (DBMS) can increase access to patient data, but most systems are limited in their ability to store and retrieve complex temporal information. The Time-Oriented Databank (TOD) model, the most widely used data model for medical database systems, associates a single time stamp with each observation. The proper analysis of most clinical data requires accounting for multiple concurrent clinical events that may alter the interpretation of the raw data. Most medical DBMSs cannot retrieve patient data indexed by multiple clinical events. We describe two logical extensions to TOD-based databases that solve a set of temporal reasoning problems we encountered in constructing medical expert systems. A key feature of both extensions is that stored data are partitioned into groupings, such as sequential clinical visits, clinical exacerbations, or other abstract events that have clinical decision-making relevance. The temporal network (TNET) is an object-oriented database that extends the temporal reasoning capabilities of ONCOCIN, a medical expert system that provides chemotherapy advice. TNET uses persistent objects to associate observations with intervals of time during which "an event of clinical interest" occurred. A second object-oriented system called the extended temporal network (ETNET), is both an extension and a simplification of TNET. Like TNET, ETNET uses persistent objects to represent relevant intervals; unlike the first system, however, ETNET contains reasoning methods (rules) that can be executed when an event "begins", and that are withdrawn when that event "concludes". TNET and ETNET capture temporal relationships among recorded information that are not represented in TOD-based databases. Although they do not solve all temporal reasoning problems found in medical decision making, these new structures enable patient database systems to encode complex temporal relationships, to store and retrieve patient data based on multiple clinical contexts and, in ETNET, to modify the reasoning methods available to an expert system based on the onset or conclusion of specific clinical events.

Diagnosis, Computer-Assisted

Microcomputer database management for surgical residents.

Surgical residents must record procedures performed and may choose to keep files of photographic slides, bibliographic references, and a curriculum vitae. Four databases that store this information are produced with an inexpensive and easily obtained microcomputer software program. A surgical procedure database is modeled after the procedure list recommended by surgical boards. This list can be viewed while one enters data, thereby enabling production of accurate and complete records. In the second database, photographic slides are assigned sequence numbers and slide content is designated using both procedure codes and key words, allowing structured and personal recall of data. Data can be printed in many report formats, including that used by the boards of surgery for final submission of reports of residents' operations at the completion of residency. The bibliographic and CV databases contain highly segmented citation data. This structure enables manipulation of data to satisfy the sequence requirements of journals or institutions to which articles or CV are submitted. Database maintenance consumes a few minutes daily and requires a minimum of experience with computers. By providing ease of access to organized data, these databases enhance the potential for critical review of clinical experience by both residents and program directors.

General Surgery

A medical record database in radiology.

A database system on the medical records of radiation therapy, computer tomographic and radioisotopic examinations of our department was created in Computing Center of Hokkaido University which has two remote terminals in the department. Old three filing systems which had been kept in three sections of our department independently since 1972 were integrated by the creation of the database. The main functions of our database management system are as follows; 1. data input through two minicomputers in the department; 2. data loading to the database from the minicomputers; 3. production of key word files and link files for generalised data handling. Seven files are defined in the database with total data of 30 Mega bytes at the end of 1981. Many programs for information retrieval and data processings were prepared and every member of the department can share both data and application programs registered. Outline and operation of the database system and some examples of data processings are reported.

Computers

Computer databases of medical school curricula.

As the pace of curriculum reform in medical education has accelerated during the past decade, so too have demands on curriculum managers to supply increasingly detailed information about the curriculum. In response, a number of schools have joined together to begin work on designs for computer databases of the curriculum. The authors describe three of the most mature curriculum database prototypes, developed by groups at the medical schools of the University of North Carolina at Chapel Hill (UNC), The University of Maryland, and the University of Miami. All three groups have employed relational database management systems to organize information about each "instructional unit" in the preclinical curriculum, including a set of keywords defining the major concepts presented. The keywords are indexed to a controlled vocabulary, either the Medical Subject Headings (MeSH) or a MeSH derivative. The UNC database also employs a textfile management system to provide users with an overview of the entire curriculum. Future work will focus on identifying a suitable controlled vocabulary; capturing content in greater contextual detail; incorporating alternative learning formats, such as problem-based learning; creating links between content items and examination questions; and capturing information generated by student-patient interactions in clinical settings. As a result of recent collaboration with the Association of American Medical Colleges, work to define a prototype national database has begun and a consortium of interested schools is addressing further development activities.

Abstracting and Indexing

Reference Sequence Browser: An R application with a user-friendly GUI to rapidly query sequence databases.

Land managers, researchers, and regulators increasingly utilize environmental DNA (eDNA) techniques to monitor species richness, presence, and absence. In order to properly develop a biological assay for eDNA metabarcoding or quantitative PCR, scientists must be able to find not only reference sequences (previously identified sequences in a genomics database) that match their target taxa but also reference sequences that match non-target taxa. Determining which taxa have publicly available sequences in a time-efficient and accurate manner currently requires computational skills to search, manipulate, and parse multiple unconnected DNA sequence databases. Our team iteratively designed a Graphic User Interface (GUI) Shiny application called the Reference Sequence Browser (RSB) that provides users efficient and intuitive access to multiple genetic databases regardless of computer programming expertise. The application returns the number of publicly accessible barcode markers per organism in the NCBI Nucleotide, BOLD, or CALeDNA CRUX Metabarcoding Reference Databases. Depending on the database, we offer various search filters such as min and max sequence length or country of origin. Users can then download the FASTA/GenBank files from the RSB web tool, view statistics about the data, and explore results to determine details about the availability or absence of reference sequences.

User-Computer Interface

A graphical query generator for clinical research databases.

Clinical research involves recording, storage and retrieval of disease-related patient data, typically using a database system. In order to facilitate ad hoc queries to clinical databases we have developed a query generator with a graphical interface. The query generator uses an object-oriented data model which is visualized by directed graphs. The main focus of our work was the definition of object-oriented user views to the partly complex data structures of a relational database. Furthermore, we tried to define graphical abstractions for all common types of queries. Thus, even for non-expert database users such as clinicians, it is easy to assemble highly complex queries for a thorough examination of the content of large research databases.

Computer Graphics

Persistent gaps and errors in reference databases impede ecologically meaningful taxonomy assignments in 18S rRNA studies: a case study of terrestrial and marine nematodes.

In metabarcoding studies, Linnaean taxonomy assignments of Operational Taxonomic Units (OTUs) or Amplicon Sequence Variants (ASVs) underpin many downstream bioinformatics analyses and ecological interpretations of environmental DNA (eDNA) datasets. However, public molecular databases (i.e., SILVA, EUKARYOME, BOLD) for most microbial metazoan phyla (nematodes, tardigrades, kinorhynchs, etc.) are sparsely populated, negatively impacting our ability to assign ecologically meaningful taxonomy to these understudied groups. Additionally, the choice of bioinformatics parameters and computational algorithms can further impact the accuracy of eDNA taxonomy assignments. Here, we use two in-silico datasets to show that taxonomy assignments using the 18S rRNA gene can be dramatically improved by curating Linnaean taxonomy strings associated with each reference sequence and closing phylogenetic gaps by improving taxon sampling. Using free-living nematodes as a case study, we applied two commonly used taxonomy assignment algorithms (BLAST+ and the QIIME2 Naïve Bayes classifier) across six iterations of the SILVA 138 reference database to evaluate the precision and accuracy of taxonomy assignments. The BLAST+ top hit with a 90% sequence similarity cutoff often returned the highest percentage of correctly assigned taxonomy at the genus level, and the QIIME2 Naïve Bayes classifier performed similarly well when paired with a reference database containing corrected taxonomy strings. Our results highlight the urgent need for phylogenetically-informed expansions of public reference databases (encompassing both genomes and common gene markers), focused on poorly sampled lineages which are now robustly recovered via eDNA metabarcoding approaches. Additional taxonomy curation efforts should be applied to popular reference databases such as SILVA, and taxon sampling could be rapidly improved by more frequent incorporation of newly published GenBank sequences linked to genus and/or species level identifications.

18S rRNA metabarcoding

Microsequences of 145 proteins recorded in the two-dimensional gel protein database of normal human epidermal keratinocytes.

Microsequencing of proteins recovered from two-dimensional (2-D) gels is being used systematically to identify proteins in the master human keratinocyte 2-D gel database. To date, about 250 protein spots recorded in human 2-D gel databases have been microsequenced and, of these, 145 are recorded in the keratinocyte database under the entry partial amino acid sequence. Coomassie Brilliant Blue-stained protein spots cut from several (up to 40) dry gels were concentrated by elution-concentration gel electrophoresis, electroblotted onto PVDF membranes and digested in situ with trypsin. Eluting peptides were separated by reversed-phase HPLC, collected individually and sequenced. Computer search using the FASTA and TFASTA programs from Genetics Computer Group indicated that 110 of the microsequenced polypeptides shared significant similarity with proteins contained in the PIR, Mipsx or GenEMBL databases. Only 35 polypeptides corresponded to hitherto unknown proteins. Peptide sequences of all 145 proteins are listed together with their coordinates (apparent molecular weight and pI) in the keratinocyte database.

Amino Acid Sequence

Workshop on two-dimensional gel protein databases.

A workshop on two-dimensional gel electrophoresis (2-DE) protein database, organized by the Committee on Data for Science and Technology (CODATA) of the International Council of Scientific Unions Task Group on Biological Macromolecules, was held at the CODATA Secretariat in Paris on March 9, 1992. Eleven scientists from eight different countries represented various aspects of 2-DE analysis--namely, cellular protein database development and protein microsequencing methodologies. The purpose of the workshop was to explore means of integrating the rapidly expanding body of information on 2-DE resolved proteins from different laboratories. A major proposal emanating from the workshop was the establishment of an intermediary or "relational" 2-DE gel protein database. This intermediary database, which would catalogue pertinent information on 2-DE resolved proteins (experimental source, 2-DE loci, biological information, etc.) could be an adjunct to, and accessed through, the existing international protein sequence databanks. It would function as a pointer for researchers to the individual 2-DE protein databases where primary and more specialized 2-DE data would be housed.

Databases, Factual

The human keratinocyte two-dimensional protein database (update 1994): towards an integrated approach to the study of cell proliferation, differentiation and skin diseases.

The master two-dimensional (2-D) gel database of human keratinocytes currently lists 3087 cellular proteins (2168 isoelectric focusing, IEF; and 919 none-quilibrium pH gradient electrophoresis, NEPHGE), many of which correspond to posttranslational modifications, 890 polypeptides have been identified (protein name, organelle components, etc.) using one or a combination of procedures that include (i) comigration with known human proteins, (ii) 2-D gel immunoblotting using specific antibodies (iii) microsequencing of Coomassie Brilliant Blue stained proteins, (iv) mass spectrometry and (v) vaccinia virus expression of full length cDNAs. These are listed both in alphabetical order and with increasing SSP number, together with their M(r), pI, cellular localization and credit to the investigator(s) that aided in the identification. Furthermore, we list 239 microsequenced proteins recorded in the database. We also report a database of proteins recovered from the medium of noncultured, unfractionated keratinocytes. This database lists 398 polypeptides (309 IEF; 89 NEPHGE) of which 76 have been identified. The aim of the comprehensive databases is to gather, through a systematic study of keratinocytes, qualitative and quantitative information on proteins and their genes that may allow us to identify abnormal patterns of gene expression and, ultimately, to pinpoint signaling pathways and components affected in various skin diseases, cancer included.

Amino Acid Sequence

A deductive database system for analyzing human nucleotide sequence data.

The analysis of the human genome is one of the most significant topics in both biology and medical science. There is a growing need for a well-designed database system for searching and analyzing the human genome data. We developed a deductive database system to search and analyze nucleotide sequence data derived from the GenBank primates data. A deductive database system is a next generation one and it contains an inference mechanism that can handle problems beyond the capabilities of classical database systems. Database queries are described in logical rules. These rules are simple even for molecular biologists who are not experts in computer programs because they are declarative and do not require the procedural commands that are usually used in computer programs. Furthermore, queries based on logical rules are powerful enough to express complicated biological problems. Particularly, recursive rules are suitable for examining secondary structures of nucleotide sequences. In our analysis of TfR's IRE, we noted five stem-and-loop structures.

Artificial Intelligence

Statistical analysis in dBASE-compatible databases.

Database management in clinical and experimental research often requires statistical analysis of the data in addition to the usual functions for storing, organizing, manipulating and reporting. With most database systems, transfer of data to a dedicated statistics package is a relatively simple task. However, many statistics programs lack the powerful features found in database management software. dBASE IV and compatible programs are currently among the most widely used database management programs. d4STAT is a utility program for dBASE, containing a collection of statistical functions and tests for data stored in the dBASE file format. By using d4STAT, statistical calculations may be performed directly on the data stored in the database without having to exit dBASE IV or export data. Record selection and variable transformations are performed in memory, thus obviating the need for creating new variables or data files. The current version of the program contains routines for descriptive statistics, paired and unpaired t-tests, correlation, linear regression, frequency tables, Mann-Whitney U-test, Wilcoxon signed rank test, a time-saving procedure for counting observations according to user specified selection criteria, survival analysis (product limit estimate analysis, log-rank test, and graphics), and normal t and chi-squared distribution functions.

Database Management Systems

Redesigning, implementing and integrating Escherichia coli genome software tools with an object-oriented database system.

This paper reports our exploratory work to redesign, implement and integrate a collection of genome software tools with an object-oriented database system. Our software tools deal with genome data from Escherichia coli K-12, a bacterium that has been studied intensively and provides richer data sets than any other living organism. The object-oriented DBMS used for the integration is ONTOS, a commercial object-oriented system from Ontologic Inc. This redesign and implementation task was performed in two steps. First, C programs were converted into C++, and then the C++ version programs were modified and integrated with an object-oriented modeling of the data to form an ONTOS database application. The first step helps us develop a conceptual view for a DBMS-independent object-oriented construct. The second step elucidates what additional DBMS-dependent modification steps are needed to provide persistency to the objects. Examples are included to illustrate steps of the redesign and implementation. Overall, the outcome of this project demonstrates that programs and data can be successfully integrated with an object-oriented database, while providing the objects with persistency and shareability. This paper includes discussions using concrete examples on what advantage the object-oriented database approach provides over the relational database approach.

Base Sequence

PATMAT: a searching and extraction program for sequence, pattern and block queries and databases.

A program has been developed that provides molecular biologists with multiple tools for searching databases, yet uses a very simple interface. PATMAT can use protein or (translated) DNA sequences, patterns or blocks of aligned proteins as queries of databases consisting of amino acid or nucleotide sequences, patterns or blocks. The ability to search databases of blocks by 'on-the-fly' conversion to scoring matrices provides a new tool for detection and evaluation of distant relationships. PATMAT uses a pull-down, menu-driven interface to carry out its multiple searching, extraction and viewing functions. Each query or database type is recognized, reported, and the appropriate search carried out, with matches and alignments reported in windows as they occur. Any of the high scoring matches can be exported to a file, viewed and recalled as a query using only a few keystrokes or mouse selections. Searches of multiple database files are carried out by user selection within a window. PATMAT runs under DOS; the searching engine also runs under UNIX.

Amino Acid Sequence

Corruption of genomic databases with anomalous sequence.

We describe evidence that DNA sequences from vectors used for cloning and sequencing have been incorporated accidentally into eukaryotic entries in the GenBank database. These incorporations were not restricted to one type of vector or to a single mechanism. Many minor instances may have been the result of simple editing errors, but some entries contained large blocks of vector sequence that had been incorporated by contamination or other accidents during cloning. Some cases involved unusual rearrangements and areas of vector distant from the normal insertion sites. Matches to vector were found in 0.23% of 20,000 sequences analyzed in GenBank Release 63. Although the possibility of anomalous sequence incorporation has been recognized since the inception of GenBank and should be easy to avoid, recent evidence suggests that this problem is increasing more quickly than the database itself. The presence of anomalous sequence may have serious consequences for the interpretation and use of database entries, and will have an impact on issues of database management. The incorporated vector fragments described here may also be useful for a crude estimate of the fidelity of sequence information in the database. In alignments with well-defined ends, the matching sequences showed 96.8% identity to vector; when poorer matches with arbitrary limits were included, the aggregate identity to vector sequence was 94.8%.

Base Sequence

The European Bioinformatics Institute (EBI) databases.

This paper describes the databases and services of the European Bioinformatics Institute (EBI). In collaboration with DDBJ and GenBank/NCBI, the EBI maintains and distributes the EMBL Nucleotide Sequence Database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence Database, in collaboration with Amos Bairoch of the University of Geneva. Over thirty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists, are also available. The EBI network services include database searching, entry retrieval, and sequence similarity searching facilities.

Amino Acid Sequence

Comparison of systematic search and database methods for constructing segments of protein structure.

Two principal methods of determining the conformation of short pieces of polypeptide backbone in proteins have been developed: using a database of known structures and systematically generating all conformations. In this paper, we compare the effectiveness of these two techniques. The completeness of the database for segments of different lengths is examined and it is found to contain most conformations for segments seven residues long, but to deteriorate rapidly for longer regions. When the database segment is to be incorporated into the rest of a structure, at least seven residues are required to build four new residues, because of the need to position the segment relative to the rest of the structure. It is found that such positioning using flanking residues results in large errors in the inserted region. We conclude that the database method is currently not effective for comparative modeling, even for short segments. The systematic search procedure is found to generate almost all structures of short segments found in proteins. In contrast to the database method, low root mean square error structures are obtained for a set of trial segments embedded in the rest of a protein structure. Thus, it should be considered the method of choice.

Computer Simulation

SubtiList: a relational database for the Bacillus subtilis genome.

In the framework of the international collaborative project aiming to sequence the whole Bacillus subtilis chromosome, we have created a relational database for managing and analysing information associated with the molecular genetics of this bacterium: SubtiList. It allows recovery of non-redundant DNA sequences of the B. subtilis genome, as well as related information, i.e. genes, proteins, etc. A logical structure has been designed with appropriate links between the different objects, and a set of procedures has been implemented for data updating and management. The database is organized around a core constituted by all known contigs of B. subtilis, i.e. sets of non-redundant sequences created from original entries in the EMBL data library. A user-friendly interface has been developed to make the database easy to consult. Sequence analysis tools have been integrated into the database, such as a program for rapid similarity searching of protein data banks, and a powerful DNA pattern searching program. Thanks to the consistency of SubtiList, we have performed a codon usage analysis by Factorial Correspondence Analysis, and a study of the distribution of the isoelectric points of known proteins of B. subtilis. The SubtiList database is available through anonymous ftp (address 'ftp.pasteur.fr' or IP number 157.99.64.12, directory '/pub/GenomeDB/SubtiList').

Algorithms