PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

ProtBuD: a database of biological unit structures of protein families and superfamilies.

MOTIVATION: Modeling of protein interactions is often possible from known structures of related complexes. It is often time-consuming to find the most appropriate template. Hypothesized biological units (BUs) often differ from the asymmetric units and it is usually preferable to model from the BUs. RESULTS: ProtBuD is a database of BUs for all structures in the Protein Data Bank (PDB). We use both the PDBs BUs and those from the Protein Quaternary Server. ProtBuD is searchable by PDB entry, the Structural Classification of Proteins (SCOP) designation or pairs of SCOP designations. The database provides the asymmetric and BU contents of related proteins in the PDB as identified in SCOP and Position-Specific Iterated BLAST (PSI-BLAST). The asymmetric unit is different from PDB and/or Protein Quaternary Server (PQS) BUs for 52% of X-ray structures, and the PDB and PQS BUs disagree on 18% of entries. AVAILABILITY: The database is provided as a standalone program and a web server from http://dunbrack.fccc.edu/ProtBuD.php.

Amino Acid Sequence↗

The Protein Data Bank and lessons in data management.

The Protein Data Bank (PDB) is a widely used biological database of macromolecular structures with a long history. This history is treated as lessons learned and is used to highlight what are believed to be the best practices important to developers of biological databases today. While the focus is on data quality, data representation and the information technology to support these data, the non-data and technology issues cannot be ignored. The role of the human factor in the form of users, collaborators, scientific society and ad hoc committees is also included.

Database Management Systems↗

Evaluation of the Biolog MicroStation system for yeast identification.

One hundred and fifty-nine isolates representing 16 genera and 53 species of yeasts were processed with the Biolog MicroStation System for yeast identification. Thirteen genera and 38 species were included in the Biolog database. For these 129 isolates, correct identifications to the species level were 13.2, 39.5 and 48.8% after 24, 48 and 72 hours incubation at 30 degrees C, respectively. Three genera and 15 species which were not included in the Biolog database were also tested. Of the 30 isolates studied, 16.7, 53.3 and 56.7% of the isolates were given incorrect names from the system's database after 24,48 and 72 h incubation at 30 degrees C, respectively. The remaining isolates of this group were not identified.

Candida↗

Assessment of variability in biomonitoring data using a large database of biological measures of exposure.

Although intra- and interindividual sources of variation in airborne exposures have been extensively studied, similar investigations examining variability in biological measures of exposure have been limited. Following a review of the world's published literature, biological monitoring data were abstracted from 53 studies that examined workers' exposures to metals, solvents, polycyclic aromatic hydrocarbons, and pesticides. Approximately 40% of the studies also reported personal sampling results, which were compiled as well. In this study, the authors evaluated the intra- and interindividual sources of variation in biological measures of exposure collected on workers employed at the same plant. In 60% of the data sets, there was more variation among workers than variation from day to day. Approximately one-fourth of the data were homogeneous with small differences among workers' mean exposure levels. However, an almost equal number of data sets exhibited moderate to extreme levels of heterogeneity in exposures among workers at the same facility. In addition, the relative magnitude of the intra- to interindividual source of variation was larger for biomarkers with short compared to long half-lives, which suggests that biomarkers with half-lives of 7 days or longer exhibit physiologic dampening of fluctuations in external levels of the workplace contaminant and thereby may offer advantages when compared to short-lived biomarkers or exposures assessed by air monitoring. The use of biological indices of exposure, however, places an additional burden on the strategy used to evaluate exposures, because data may be serially correlated as evidenced in this study, which could result in biased estimates of the variance components if autocorrelation is undetected or ignored in the statistical analyses.

Air Pollutants, Occupational↗

Evaluation of text data mining for database curation: lessons learned from the KDD Challenge Cup.

MOTIVATION: The biological literature is a major repository of knowledge. Many biological databases draw much of their content from a careful curation of this literature. However, as the volume of literature increases, the burden of curation increases. Text mining may provide useful tools to assist in the curation process. To date, the lack of standards has made it impossible to determine whether text mining techniques are sufficiently mature to be useful. RESULTS: We report on a Challenge Evaluation task that we created for the Knowledge Discovery and Data Mining (KDD) Challenge Cup. We provided a training corpus of 862 articles consisting of journal articles curated in FlyBase, along with the associated lists of genes and gene products, as well as the relevant data fields from FlyBase. For the test, we provided a corpus of 213 new ('blind') articles; the 18 participating groups provided systems that flagged articles for curation, based on whether the article contained experimental evidence for gene expression products. We report on the evaluation results and describe the techniques used by the top performing groups.

Abstracting and Indexing↗

Rutabaga by any other name: extracting biological names.

As the pace of biological research accelerates, biologists are becoming increasingly reliant on computers to manage the information explosion. Biologists communicate their research findings by relying on precise biological terms; these terms then provide indices into the literature and across the growing number of biological databases. This article examines emerging techniques to access biological resources through extraction of entity names and relations among them. Information extraction has been an active area of research in natural language processing and there are promising results for information extraction applied to news stories, e.g., balanced precision and recall in the 93-95% range for identifying person, organization and location names. But these results do not seem to transfer directly to biological names, where results remain in the 75-80% range. Multiple factors may be involved, including absence of shared training and test sets for rigorous measures of progress, lack of annotated training data specific to biological tasks, pervasive ambiguity of terms, frequent introduction of new terms, and a mismatch between evaluation tasks as defined for news and real biological problems. We present evidence from a simple lexical matching exercise that illustrates some specific problems encountered when identifying biological names. We conclude by outlining a research agenda to raise performance of named entity tagging to a level where it can be used to perform tasks of biological importance.

Abstracting and Indexing↗

Improved sensitivity of biological sequence database searches.

We have increased the sensitivity of DNA and protein sequence database searches by allowing similar but non-identical amino acids or nucleotides to match. In addition, one can match k-tuples or words instead of matching individual residues in order to speed the search. A matching matrix species which k-tuples match each other. The matching matrix can be calculated from a similarity matrix of amino acids and a threshold of similarity required for matching. This permits amino acid similarity matrices or replacement matrices (PAM matrices) to be used in the first step of a sequence comparison rather than in a secondary scoring phase. The concept of matching non-identical k-tuples also increases the power of DNA database searches. For example, a matrix that specifies that any 3-tuple in a DNA sequence can match any other 3-tuple encoding the same amino acid permits a DNA database search using a DNA query sequence for regions that would encode a similar amino acid sequence.

Amino Acid Sequence↗