PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Development of the Overlapping Oligonucleotide Database and its application to signal sequence search of the human genome.

We have developed ODS (Overlapping Oligonucleotide Database for Signal Sequence Search)--the first relational database that integrates information on biological features into the search for signal sequences. In existing biological sequence databases, even relational ones, retrieving nucleotide sequences based on their biological features involves much labour and time or even the development of a new program. GenBank sequence data, including FEATURES records, are organized into three relational tables in ODS. Nucleotide sequences are transformed into overlapping oligonucleotides in order to facilitate the signal sequence search rapidly without the need for specific alignment programs. This transformation leads to a one-to-one correspondence between the nucleotide sequence and its biological feature. The signal sequence search by ODS is done in SQL queries and ODS obviates the need for molecular biologists to write computer programs. The application of ODS to searches of promoter regions revealed putative cis-acting elements and basic statistical analyses of occurrences of oligonucleotides showed interesting findings concerning the 'cg' dinucleotide.

Computer Systems↗

Oryzabase. An integrated biological and genome information database for rice.

The aim of Oryzabase is to create a comprehensive view of rice (Oryza sativa) as a model monocot plant by integrating biological data with molecular genomic information (http://www.shigen.nig.ac.jp/rice/oryzabase/top/top.jsp). The database contains information about rice development and anatomy, rice mutants, and genetic resources, especially for wild varieties of rice. The anatomical description of rice development is unique and is the first known representation for rice. Developmental and anatomical descriptions include in situ gene expression data serving as stage and tissue markers. The systematic presentation of a large number of rice mutant and mutant trait genes is indispensable, as is description of research in wild strains, core collections, and their detailed characterization. Several genetic, physical, and expression maps with full genome and cDNA sequences are also combined with biological data in Oryzabase. These datasets, when pooled together, could provide a useful tool for gaining greater knowledge about the life cycle of rice, the relationship between phenotype and gene function, and rice genetic diversity. For exchanging community information, Oryzabase publishes the Rice Genetics Newsletter organized by the Rice Genetics Cooperative and provides a mailing service, rice-e-net/rice-net.

Chromosome Mapping↗

Large scale hierarchical clustering of protein sequences.

BACKGROUND: Searching a biological sequence database with a query sequence looking for homologues has become a routine operation in computational biology. In spite of the high degree of sophistication of currently available search routines it is still virtually impossible to identify quickly and clearly a group of sequences that a given query sequence belongs to. RESULTS: We report on our developments in grouping all known protein sequences hierarchically into superfamily and family clusters. Our graph-based algorithms take into account the topology of the sequence space induced by the data itself to construct a biologically meaningful partitioning. We have applied our clustering procedures to a non-redundant set of about 1,000,000 sequences resulting in a hierarchical clustering which is being made available for querying and browsing at http://systers.molgen.mpg.de/. CONCLUSIONS: Comparisons with other widely used clustering methods on various data sets show the abilities and strengths of our clustering methods in producing a biologically meaningful grouping of protein sequences.

Algorithms↗

Databases and software for the comparison of prokaryotic genomes.

The explosion in the number of complete genomes over the past decade has spawned a new and exciting discipline, that of comparative genomics. To exploit the full potential of this approach requires the development of novel algorithms, databases and software which are sophisticated enough to draw meaningful comparisons between complete genome sequences and are widely accessible to the scientific community at large. This article reviews progress towards the development of computational tools and databases for organizing and extracting biological meaning from the comparison of large collections of genomes.

Computational Biology↗

Congruence of tissue expression profiles from Gene Expression Atlas, SAGEmap and TissueInfo databases.

BACKGROUND: Extracting biological knowledge from large amounts of gene expression information deposited in public databases is a major challenge of the postgenomic era. Additional insights may be derived by data integration and cross-platform comparisons of expression profiles. However, database meta-analysis is complicated by differences in experimental technologies, data post-processing, database formats, and inconsistent gene and sample annotation. RESULTS: We have analysed expression profiles from three public databases: Gene Expression Atlas, SAGEmap and TissueInfo. These are repositories of oligonucleotide microarray, Serial Analysis of Gene Expression and Expressed Sequence Tag human gene expression data respectively. We devised a method, Preferential Expression Measure, to identify genes that are significantly over- or under-expressed in any given tissue. We examined intra- and inter-database consistency of Preferential Expression Measures. There was good correlation between replicate experiments of oligonucleotide microarray data, but there was less coherence in expression profiles as measured by Serial Analysis of Gene Expression and Expressed Sequence Tag counts. We investigated inter-database correlations for six tissue categories, for which data were present in the three databases. Significant positive correlations were found for brain, prostate and vascular endothelium but not for ovary, kidney, and pancreas. CONCLUSION: We show that data from Gene Expression Atlas, SAGEmap and TissueInfo can be integrated using the UniGene gene index, and that expression profiles correlate relatively well when large numbers of tags are available or when tissue cellular composition is simple. Finally, in the case of brain, we demonstrate that when PEM values show good correlation, predictions of tissue-specific expression based on integrated data are very accurate.

Brain↗

Assessment of the impact of manual curation in BioCyc.

INTRODUCTION: BioCyc is an extensive collection of databases of genomic and pathway information for microorganisms and model eukaryotes. These organismal databases integrate diverse biological data by combining computationally inferred information, data imported from other databases, and, for selected organisms, literature-based manual curation. This study investigates the magnitude and significance of annotation changes performed during the curation of 10 prokaryotic genomes to better understand the rate of erroneous annotations and the value of BioCyc curation. METHODS: We identified curation changes by finding cases where the annotation of the protein at the start of the curation process differed from its annotation at the end of the process. RESULTS: We found that across a sample of curated databases (n = 10), the annotation of 6,753, or 25.6% of the proteins in the pooled protein dataset (n = 26,126) were modified. Assessment of considerable sampling fractions of these proteins found that a median of 62% (mean of 52.9%) represented functionally informative name changes, rather than stylistic annotation changes. These results were then extrapolated to total proteins with name changes with uncertainty quantified via finite population correction, indicating that most Tier 2 Biocyc PGDBs received hundreds of functionally informative name changes during manual curation. On average 363, or13% (±5.4% SD) of the proteins encoded in each genome received functionally informative annotation changes, ranging from 5.3% (Streptococcus pneumoniae D39V) to 22.7% (Staphylococcus aureus NCTC 8325). DISCUSSION: These findings demonstrate a substantial improvement in the accuracy of manually curated BioCyc databases compared with automated annotation pipelines. This result is particularly impactful as the rate of downstream propagation of erroneous annotations across biological databases can significantly compromise scientific discovery.

annotation errors↗

Biological sequences integrated: a relational database approach.

Over the last decade the modeling and the storage of biological data has been a topic of wide interest for scientists dealing with biological and biomedical research. Currently most data is still stored in text files which leads to data redundancies and file chaos. In this paper we show how to use relational modeling techniques and relational database technology for modeling and storing biological sequence data, i.e. for data maintained in collections like EMBL or SWISS-PROT to better serve the needs for these application domains. For this reason we propose a two step approach. First, we model the structure (and therefore the meaning of the) data using an Entity-Relationship approach. The ER model leads to a clean design of a relational database schema for storing and retrieving the DNA and protein data extracted from various sources. Our approach provides the clean basis for building complex biological applications that are more amenable to changes and software ports than their file-base counterparts.

Computational Biology↗

Development of human protein reference database as an initial platform for approaching systems biology in humans.

Human Protein Reference Database (HPRD) is an object database that integrates a wealth of information relevant to the function of human proteins in health and disease. Data pertaining to thousands of protein-protein interactions, posttranslational modifications, enzyme/substrate relationships, disease associations, tissue expression, and subcellular localization were extracted from the literature for a nonredundant set of 2750 human proteins. Almost all the information was obtained manually by biologists who read and interpreted >300,000 published articles during the annotation process. This database, which has an intuitive query interface allowing easy access to all the features of proteins, was built by using open source technologies and will be freely available at http://www.hprd.org to the academic community. This unified bioinformatics platform will be useful in cataloging and mining the large number of proteomic interactions and alterations that will be discovered in the postgenomic era.

BRCA1 Protein↗

PANTHER: a browsable database of gene products organized by biological function, using curated protein family and subfamily classification.

The PANTHER database was designed for high-throughput analysis of protein sequences. One of the key features is a simplified ontology of protein function, which allows browsing of the database by biological functions. Biologist curators have associated the ontology terms with groups of protein sequences rather than individual sequences. Statistical models (Hidden Markov Models, or HMMs) are built from each of these groups. The advantage of this approach is that new sequences can be automatically classified as they become available. To ensure accurate functional classification, HMMs are constructed not only for families, but also for functionally distinct subfamilies. Multiple sequence alignments and phylogenetic trees, including curator-assigned information, are available for each family. The current version of the PANTHER database includes training sequences from all organisms in the GenBank non-redundant protein database, and the HMMs have been used to classify gene products across the entire genomes of human, and Drosophila melanogaster. The ontology terms and protein families and subfamilies, as well as Drosophila gene c;assifications, can be browsed and searched for free. Due to outstanding contractual obligations, access to human gene classifications and to protein family trees and multiple sequence alignments will temporarily require a nominal registration fee. PANTHER is publicly available on the web at http://panther.celera.com.

Animals↗

Proteomic 2DE database for spot selection, automated annotation, and data analysis.

We present a software solution that enables faster and more accurate data analysis of 2DE/MALDI TOF MS data. The software supports data analysis through a number of automated data selection functions and advanced graphical tools. Once protein identities are determined using MALDI TOF MS, automated data retrieval from online databases provides biological information. The software, called 2DDB, reduces analysis time to a fraction without losing any quality compared to more manual data analysis. The database contains over 100,000 data entries, and selected parts can be reached at http://2ddb.org.

Animals↗

Relational databases: a transparent framework for encouraging biology students to think informatically.

We discuss how relational databases constitute an ideal framework for representing and analyzing large-scale genomic data sets in biology. As a case study, we describe a Drosophila splice-site database that we recently developed at Wesleyan University for use in research and teaching. The database stores data about splice sites computed by a custom algorithm using Drosophila cDNA transcripts and genomic DNA and supports a set of procedures for analyzing splice-site sequence space. A generic Web interface permits the execution of the procedures with a variety of parameter settings and also supports custom structured query language queries. Moreover, new analytical procedures can be added by updating special metatables in the database without altering the Web interface. The database provides a powerful setting for students to develop informatic thinking skills.

Algorithms↗

Fibrinogen depleting agents for acute ischaemic stroke.

BACKGROUND: Fibrinogen depleting agents reduce fibrinogen in blood plasma, reduce blood viscosity and hence increase blood flow. This may help remove the blood clot blocking the artery and re-establish blood flow to the affected area of the brain after an ischaemic stroke. The risk of haemorrhage may be less than with thrombolytic agents. OBJECTIVES: The objective of this review was to assess the effect of fibrinogen depleting agents in people with acute ischaemic stroke. SEARCH STRATEGY: We searched the Cochrane Stroke Group trials register, Embase (to January 1996) and the Index of scientific and technical proceedings (1982 to 1996). We handsearched 10 Chinese journals and the proceedings of the 4th Chinese Stroke Conference (1995). We contacted Chinese and Japanese researchers, and drug companies. SELECTION CRITERIA: Randomised and quasi-randomised trials of fibrinogen depleting agents started within 14 days of stroke onset, compared with control in patients with definite or possible ischaemic stroke. DATA COLLECTION AND ANALYSIS: Two reviewers independently applied the inclusion criteria, assessed trial quality and extracted the data. MAIN RESULTS: Three trials involving 182 patients were included. All trials tested ancrod. Allocation concealment was adequate in two trials. Fibrinolytic therapy appeared to reduce the number of deaths during the scheduled treatment period (odds ratio 0.33, 95% confidence interval 0.13 to 0.85). There was no difference in death from all causes at the end of follow-up (odds ratio 0.57, 95% confidence interval 0.27 to 1.23). The risk of haemorrhage appeared to be small, but there were too few data to be conclusive. One trial in 132 people showed a statistically non-significant decrease in the odds of being dead or disabled at the end of follow-up (odds ratio 0.52, 95% confidence interval 0.26 to 1.03). REVIEWER'S CONCLUSIONS: Although ancrod appears to be promising, it is not possible to draw reliable conclusions from the available data.

Ancrod↗