PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A protocol for maintaining multidatabase referential integrity.

The bioinformatics community is becoming increasingly reliant on the creation of links among biological databases (DBs) as a foundation for DB interoperability. For example, a link might be created from a protein in one DB (such as PIR), to a gene in another DB (such as GDB), by storing the unique identifier (id) of the gene object within an attribute of the protein object. User interfaces can then support navigation from the protein to the gene, and multiDB queries can join the protein with the gene. The unique id of the gene is serving as a foreign key. However, a variety of factors, such as changes in the underlying biology, can cause object ids to become invalid, thus producing invalid links among DBs. Invalid links are a violation of multidatabase referential integrity. We propose a network protocol whereby a database administrator can provide information about changes to the identifiers of objects in their database via Internet, to allow other databases to maintain referential integrity. We request comments from the bioinformatics community for the purpose of building a consensus on the proposed protocol.

Computational Biology↗

Development of a computer system in search of antifertility drug from indigenous plants.

Recent advancement in computer technology has increased the application of database in identification and verification systems used in research of indigenous drugs. To overcome some of the drawbacks of existing knowledge (systems) we have developed a computer version of a key for interfacing with pre-existing knowledgebase. Appearance of different types of biological databases on the World Wide Web (WWW) and hypertext links between them has made a large amount of information easily accessible to biologists. Storage of explicitly specified biological relationship between different entities as discrete entries can allow additional identification capabilities. These systems are designed to detect the identity of a natural product when it is unknown or to verify the product identity when an antifertility products is provided. We have built a database about the cytoskeleton that explores the natural product approach as antifertility drug and the gap in existing knowledge. The stored information is displayed along with other retrieved information. The system is capable to search for entities with specified properties. It is extensible so that new types of relationships may be incorporated.

Contraceptive Agents↗

Novel techniques for visualising biological information.

The major challenge facing the bioinformatics community is the continuing increase in the number, size and complexity of biological databases with which it must contend. The goal of the research discussed herein is the development and utilisation of techniques that allow researchers to extract new and useful information from these burgeoning information resources using advanced visualisation methods and paradigms, coupled with distributed object technologies that allow communications between applications and remote databases. Visualisation has roles not only in analysis, but also in building more user-friendly interfaces, implementing methods to navigate large information spaces intuitively and powerful techniques to browse and query data. By using platform-independent object-oriented programming languages, these resources may be developed as reusable pieces of software componentry with their methods and interfaces defined fully, and then distributed through organisations such as the bioWidget Consortium. The widget and object-oriented approach is a powerful paradigm in developing new applications from existing components. Development time is reduced and greater time is spent on analysing these data, rather than in the writing of monolithic applications. More powerful applications can be constructed from components interacting in concert and offers the opportunity of a new generation of bioinformatics tools.

Base Composition↗

The tissue microarray data exchange specification: implementation by the Cooperative Prostate Cancer Tissue Resource.

BACKGROUND: Tissue Microarrays (TMAs) have emerged as a powerful tool for examining the distribution of marker molecules in hundreds of different tissues displayed on a single slide. TMAs have been used successfully to validate candidate molecules discovered in gene array experiments. Like gene expression studies, TMA experiments are data intensive, requiring substantial information to interpret, replicate or validate. Recently, an open access Tissue Microarray Data Exchange Specification has been released that allows TMA data to be organized in a self-describing XML document annotated with well-defined common data elements. While this specification provides sufficient information for the reproduction of the experiment by outside research groups, its initial description did not contain instructions or examples of actual implementations, and no implementation studies have been published. The purpose of this paper is to demonstrate how the TMA Data Exchange Specification is implemented in a prostate cancer TMA. RESULTS: The Cooperative Prostate Cancer Tissue Resource (CPCTR) is funded by the National Cancer Institute to provide researchers with samples of prostate cancer annotated with demographic and clinical data. The CPCTR now offers prostate cancer TMAs and has implemented a TMA database conforming to the new open access Tissue Microarray Data Exchange Specification. The bulk of the TMA database consists of clinical and demographic data elements for 299 patient samples. These data elements were extracted from an Excel database using a transformative Perl script. The Perl script and the TMA database are open access documents distributed with this manuscript. CONCLUSIONS: TMA databases conforming to the Tissue Microarray Data Exchange Specification can be merged with other TMA files, expanded through the addition of data elements, or linked to data contained in external biological databases. This article describes an open access implementation of the TMA Data Exchange Specification and provides detailed guidance to researchers who wish to use the Specification.

Confidentiality↗

Automated discovery of structural signatures of protein fold and function.

There are constraints on a protein sequence/structure for it to adopt a particular fold. These constraints could be either a local signature involving particular sequences or arrangements of secondary structure or a global signature involving features along the entire chain. To search systematically for protein fold signatures, we have explored the use of Inductive Logic Programming (ILP). ILP is a machine learning technique which derives rules from observation and encoded principles. The derived rules are readily interpreted in terms of concepts used by experts. For 20 populated folds in SCOP, 59 rules were found automatically. The accuracy of these rules, which is defined as the number of true positive plus true negative over the total number of examples, is 74% (cross-validated value). Further analysis was carried out for 23 signatures covering 30% or more positive examples of a particular fold. The work showed that signatures of protein folds exist, about half of rules discovered automatically coincide with the level of fold in the SCOP classification. Other signatures correspond to homologous family and may be the consequence of a functional requirement. Examination of the rules shows that many correspond to established principles published in specific literature. However, in general, the list of signatures is not part of standard biological databases of protein patterns. We find that the length of the loops makes an important contribution to the signatures, suggesting that this is an important determinant of the identity of protein folds. With the expansion in the number of determined protein structures, stimulated by structural genomics initiatives, there will be an increased need for automated methods to extract principles of protein folding from coordinates.

Algorithms↗

Statistical Viewer: a tool to upload and integrate linkage and association data as plots displayed within the Ensembl genome browser.

BACKGROUND: To facilitate efficient selection and the prioritization of candidate complex disease susceptibility genes for association analysis, increasingly comprehensive annotation tools are essential to integrate, visualize and analyze vast quantities of disparate data generated by genomic screens, public human genome sequence annotation and ancillary biological databases. We have developed a plug-in package for Ensembl called "Statistical Viewer" that facilitates the analysis of genomic features and annotation in the regions of interest defined by linkage analysis. RESULTS: Statistical Viewer is an add-on package to the open-source Ensembl Genome Browser and Annotation System that displays disease study-specific linkage and/or association data as 2 dimensional plots in new panels in the context of Ensembl's Contig View and Cyto View pages. An enhanced upload server facilitates the upload of statistical data, as well as additional feature annotation to be displayed in DAS tracts, in the form of Excel Files. The Statistical View panel, drawn directly under the ideogram, illustrates lod score values for markers from a study of interest that are plotted against their position in base pairs. A module called "Get Map" easily converts the genetic locations of markers to genomic coordinates. The graph is placed under the corresponding ideogram features a synchronized vertical sliding selection box that is seamlessly integrated into Ensembl's Contig- and Cyto- View pages to choose the region to be displayed in Ensembl's "Overview" and "Detailed View" panels. To resolve Association and Fine mapping data plots, a "Detailed Statistic View" plot corresponding to the "Detailed View" may be displayed underneath. CONCLUSION: Features mapping to regions of linkage are accentuated when Statistic View is used in conjunction with the Distributed Annotation System (DAS) to display supplemental laboratory information such as differentially expressed disease genes in private data tracks. Statistic View is a novel and powerful visual feature that enhances Ensembl's utility as valuable resource for integrative genomic-based approaches to the identification of candidate disease susceptibility genes. At present there are no other tools that provide for the visualization of 2-dimensional plots of quantitative data scores against genomic coordinates in the context of a primary public genome annotation browser.

Chromosome Mapping↗

Visualizations for taxonomic and phylogenetic trees.

MOTIVATION: Despite substantial efforts to develop and populate the back-ends of biological databases, front-ends to these systems often rely on taxonomic expertise. This research applies techniques from human-computer interaction research to the biodiversity domain. RESULTS: We developed an interactive node-link tool, TaxonTree, illustrating the value of a carefully designed interaction model, animation, and integrated searching and browsing towards retrieval of biological names and other information. Users tested the tool using a new, large integrated dataset of animal names with phylogenetic-based and classification-based tree structures. These techniques also translated well for a tool, DoubleTree, to allow comparison of trees using coupled interaction. Our approaches will be useful not only for biological data but as general portal interfaces.

Algorithms↗

Extension and integration of the gene ontology (GO): combining GO vocabularies with external vocabularies.

Structured vocabulary development enhances the management of information in biological databases. As information grows, handling the complexity of vocabularies becomes difficult. Defined methods are needed to manipulate, expand and integrate complex vocabularies. The Gene Ontology (GO) project provides the scientific community with a set of structured vocabularies to describe domains of molecular biology. The vocabularies are used for annotation of gene products and for computational annotation of sequence data sets. The vocabularies focus on three concepts universal to living systems, biological process, molecular function and cellular component. As the vocabularies expand to incorporate terms needed by diverse annotation communities, species-specific terms become problematic. In particular, the use of species-specific anatomical concepts remains unresolved. We present a method for expansion of GO into areas outside of the three original universal concept domains. We combine concepts from two orthogonal vocabularies to generate a larger, more specific vocabulary. The example of mammalian heart development is presented because it addresses two issues that challenge GO; inclusion of organism-specific anatomical terms, and proliferation of terms and relationships. The combination of concepts from orthogonal vocabularies provides a robust representation of relevant terms and an opportunity for evaluation of hypothetical concepts.

Animals↗

BIOZON: a hub of heterogeneous biological data.

Biological entities are strongly related and mutually dependent on each other. Therefore, there is a growing need to corroborate and integrate data from different resources and aspects of biological systems in order to analyze them effectively. Biozon is a unified biological database that integrates heterogeneous data types such as proteins, structures, domain families, protein-protein interactions and cellular pathways, and establishes the relationships between them. All data are integrated on to a single graph schema centered around the non-redundant set of biological objects that are shared by each source. This integration results in a highly connected graph structure that provides a more complete picture of the known context of a given object that cannot be determined from any one source. Currently, Biozon integrates roughly 2 million protein sequences, 42 million DNA or RNA sequences, 32,000 protein structures, 150,000 interactions and more from sources such as GenBank, UniProt, Protein Data Bank (PDB) and BIND. Biozon augments source data with locally derived data such as 5 billion pairwise protein alignments and 8 million structural alignments. The user may form complex cross-type queries on the graph structure, add similarity relations to form fuzzy queries and rank the results based on analysis of the edge structure similar to Google PageRank, online at Biozon.org.

Computer Graphics↗

The EMBL Nucleotide Sequence Database.

The EMBL Nucleotide Sequence Database (http://www.ebi.ac.uk/embl), maintained at the European Bioinformatics Institute (EBI) near Cambridge, UK, is a comprehensive collection of nucleotide sequences and annotation from available public sources. The database is part of an international collaboration with DDBJ (Japan) and GenBank (USA). Data are exchanged daily between the collaborating institutes to achieve swift synchrony. Webin is the preferred tool for individual submissions of nucleotide sequences, including Third Party Annotation (TPA) and alignments. Automated procedures are provided for submissions from large-scale sequencing projects and data from the European Patent Office. New and updated data records are distributed daily and the whole EMBL Nucleotide Sequence Database is released four times a year. Access to the sequence data is provided via ftp and several WWW interfaces. With the web-based Sequence Retrieval System (SRS) it is also possible to link nucleotide data to other specialist molecular biology databases maintained at the EBI. Other tools are available for sequence similarity searching (e.g. FASTA and BLAST). Changes over the past year include the removal of the sequence length limit, the launch of the EMBLCDSs dataset, extension of the Sequence Version Archive functionality and the revision of quality rules for TPA data.

Base Sequence↗

Hubs of knowledge: using the functional link structure in Biozon to mine for biologically significant entities.

BACKGROUND: Existing biological databases support a variety of queries such as keyword or definition search. However, they do not provide any measure of relevance for the instances reported, and result sets are usually sorted arbitrarily. RESULTS: We describe a system that builds upon the complex infrastructure of the Biozon database and applies methods similar to those of Google to rank documents that match queries. We explore different prominence models and study the spectral properties of the corresponding data graphs. We evaluate the information content of principal and non-principal eigenspaces, and test various scoring functions which combine contributions from multiple eigenspaces. We also test the effect of similarity data and other variations which are unique to the biological knowledge domain on the quality of the results. Query result sets are assessed using a probabilistic approach that measures the significance of coherence between directly connected nodes in the data graph. This model allows us, for the first time, to compare different prominence models quantitatively and effectively and to observe unique trends. CONCLUSION: Our tests show that the ranked query results outperform unsorted results with respect to our significance measure and the top ranked entities are typically linked to many other biological entities. Our study resulted in a working ranking system of biological entities that was integrated into Biozon at http://biozon.org.

Abstracting and Indexing↗

RINGdb: an integrated database for G protein-coupled receptors and regulators of G protein signaling.

BACKGROUND: Many marketed therapeutic agents have been developed to modulate the function of G protein-coupled receptors (GPCRs). The regulators of G-protein signaling (RGS proteins) are also being examined as potential drug targets. To facilitate clinical and pharmacological research, we have developed a novel integrated biological database called RINGdb to provide comprehensive and organized RGS protein and GPCR information. RESULTS: RINGdb contains information on mutations, tissue distributions, protein-protein interactions, diseases/disorders and other features, which has been automatically collected from the Internet and manually extracted from the literature. In addition, RINGdb offers various user-friendly query functions to answer different questions about RGS proteins and GPCRs such as their possible contribution to disease processes, the putative direct or indirect relationship between RGS proteins and GPCRs. RINGdb also integrates organized database cross-references to allow users direct access to detailed information. The database is now available at http://ringdb.csie.ncu.edu.tw/ringdb/. CONCLUSION: RINGdb is the only integrated database on the Internet to provide comprehensive RGS protein and GPCR information. This knowledge base will be useful for clinical research, drug discovery and GPCR signaling pathway research.

Amino Acid Sequence↗

NIDDK data repository: a central collection of clinical trial data.

BACKGROUND: The National Institute of Diabetes and Digestive and Kidney Diseases have established central repositories for the collection of DNA, biological samples, and clinical data to be catalogued at a single site. Here we present an overview of the site which stores the clinical data and links to biospecimens. DESCRIPTION: The NIDDK Data repository is a web-enabled resource cataloguing clinical trial data and supporting information from NIDDK supported studies. The Data Repository allows for the co-location of multiple electronic datasets that were created as part of clinical investigations. The Data Repository does not serve the role of a Data Coordinating Center, but rather as a warehouse for the clinical findings once the trials have been completed. Because both biological and genetic samples are collected from many of the studies, a data management system for the cataloguing and retrieval of samples was developed. CONCLUSION: The Data Repository provides a unique resource for researchers in the clinical areas supported by NIDDK. In addition to providing a warehouse of data, Data Repository staff work with the users to educate them on the datasets as well as assist them in the acquisition of multiple data sets for cross-study analysis. Unlike the majority of biological databases, the Data Repository acts both as a catalogue for data, biosamples, and genetic materials and as a central processing point for the requests for all biospecimens. Due to regulations on the use of clinical data, the ultimate release of that data is governed under NIDDK data release policies. The Data Repository serves as the conduit for such requests.

Access to Information↗

Yeast acyl-CoA synthetases at the crossroads of fatty acid metabolism and regulation.

Acyl-CoA synthetases (ACSs) are a family of enzymes that catalyze the thioesterification of fatty acids with coenzymeA to form activated intermediates, which play a fundamental role in lipid metabolism and homeostasis of lipid-related processes. The products of the ACS enzyme reaction, acyl-CoAs, are required for complex lipid synthesis, energy production via beta-oxidation, protein acylation and fatty-acid dependent transcriptional regulation. ACS enzymes are also necessary for fatty acid import into cells by the process of vectorial acylation. The yeast Saccharomyces cerevisiae has four long chain ACS enzymes designated Faa1p through Faa4p, one very long chain ACS named Fat1p and one ACS, Fat2p, for which substrate specificity has not been defined. Pivotal roles have been defined for Faa1p and Faa4p in fatty acid import, beta-oxidation and transcriptional control mediated by the transcription factors Oaf1p/Pip2p and Mga2p/Spt23p. Fat1p is a bifunctional protein required for fatty acid transport of long chain fatty acids, as well as activation of very long chain fatty acids. This review focuses on the various roles yeast ACS enzymes play in cellular metabolism targeting especially the functions of specific isoforms in fatty acid transport, metabolism and energy production. We will also present evidence from directed experimentation, as well as information obtained by mining the molecular biological databases suggesting the long chain ACS enzymes are required in protein acylation, vesicular trafficking, signal transduction pathways and cell wall synthesis.

Acyl Coenzyme A↗

Development of the receptor database (RDB): application to the endocrine disruptor problem.

MOTIVATION: To represent various aspects of receptors effectively, we developed the receptor database (RDB), using an object-oriented database management system ACEDB and the Internet/WWW technology. RESULTS: RDB was constructed so that the system collects data items such as attributes of proteins from distributed data sources of the Internet, and so that it provides various viewing tools effectively, depending on different types of receptor data. Such sources include standard international biological databases such as the up-to-date database of PIR, Swiss Prot, PDB, GenBank and GDB. Application to the endocrine disruptor problem is presented. AVAILABILITY: RDB is available through the Internet at http://impact.nihs.go.jp/RDB.html.

Amino Acid Sequence↗

[The thirty years of Acta Genetica Sinica].

Acta Genetica Sinica (AGS) is sponsored by the Genetics Society of China and the Institute of Genetics and Developmental Biology of Chinese Academy of Sciences, and is published by Science Press. The journal is a leading national academic periodical and one of the Chinese key periodicals of natural sciences. Currently, AGS is being indexed by several well-known domestic and international indexing systems, such as the American Chemical Digest (CA), BIOSIS database, Biological Digest (BA), Medical Index and Russian Digest (P [symbol: see text]). Papers in the areas of genetics, developmental biology, cell molecular biology and evolution are regularly published by AGS.

China↗

Text mining for metabolic pathways, signaling cascades, and protein networks.

The complexity of the information stored in databases and publications on metabolic and signaling pathways, the high throughput of experimental data, and the growing number of publications make it imperative to provide systems to help the researcher navigate through these interrelated information resources. Text-mining methods have started to play a key role in the creation and maintenance of links between the information stored in biological databases and its original sources in the literature. These links will be extremely useful for database updating and curation, especially if a number of technical problems can be solved satisfactorily, including the identification of protein and gene names (entities in general) and the characterization of their types of interactions. The first generation of openly accessible text-mining systems, such as iHOP (Information Hyperlinked over Proteins), provides additional functions to facilitate the reconstruction of protein interaction networks, combine database and text information, and support the scientist in the formulation of novel hypotheses. The next challenge is the generation of comprehensive information regarding the general function of signaling pathways and protein interaction networks.

Animals↗

From genes to whole organs: connecting biochemistry to physiology.

The successful analysis of physiological processes requires quantitative understanding of the functional interactions between the key components of cells, organs and systems, and how these interactions change in disease states. This information does not reside in the genome, or even in the individual proteins that genes code for. There is therefore no alternative to copying nature and computing these interactions to determine the logic of healthy and diseased states. The rapid growth in biological databases, models of cells, tissues and organs, and in computing power has made it possible to explore functionality all the way from the level of genes to whole organs and systems. Examples are given of genetic modifications of the Na+ channel protein in the heart that predispose people to ventricular fibrillation, and of multiple target therapy in drug development. Complexity in biological systems also arises from tissue and organ geometry. This is illustrated using modelling of the whole heart.

Animals↗