PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Database links are a foundation for interoperability.

Several techniques are being introduced into the bioinformatics community to permit interoperation between molecular biology databases (DBs). The common factor to these approaches is the creation of links between entities in different DBs. Links can connect pieces of information about a single protein that are partitioned across multiple DBs, and can also encode relationships between different biological entities, such as relationships between an enzyme, its gene and its catalytic activity. This article provides an overview of the DB-interoperation problem, and offers several solutions. It discusses how links are used in molecular biology DBs, and describes the potential stumbling blocks when DB links are created and used.

Biotechnology↗

BLAST2SRS, a web server for flexible retrieval of related protein sequences in the SWISS-PROT and SPTrEMBL databases.

SRS (Sequence Retrieval System) is a widely used keyword search engine for querying biological databases. BLAST2 is the most widely used tool to query databases by sequence similarity search. These tools allow users to retrieve sequences by shared keyword or by shared similarity, with many public web servers available. However, with the increasingly large datasets available it is now quite common that a user is interested in some subset of homologous sequences but has no efficient way to restrict retrieval to that set. By allowing the user to control SRS from the BLAST output, BLAST2SRS (http://blast2srs.embl.de/) aims to meet this need. This server therefore combines the two ways to search sequence databases: similarity and keyword.

Animals↗

Genome-scale Gene Expression Analysis and Pathway Reconstruction in KEGG.

The massively parallel hybridization technologies by DNA chips and microarrays make it possible to monitor expression patterns of the whole set of genes in a genome under various conditions. The vast amount of data generated by such technologies necessitates the development of a new database management system that integrates expression data with other molecular biology databases and various analysis tools. We report here an extension of our KEGG (Kyoto Encyclopedia of Genes and Genomes) and DBGET/LinkDB systems for analyzing gene expression data in conjunction with pathway information and genomic information. It is now possible to make use of expression data for the reconstruction of pathways from the complete genome sequences.

Journal Article↗

Simplified user poll and experience report language (SUPER): implementation and application.

Biological computing is generally organized as standalone implementation on a PC-type computer or on a central facility (e.g. a university computer center). Services provided by central facilities need to be tailored to the user community. Unless very work-intensive individual contacts are used, the feedback must be collected with generalized tools, such as questionnaires distributed in the form of a newsletter. We have developed a method to have such polls automated and tailored, as well as having multiple-choice questions combined with branching after fundamental questions. As the evaluation of the results needs to know the questions asked, we have also included a method to process the answers and give detailed tables on the answers. SUPER was applied in a poll to query the academic usership in Switzerland on the usage of molecular biology databases.

Biology↗

SYSTOMONAS--an integrated database for systems biology analysis of Pseudomonas.

To provide an integrated bioinformatics platform for a systems biology approach to the biology of pseudomonads in infection and biotechnology the database SYSTOMONAS (SYSTems biology of pseudOMONAS) was established. Besides our own experimental metabolome, proteome and transcriptome data, various additional predictions of cellular processes, such as gene-regulatory networks were stored. Reconstruction of metabolic networks in SYSTOMONAS was achieved via comparative genomics. Broad data integration is realized using SOAP interfaces for the well established databases BRENDA, KEGG and PRODORIC. Several tools for the analysis of stored data and for the visualization of the corresponding results are provided, enabling a quick understanding of metabolic pathways, genomic arrangements or promoter structures of interest. The focus of SYSTOMONAS is on pseudomonads and in particular Pseudomonas aeruginosa, an opportunistic human pathogen. With this database we would like to encourage the Pseudomonas community to elucidate cellular processes of interest using an integrated systems biology strategy. The database is accessible at http://www.systomonas.de.

Bacterial Proteins↗

New methods for joint analysis of biological networks and expression data.

SUMMARY: Biological networks, such as protein interaction, regulatory or metabolic networks, derived from public databases, biological experiments or text mining can be useful for the analysis of high-throughput experimental data. We present two algorithms embedded in the ToPNet application that show promising performance in analyzing expression data in the context of such networks. First, the Significant Area Search algorithm detects subnetworks consisting of significantly regulated genes. These subnetworks often provide hints on which biological processes are affected in the measured conditions. Second, Pathway Queries allow detection of networks including molecules that are not necessarily significantly regulated, such as transcription factors or signaling proteins. Moreover, using these queries, the user can formulate biological hypotheses and check their validity with respect to experimental data. All resulting networks and pathways can be explored further using the interactive analysis tools provided by ToPNet program.

Algorithms↗

A real-time and dynamic biological information retrieval and analysis system (BIRAS).

The aim of this study is to design a biological information retrieval and analysis system (BIRAS) based on the Internet. Using the specific network protocol, BIRAS system could send and receive information from the Entrez search and retrieval system maintained by National Center for Biotechnology Information (NCBI) in USA. The literatures, nucleotide sequence, protein sequences, and other resources according to the user-defined term could then be retrieved and sent to the user by pop up message or by E-mail informing automatically using BIRAS system. All the information retrieving and analyzing processes are done in real-time. As a robust system for intelligently and dynamically retrieving and analyzing on the user-defined information, it is believed that BIRAS would be extensively used to retrieve specific information from large amount of biological databases in now days. The program is available on request from the corresponding author.

Animals↗

A compression mechanism for sequence databases to improve the efficiency of conventional tools.

This paper describes a method to compress molecular biology databases that are characterized by an increasing proportion of data derived from genome projects. The performance of our tool has been tested on various data files of the EMBL nucleotide sequence database. The best compression ratios were achieved on EST (Expressed Sequence Tags) data, typically derived from large-scale sequence projects. The compression of sequence database updates was tested in combination with the common Unix compression program 'compress'. Our tool improved the efficiency of 'compress' on average by 16%.

Base Sequence↗

PPMdb: a plant plasma membrane database.

PPMdb is a proteome database dedicated to proteins from plant plasma membranes. It provides comprehensive two-dimensional polyacrylamide gel electrophoresis (2-D PAGE) maps, partial amino acid sequences and expression data. All this information is gathered and structured in a relational database, after being analyzed and annotated. PPMdb includes active links to related biological databases (EMBL, GenBank, GenPep, and SWISS-PROT and TrEMBL) as well as to MEDLINE abstracts. Information on specific protein spots can be displayed by clicking on the 2-D maps. In addition, users can query the database by accession number, protein name, pI and MW, and cellular location. Access to PPMdb is available at the following URL: http://sphinx.rug. ac.be:8080.

Amino Acid Sequence↗

SIR: a simple indexing and retrieval system for biological flat file databases.

SUMMARY: SIR is a Simple Indexing and Retrieval tool for indexing and searching biological flat file databases. SIR is a cross-platform solution entirely written in Python. Since the package is very small and installation is trivial, this would be an ideal solution for database providers to provide a custom retrieval tool to access them. AVAILABILITY: The modules will be made available at http://www.EMBLHeidelberg.de/~chenna/PySAT/sir.html

Abstracting and Indexing↗

Biological Macromolecule Crystallization Database, Version 3.0: new features, data and the NASA archive for protein crystal growth data.

Version 3.0 of the NIST/NASA/CARB Biological Macromolecule Crystallization Database (BMCD) includes crystal and crystallization data on all forms of biological macromolecules which have produced crystals suitable for X-ray diffraction studies. The data include summary information on each of the macromolecules, crystal data, crystallization conditions and comments about the crystallization procedure if it varies from the traditional methods employed for crystal growth. The database-management software maintains continuity with previous versions providing similar search procedures and displays. Version 3.0 of the BMCD includes protocols and results of crystallization experiments undertaken in space. These new data are comprised of both the NASA Protein Crystal Growth Archive, which includes information on all NASA-sponsored protein crystal growth experiments, and data describing other internationally sponsored microgravity macromolecule crystallization studies. The entries for the space growth crystallization experiments contain the crystallization protocols, apparatus descriptions, flight summary data, indication of success or failure of the experiments, references, etc. Other new features of the BMCD include the addition of crystallization procedures for small peptides and cross references to other structural biology databases.

Journal Article↗

The EBI SRS server--recent developments.

MOTIVATION: The current data explosion is intractable without advanced data management systems. The numerous data sets become really useful when they are interconnected under a uniform interface--representing the domain knowledge. The SRS has become an integration system for both data retrieval and applications for data analysis. It provides capabilities to search multiple databases by shared attributes and to query across databases fast and efficiently. RESULTS: Here we present recent developments at the EBI SRS server (http://srs.ebi.ac.uk). The EBI SRS server contains today more than 130 biological databases and integrates more than 10 applications. It is a central resource for molecular biology data as well as a reference server for the latest developments in data integration. One of the latest additions to the EBI SRS server is the InterPro database-Integrated Resource of Protein Domains and Functional Sites. Distributed in XML format it became a turning point in low level XML-SRS integration. We present InterProScan as an example of data analysis applications, describe some advanced features of SRS6, and introduce the SRSQuickSearch JavaScript interfaces to SRS.

Computational Biology↗

DBGET/LinkDB: an integrated database retrieval system.

The integrated database retrieval system DBGET/LinkDB is the backbone of the Japanese GenomeNet service. DBGET is used to search and extract entries from a wide range of molecular biology databases, while LinkDB is used to search and compute links between entries in different databases. DBGET/LinkDB is designed to be a network distributed database system with an open architecture, which is suitable for incorporating local databases or establishing a specialized server environment. It also has an advantage of simple architecture allowing rapid daily updates of all the major databases. The WWW version of DBGET/LinkDB at GenomeNet is integrated with other search tools, such as BLAST, FASTA and MOTIF, and with local helper applications, such as RasMol. In addition to factual links between database entries, LinkDB is being extended to included similarity links and biological links toward computerization of logical reasoning processes.

Databases, Factual↗

Visualisation and navigation methods for typed protein-protein interaction networks.

Protein-protein interactions form large and complex networks. Their visualisation can aid biologists in gaining new insights about the processes in cells and is, therefore, very useful for building sophisticated research tools. Often standard force-directed graph drawing algorithms are used for the visualisation of these networks. However, currently available visual interfaces to biological databases only show general interactions and cannot cope well with more complex networks with different types of interactions. This paper presents a new approach to the visual analysis of protein-protein interaction networks. It uses a combination of circular and force-directed graph drawing algorithms to compute visual representations of protein networks depending on the type of the selected interaction. Smooth transitions between subsequent drawings enable users to explore different functional clusters in these networks without getting lost in the entire network. The visualisation system has been tested with data from the BRITE database.

Algorithms↗

Functional cartography of complex metabolic networks.

High-throughput techniques are leading to an explosive growth in the size of biological databases and creating the opportunity to revolutionize our understanding of life and disease. Interpretation of these data remains, however, a major scientific challenge. Here, we propose a methodology that enables us to extract and display information contained in complex networks. Specifically, we demonstrate that we can find functional modules in complex networks, and classify nodes into universal roles according to their pattern of intra- and inter-module connections. The method thus yields a 'cartographic representation' of complex networks. Metabolic networks are among the most challenging biological networks and, arguably, the ones with most potential for immediate applicability. We use our method to analyse the metabolic networks of twelve organisms from three different superkingdoms. We find that, typically, 80% of the nodes are only connected to other nodes within their respective modules, and that nodes with different roles are affected by different evolutionary constraints and pressures. Remarkably, we find that metabolites that participate in only a few reactions but that connect different modules are more conserved than hubs whose links are mostly within a single module.

Adenosine Triphosphate↗

Multi-scale methodology: a key to deciphering systems biology.

Presently, it is widely accepted complex systems couldn't be comprehended by studying parts in isolation without examining integrative and emergent properties, and system-level understanding thus has become the focus in biological science. However, it should also be noted that common systematic analysis was restricted to large-scale analysis at a certain level, while the facts that the nature of complex systems is their multi-scale structures was usually neglected or ignored. Therefore, this paper described a multi-scale methodology to investigate the nature of biological complexity and prospected this methodology could lead to a promising revolution in current system-level understanding and the integration of molecular biology databases.

Animals↗

ANDY: a general, fault-tolerant tool for database searching on computer clusters.

SUMMARY: ANDY (seArch coordination aND analYsis) is a set of Perl programs and modules for distributing large biological database searches, and in general any sequence of commands, across the nodes of a Linux computer cluster. ANDY is compatible with several commonly used distributed resource management (DRM) systems, and it can be easily extended to new DRMs. A distinctive feature of ANDY is the choice of either dedicated or fair-use operation: ANDY is almost as efficient as single-purpose tools that require a dedicated cluster, but it runs on a general-purpose cluster along with any other jobs scheduled by a DRM. Other features include communication through named pipes for performance, flexible customizable routines for error-checking and summarizing results, and multiple fault-tolerance mechanisms. AVAILABILITY: ANDY is freely available and can be obtained from http://compbio.berkeley.edu/proj/andy. SUPPLEMENTARY INFORMATION: Supplemental data, figures, and a more detailed overview of the software are found at http://compbio.berkeley.edu/proj/andy.

Computing Methodologies↗

Microarray annotation and biological information on function.

OBJECTIVES: Many methods for statistical analysis of gene expression studies by DNA microarrays produce lists of genes as output. To understand gene lists in terms of traditional biology, e.g. which pathways may be affected, it is necessary to get appropriate annotations for the probes on an array. METHODS: Problems arise with the different sources that have been used by manufacturers to design microarray probes, and their association to biological entities like genes, transcripts and proteins. Function annotation is of crucial importance, and systems like Gene Ontology can be used for this purpose. It arranges annotation terms in a hierarchical manner and thus makes annotations in a gene list amenable to automated analysis. RESULTS: Several methods for analyses of gene function are described. The hierarchical nature of systems like Gene Ontology particularly suggests using methods from graph theory. CONCLUSIONS: The main problem in annotating microarray probes and inferring affected functional modules is the incompleteness and degree of error in current biological databases. Initial approaches to make use of functional annotation exist, but have to be extended, in particular with respect to estimating the statistical significance of results.

Computational Biology↗