PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Current Comparative Table (CCT) automates customized searches of dynamic biological databases.

The Current Comparative Table (CCT) software program enables working biologists to automate customized bioinformatics searches, typically of remote sequence or HMM (hidden Markov model) databases. CCT currently supports BLAST, hmmpfam and other programs useful for gene and ortholog identification. The software is web based, has a BioPerl core and can be used remotely via a browser or locally on Mac OS X or Linux machines. CCT is particularly useful to scientists who study large sets of molecules in today's evolving information landscape because it color-codes all result files by age and highlights even tiny changes in sequence or annotation. By empowering non-bioinformaticians to automate custom searches and examine current results in context at a glance, CCT allows a remote database submission in the evening to influence the next morning's bench experiment. A demonstration of CCT is available at http://orb.public.stolaf.edu/CCTdemo and the open source software is freely available from http://sourceforge.net/projects/orb-cct.

Computational Biology↗

Hash function performance on different biological databases.

Open hashing is used to demonstrate the effectiveness of several hashing functions for the uniform distribution of biological records. The three types of database tested include (1) genetic nomenclature, mutation sites and strain names, (2) surnames extracted from literature files and (3) a set of 1000 numeric ASCII strings. Several hash functions (hashpjw, hashcrc and hashquad) showed considerable versatility on all data sets examined while two hash functions, hashsum and hashsmc, performed poorly, on the same databases.

Clinical Laboratory Information Systems↗

BioMart and Bioconductor: a powerful link between biological databases and microarray data analysis.

biomaRt is a new Bioconductor package that integrates BioMart data resources with data analysis software in Bioconductor. It can annotate a wide range of gene or gene product identifiers (e.g. Entrez-Gene and Affymetrix probe identifiers) with information such as gene symbol, chromosomal coordinates, Gene Ontology and OMIM annotation. Furthermore biomaRt enables retrieval of genomic sequences and single nucleotide polymorphism information, which can be used in data analysis. Fast and up-to-date data retrieval is possible as the package executes direct SQL queries to the BioMart databases (e.g. Ensembl). The biomaRt package provides a tight integration of large, public or locally installed BioMart databases with data analysis in Bioconductor creating a powerful environment for biological data mining.

Algorithms↗

Metric-space indexes as a basis for scalable biological databases.

Biochemical databases will be best served by the development of new specialized database management systems whose storage managers are based on metric-space indexing techniques and the development a database query languages that embody semantics derived from biochemical models of similarity and evolution. Important biochemical data types cannot be effectively mapped to low dimensional coordinate systems on which O(log n) indexing methods rely. It is clear from an abundance of bioinformatic discoveries that biochemical data is not random and exhibits interesting structure with respect to clustering. Metric-space indexing exploits a data set's intrinsic clustering to speed the execution of similarity queries, even when the data cannot be mapped to a coordinate system. Database management systems that seamlessly integrate semantically rich query languages with a metric-storage and retrieval mechanism will allow biologists to simply and concisely develop informatic studies that have traditionally been large and labor intensive.

Computational Biology↗

Metadata-based generation and management of knowledgebases from molecular biological databases.

Present-day knowledge-based systems (or expert systems) and databases constitute 'islands of computing' with little or no connection to each other. The use of software to provide a communication channel between the two, and to integrate their separate functions, is particularly attractive in certain data-rich domains where there are already pre-existing database systems containing the data required by the relevant knowledge-based system. Our evolving program, GENPRO, provides such a communication channel. The original methodology has been extended to provide interactive Prolog clause input with syntactic and semantic verification. This enables automatic generation of clauses from the source database, together with complete management of subsequent interfacing to the specified knowledge-based system. The particular data-rich domain used in this paper is protein structure, where processes which require reasoning (modelled by knowledge-based systems), such as the inference of protein topology, protein model-building and protein structure prediction, often require large amounts of raw data (i.e., facts about particular proteins) in the form of logic programming ground clauses. These are generated in the proper format by use of the concept of metadata.

Artificial Intelligence↗

Gene name identification and normalization using a model organism database.

Biology has now become an information science, and researchers are increasingly dependent on expert-curated biological databases to organize the findings from the published literature. We report here on a series of experiments related to the application of natural language processing to aid in the curation process for FlyBase. We focused on listing the normalized form of genes and gene products discussed in an article. We broke this into two steps: gene mention tagging in text, followed by normalization of gene names. For gene mention tagging, we adopted a statistical approach. To provide training data, we were able to reverse engineer the gene lists from the associated articles and abstracts, to generate text labeled (imperfectly) with gene mentions. We then evaluated the quality of the noisy training data (precision of 78%, recall 88%) and the quality of the HMM tagger output trained on this noisy data (precision 78%, recall 71%). In order to generate normalized gene lists, we explored two approaches. First, we explored simple pattern matching based on synonym lists to obtain a high recall/low precision system (recall 95%, precision 2%). Using a series of filters, we were able to improve precision to 50% with a recall of 72% (balanced F-measure of 0.59). Our second approach combined the HMM gene mention tagger with various filters to remove ambiguous mentions; this approach achieved an F-measure of 0.72 (precision 88%, recall 61%). These experiments indicate that the lexical resources provided by FlyBase are complete enough to achieve high recall on the gene list task, and that normalization requires accurate disambiguation; different strategies for tagging and normalization trade off recall for precision.

Abstracting and Indexing↗

Non-sequence databases for biological activity and physicochemical properties.

A biological activity database and a physicochemical property database are described. They are intended to complement the protein sequence database of PIR-International. The Biological Activity Database and the Physicochemical Property Database contain information regarding the biological activity and the physicochemical properties of proteins, respectively. In addition they also provide information about wild-type molecules with which information concerning variant molecules may be compared. Data on artificial variant molecules are stored in the Artificial Variant Database which is described separately.

Amino Acid Sequence↗

Web-based access to mouse models of human cancers: the Mouse Tumor Biology (MTB) Database.

The Mouse Tumor Biology (MTB) Database serves as a curated, integrated resource for information about tumor genetics and pathology in genetically defined strains of mice (i.e., inbred, transgenic and targeted mutation strains). Sources of information for the database include the published scientific literature and direct data submissions by the scientific community. Researchers access MTB using Web-based query forms and can use the database to answer such questions as 'What tumors have been reported in transgenic mice created on a C57BL/6J background?', 'What tumors in mice are associated with mutations in the Trp53 gene?' and 'What pathology images are available for tumors of the mammary gland regardless of genetic background?'. MTB has been available on the Web since 1998 from the Mouse Genome Informatics web site (http://www.informatics.jax.org). We have recently implemented a number of enhancements to MTB including new query options, redesigned query forms and results pages for pathology and genetic data, and the addition of an electronic data submission and annotation tool for pathology data.

Animals↗

Functional bioinformatics: the cellular response database.

Biological Scientists function in an increasingly data rich environment. The emerging field of bioinformatics is attempting to insure that this flow of information can be structured to support the generation of significant biological hypothesis and ultimately new knowledge. To date, most of the current databases have focused on protein and nucleic acid sequence information as the principle type of data stored for further interpretation. In this paper, we describe the Cellular Response Database. This database stores functional information regarding the changes of cellular gene expression associated with various stimuli, and supports queries linking cell types, expressed genes, and inducers. The database is designed to support information-intensive queries to aid in the determination of biological function, and is flexible enough to allow the storage of a broad range of experimental data such as cytotoxicity data, immunoassays of target gene protein expression, and others.

Cell Physiological Phenomena↗

Describing biological protein interactions in terms of protein states and state transitions: the LiveDIP database.

Biological protein-protein interactions differ from the more general class of physical interactions; in a biological interaction, both proteins must be in their proper states (e.g. covalently modified state, conformational state, cellular location state, etc.). Also in every biological interaction, one or both interacting molecules undergo a transition to a new state. This regulation of protein states through protein-protein interactions underlies many dynamic biological processes inside cells. Therefore, understanding biological interactions requires information on protein states. Toward this goal, DIP (the Database of Interacting Proteins) has been expanded to LiveDIP, which describes protein interactions by protein states and state transitions. This additional level of characterization permits a more complete picture of the protein-protein interaction networks and is crucial to an integrated understanding of genome-scale biology. The search tools provided by LiveDIP, Pathfinder, and Batch Search allow users to assemble biological pathways from all the protein-protein interactions collated from the scientific literature in LiveDIP. Tools have also been developed to integrate the protein-protein interaction networks of LiveDIP with large scale genomic data such as microarray data. An example of these tools applied to analyzing the pheromone response pathway in yeast suggests that the pathway functions in the context of a complex protein-protein interaction network. Seven of the eleven proteins involved in signal transduction are under negative or positive regulation of up to five other proteins through biological protein-protein interactions. During pheromone response, the mRNA expression levels of these signaling proteins exhibit different time course profiles. There is no simple correlation between changes in transcription levels and the signal intensity. This points to the importance of proteomic studies to understand how cells modulate and integrate signals. Integrating large scale, yeast two-hybrid data with mRNA expression data suggests biological interactions that may participate in pheromone response. These examples illustrate how LiveDIP provides data and tools for biological pathway discovery and pathway analysis.

Databases, Protein↗

BBID: the biological biochemical image database.

The Biological Biochemical Image Database is a WWW accessible relational database of archived images from research articles that describe regulatory pathways of higher eukaryotes. Pathway information is annotated and can be queried in the study of complex gene expression. In this way, complex regulatory pathways can be tested empirically in an efficient manner in the context of large-scale gene-expression systems.

Computer Graphics↗