PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

AnoBase: a genetic and biological database of anophelines.

AnoBase (http://www.anobase.org) is an integrated, relational database of basic biological and genetic data on anopheline species, with a particular emphasis on Anopheles gambiae. It has been designed as an information source and research support tool for the broad vector biology community. Although AnoBase is not a primary genomic database that develops and provides tools to access the genome of the malaria mosquito, it nevertheless contains several sections that offer data of genomic interest such as in situ hybridization images, an integrated gene tool and direct online access to AnoXcel, the proteomic database of An. gambiae. Moreover, AnoBase also contains information on non-gambiae mosquito species and a novel section on studies related to insecticide resistance.

Animals↗

On optimizing distance-based similarity search for biological databases.

Similarity search leveraging distance-based index structures is increasingly being used for both multimedia and biological database applications. We consider distance-based indexing for three important biological data types, protein k-mers with the metric PAM model, DNA k-mers with Hamming distance and peptide fragmentation spectra with a pseudo-metric derived from cosine distance. To date, the primary driver of this research has been multimedia applications, where similarity functions are often Euclidean norms on high dimensional feature vectors. We develop results showing that the character of these biological workloads is different from multimedia workloads. In particular, they are not intrinsically very high dimensional, and deserving different optimization heuristics. Based on MVP-trees, we develop a pivot selection heuristic seeking centers and show it outperforms the most widely used corner seeking heuristic. Similarly, we develop a data partitioning approach sensitive to the actual data distribution in lieu of median splits.

Algorithms↗

An extensible network query unification system for biological databases.

Database federation enables biological researchers to utilize resources more effectively, creating an environment in which the researcher can query multiple data sources without spending time learning new query mechanisms or issuing redundant queries which need to be integrated. Several mechanisms exist to federate databases. The ENQUire system is a network database federation system which uses a World-Wide-Web (WWW) interface to connect the users to various databases. Generic queries entered via a query generator form are sent in parallel to multiple databases, and the results are presented to the user in a unified format. All forms building, query generation, and results translation is done on the fly, and individual database translation modules can be added dynamically. ENQUire is a flexible answer to the problems of database federation on the WWW.

Computer Communication Networks↗

The Mouse Tumor Biology Database: a public resource for cancer genetics and pathology of the mouse.

Developing genetic mouse models for cancer research has been recognized as an "exceptional opportunity" by the National Cancer Institute. The establishment of bioinformatics resources to facilitate access to published and unpublished data on the genetics and pathology of cancer in different strains of the laboratory mouse is critical to developing and using mouse models of human disease. In this article, we review the Mouse Tumor Biology Database (MTB), a public resource for information on cancer genetics, epidemiology, and pathology in genetically defined mice. We outline current content, data acquisition strategies, and query mechanisms for MTB. MTB is accessible on-line at http://tumor.informatics.jax.org.

Animals↗

Bioinformatic analysis of autism positional candidate genes using biological databases and computational gene network prediction.

Common genetic disorders are believed to arise from the combined effects of multiple inherited genetic variants acting in concert with environmental factors, such that any given DNA sequence variant may have only a marginal effect on disease outcome. As a consequence, the correlation between disease status and any given DNA marker allele in a genomewide linkage study tends to be relatively weak and the implicated regions typically encompass hundreds of positional candidate genes. Therefore, new strategies are needed to parse relatively large sets of 'positional' candidate genes in search of actual disease-related gene variants. Here we use biological databases to identify 383 positional candidate genes predicted by genomewide genetic linkage analysis of a large set of families, each with two or more members diagnosed with autism, or autism spectrum disorder (ASD). Next, we seek to identify a subset of biologically meaningful, high priority candidates. The strategy is to select autism candidate genes based on prior genetic evidence from the allelic association literature to query the known transcripts within the 1-LOD (logarithm of the odds) support interval for each region. We use recently developed bioinformatic programs that automatically search the biological literature to predict pathways of interacting genes (PATHWAYASSIST and GENEWAYS). To identify gene regulatory networks, we search for coexpression between candidate genes and positional candidates. The studies are intended both to inform studies of autism, and to illustrate and explore the increasing potential of bioinformatic approaches as a compliment to linkage analysis.

Autistic Disorder↗

Object-oriented parsing of biological databases with Python.

MOTIVATION: While database activities in the biological area are increasing rapidly, rather little is done in the area of parsing them in a simple and object-oriented way. RESULTS: We present here an elegant, simple yet powerful way of parsing biological flat-file databases. We have taken EMBL, SWISSPROT and GENBANK as examples. EMBL and SWISS-PROT do not differ much in the format structure. GENBANK has a very different format structure than EMBL and SWISS-PROT. Extracting the desired fields in an entry (for example a sub-sequence with an associated feature) for later analysis is a constant need in the biological sequence-analysis community: this is illustrated with tools to make new splice-site databases. The interface to the parser is abstract in the sense that the access to all the databases is independent from their different formats, since parsing instructions are hidden.

Databases, Factual↗

Recent developments in biological sequence databases.

Biological sequence databases are currently being re-engineered to make them more efficient and easier to use. This re-engineering is also providing an infrastructure to make it easier to interrogate and integrate data from different sources. The net result of this effort should be a great improvement in the power and availability of bioinformatics resources to the general biology community.

Databases, Factual↗

Integr8: enhanced inter-operability of European molecular biology databases.

OBJECTIVES: The increasing production of molecular biology data in the post-genomic era, and the proliferation of databases that store it, require the development of an integrative layer in database services to facilitate the synthesis of related information. The solution of this problem is made more difficult by the absence of universal identifiers for biological entities, and the breadth and variety of available data. METHODS: Integr8 was modelled using UML (Universal Modelling Language). Integr8 is being implemented as an n-tier system using a modern object-oriented programming language (Java). An object-relational mapping tool, OJB, is being used to specify the interface between the upper layers and an underlying relational database. RESULTS: The European Bioinformatics Institute is launching the Integr8 project. Integr8 will be an automatically populated database in which we will maintain stable identifiers for biological entities, describe their relationships with each other (in accordance with the central dogma of biology), and store equivalences between identified entities in the source databases. Only core data will be stored in Integr8, with web links to the source databases providing further information. CONCLUSIONS: Integr8 will provide the integrative layer of the next generation of bioinformatics services from the EBI. Web-based interfaces will be developed to offer gene-centric views of the integrated data, presenting (where known) the links between genome, proteome and phenotype.

Computational Biology↗

Mouse Tumor Biology Database (MTB): status update and future directions.

The Mouse Tumor Biology (MTB) database provides access to data about endogenously arising tumors (both spontaneous and induced) in genetically defined mice (inbred, hybrid, mutant and genetically engineered mice). Data include information on the frequency and latency of mouse tumors, pathology reports and images, genomic changes occurring in the tumors, genetic (strain) background and literature or contributor citations. Data are curated from the primary literature or submitted directly from researchers. MTB is accessed via the Mouse Genome Informatics web site (http://www.informatics.jax.org). Integrated searches of MTB are enabled through use of multiple controlled vocabularies and by adherence to standardized nomenclature, when available. Recently MTB has been redesigned and its database infrastructure replaced with a robust relational database management system (RDMS). Web interface improvements include a new advanced query form and enhancements to already existing search capabilities. The Tumor Frequency Grid has been revised to enhance interactivity, providing an overview of reported tumor incidence across mouse strains and an entrée into the database. A new pathology data submission tool allows users to submit, edit and release data to the MTB system.

Animals↗

A profile for molecular biology databases and information resources.

This paper examines the requirements for building database management systems and multi-database information resources to support molecular biology research. The paper profiles the most important features of 16 integrated resources and 102 databases related to molecular biology research. The aspects surveyed in this paper include the nature of information in these databases, their sizes, update properties, cross-references, database management system heterogeneity, geographical distribution, data quality, use of temporal information and level of interpretation. The paper also comments on the access patterns to these databases. Since not all these aspects were available for all databases, specific comparisons sometimes compare fewer than the full 102 databases. Consequently, the same set of databases is not necessarily always being compared with respect to every aspect. The paper is organized primarily according to these comparison aspects and ends with some concluding remarks.

Databases, Bibliographic↗

Supporting taxonomic names in cell and molecular biology databases.

Groups of organisms require labels or names to refer to them; however, the idea of a single static name index, although tempting for its simplicity, is both impractical and unadvisable as a basis for referring to organisms for which data has been collected and stored for analyses and sharing. The relevant issues are described and some of the challenges facing database researchers are discussed.

Classification↗

Molecular biological databases--present and future.

The importance of databases as a research tool in molecular biology is growing steadily, and a wide range of databases relevant to genome research is currently available. However, the design of current databases is inadequate for accurate representation and analysis of the results of large-scale genome mapping and sequencing projects. A new generation of databases is required to master the challenges of the future.

Animals↗