PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Sharing neuroimaging studies of human cognition.

After more than a decade of collecting large neuroimaging datasets, neuroscientists are now working to archive these studies in publicly accessible databases. In particular, the fMRI Data Center (fMRIDC), a high-performance computing center managed by computer and brain scientists, seeks to catalogue and openly disseminate the data from published fMRI studies to the community. This repository enables experimental validation and allows researchers to combine and examine patterns of brain activity beyond that of any single study. As with some biological databases, early scientific, technical and sociological concerns hindered initial acceptance of the fMRIDC. However, with the continued growth of this and other neuroscience archives, researchers are recognizing the potential of such resources for identifying new knowledge about cognitive and neural activity. Thus, the field of neuroimaging is following the lead of biology and chemistry, mining its accumulating body of knowledge and moving toward a 'discovery science' of brain function.

Brain↗

Genomic comparison using data mining techniques based on a possibilistic fuzzy sets model.

Current copiousness of genomic information stored in biological databases [Mar Albà, M., Lee, M., Pearl, D., Shepherd, F.M.G., Martin, A.J., Orengo, N., Kellam, C.A., 2001. P. VIDA: a virus database system for the organisation of virus genome open reading frames. Nuleic Acids Res. 133-136] makes ultimately feasible the proposal for an application of knowledge management aimed to discover general rules in subcellular phenomena. The goal of this work is primarily to discover relationships between genes by microarray analysis. The tools exploited come from clustering techniques and are mainly based on Knowledge Discovery in Databases (KDD) concepts [Fayyad, U., Piatetsky-Shapiro, G., Smyth, P., 1996. From data mining to knowledge discovery in databases. AI Magazine 17(3), 37-54]. Starting from a data set, each element can be represented by a characteristic matrix, which sums up all data attributes. In this case data mining is oriented to perform a Pattern Recognition of related sequences, hidden in databases [Hand, D.J., Nicholas, A., 2005. Heard finding groups in gene expression data. J. Biomed. Biotechnol. 215-225]. Following a bottom up approach, the next refinement is to compare retrieved data to gather similar features, by dedicated clustering algorithms [Kaufman, L., Rousseeuw, P.J., 1990. Finding groups in data. An Introduction to Cluster Analysis. John Wiley & Sons, New York; Forman, G., Zhang, B., 2000. Distributed Data clustering can be efficient and exact HP. Laboratories Palo Alto HPL-2000, p. 158], driven by fuzzy logic, allowing us to perceive by intuition a common denominator for various genomic families and to anticipate likely future developments.

Algorithms↗

The European Bioinformatics Institute (EBI) databases.

This paper describes the databases and services of the European Bioinformatics Institute (EBI). In collaboration with DDBJ and GenBank/NCBI, the EBI maintains and distributes the EMBL Nucleotide Sequence Database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence Database, in collaboration with Amos Bairoch of the University of Geneva. Over thirty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists, are also available. The EBI network services include database searching, entry retrieval, and sequence similarity searching facilities.

Amino Acid Sequence↗

The European Bioinformatics Institute (EBI) databases.

The European Bioinformatics Institute (EBI) maintains and distributes the EMBL Nucleotide Sequence database, Europe's primary nucleotide sequence data resource. The EBI also maintains and distributes the SWISS-PROT Protein Sequence database, in collaboration with Amos Bairoch of the University of Geneva. Over fifty additional specialist molecular biology databases, as well as software and documentation of interest to molecular biologists are available. The EBI network services include database searching and sequence similarity searching facilities.

Amino Acid Sequence↗

Information resources at the National Center for Biotechnology Information.

The National Center for Biotechnology Information (NCBI), part of the National Library of Medicine, was established in 1988 to perform basic research in the field of computational molecular biology as well as build and distribute molecular biology databases. The basic research has led to new algorithms and analysis tools for interpreting genomic data and has been instrumental in the discovery of human disease genes for neurofibromatosis and Kallmann syndrome. The principal database responsibility is the National Institutes of Health (NIH) genetic sequence database, GenBank. NCBI, in collaboration with international partners, builds, distributes, and provides online and CD-ROM access to over 112,000 DNA sequences. Another major program is the integration of multiple sequences databases and related bibliographic information and the development of network-based retrieval systems for Internet access.

Animals↗

Convergent analysis of cDNA and short oligomer microarrays, mouse null mutants and bioinformatics resources to study complex traits.

Gene expression data sets have recently been exploited to study genetic factors that modulate complex traits. However, it has been challenging to establish a direct link between variation in patterns of gene expression and variation in higher order traits such as neuropharmacological responses and patterns of behavior. Here we illustrate an approach that combines gene expression data with new bioinformatics resources to discover genes that potentially modulate behavior. We have exploited three complementary genetic models to obtain convergent evidence that differential expression of a subset of genes and molecular pathways influences ethanol-induced conditioned taste aversion (CTA). As a first step, cDNA microarrays were used to compare gene expression profiles of two null mutant mouse lines with difference in ethanol-induced aversion. Mice lacking a functional copy of G protein-gated potassium channel subunit 2 (Girk2) show a decrease in the aversive effects of ethanol, whereas preproenkephalin (Penk) null mutant mice show the opposite response. We hypothesize that these behavioral differences are generated in part by alterations in expression downstream of the null alleles. We then exploited the WebQTL databases to examine the genetic covariance between mRNA expression levels and measurements of ethanol-induced CTA in BXD recombinant inbred (RI) strains. Finally, we identified a subset of genes and functional groups associated with ethanol-induced CTA in both null mutant lines and BXD RI strains. Collectively, these approaches highlight the phosphatidylinositol signaling pathway and identify several genes including protein kinase C beta isoform and preproenkephalin in regulation of ethanol- induced conditioned taste aversion. Our results point to the increasing potential of the convergent approach and biological databases to investigate genetic mechanisms of complex traits.

Animals↗

Web-based information retrieval system for the prediction of metabolic pathways.

Analysis of metabolic pathways is a central topic in understanding the relationship between genotype and phenotype. The rapid accumulation of biological data provides the possibility of studying metabolic pathways both at the genomic and metabolic levels. Our motivation is to develop a conceptual framework and computational system that will allow retrieval of metabolic information and prediction of metabolic pathways. In this paper, we introduce a metabolic pathway prediction framework that extracts metabolic information from biological databases via the Internet, and builds metabolic pathways with data sources of genes, sequences, enzymes, metabolites, etc. It provides an easy-to-use interface to retrieve, display, and manipulate metabolic information. The system has been implemented into PathAligner, available at http://bibiserv.techfak.uni-bielefeld. de/pathaligner/.

Computer Simulation↗

A case study in pathway knowledgebase verification.

BACKGROUND: Biological databases and pathway knowledge-bases are proliferating rapidly. We are developing software tools for computer-aided hypothesis design and evaluation, and we would like our tools to take advantage of the information stored in these repositories. But before we can reliably use a pathway knowledge-base as a data source, we need to proofread it to ensure that it can fully support computer-aided information integration and inference. RESULTS: We design a series of logical tests to detect potential problems we might encounter using a particular knowledge-base, the Reactome database, with a particular computer-aided hypothesis evaluation tool, HyBrow. We develop an explicit formal language from the language implicit in the Reactome data format and specify a logic to evaluate models expressed using this language. We use the formalism of finite model theory in this work. We then use this logic to formulate tests for desirable properties (such as completeness, consistency, and well-formedness) for pathways stored in Reactome. We apply these tests to the publicly available Reactome releases (releases 10 through 14) and compare the results, which highlight Reactome's steady improvement in terms of decreasing inconsistencies. We also investigate and discuss Reactome's potential for supporting computer-aided inference tools. CONCLUSION: The case study described in this work demonstrates that it is possible to use our model theory based approach to identify problems one might encounter using a knowledge-base to support hypothesis evaluation tools. The methodology we use is general and is in no way restricted to the specific knowledge-base employed in this case study. Future application of this methodology will enable us to compare pathway resources with respect to the generic properties such resources will need to possess if they are to support automated reasoning.

Algorithms↗

A fact database for toxicological data at the National Institute of Hygienic Sciences, Japan.

The computerized fact database for the toxicity data of chemicals was constructed at the National Institute of Hygienic Sciences, Tokyo, Japan (biological database, BL-DB). The BL-DB stores data on mutagenicity, teratogenicity, carcinogenicity, and other toxicological tests of chemicals that appeared in the scientific literature. The BL-DB includes information about chemical identification, test system, results of the assays, and a bibliography. The system consists of five modules: data collection, data maintenance, data search, data downloading, and backup. ADABAS is used as a core database management system. Many kinds of test data are stored with the same formats; therefore, users can retrieve data of different toxicological data by the same manner. A user of the BL-DB can use about 50 kinds of commands to interact with the system, and the majority of fields are defined as search fields, thereby facilitating retrieval of target data through many ways. Currently, there are mainly data for the mutagenicity, especially on the Salmonella/microsome assay and the rodent micronucleus assay. These data can be retrieved and used for structure-activity relationship studies.

Abnormalities, Drug-Induced↗

A relational database of transcription factors.

Recent advances in the understanding of eukaryotic gene regulation have produced an extensive body of transcriptionally-related sequence information in the biological literature, and have created a need for computing structures that organize and manage this information. The 'relational model' represents an approach that is finding increasing application in the design of biological databases. This report describes the compilation of information regarding eukaryotic transcription factors, the organization of this information into five tables, the computational applications of the resultant relational database that are of theoretical as well as experimental interest, and possible avenues of further development.

Amino Acid Sequence↗

The future: putting Humpty-Dumpty together again.

Successful biological analysis requires that we understand the functional interactions between key components of cells, organs and systems, and how these interactions change in disease. This information resides neither in the genome nor in the individual proteins that genes encode. It lies at the level of protein interactions within the context of sub-cellular, cellular, tissue, organ and system structures. There is therefore no alternative to copying Nature and computing these interactions to determine the logic of healthy and diseased states. The rapid growth in biological databases, models of cells, tissues and organs, and the development of powerful computing hardware and algorithms have made it possible to explore functionality in a quantitative manner all the way from the level of genes to the physiological function of whole organs and regulatory systems. Systems biology of the 21st century is set to become highly quantitative, and therefore one of the most computer-intensive disciplines.

Algorithms↗

Molecular immunology databases and data repositories.

Over recent years databases have become an extremely important resource for biomedical research. Immunology research is increasingly dependent on access to extensive biological databases to extract existing information, plan experiments, and analyse experimental results. This review describes 15 immunological databases that have appeared over the last 30 years. In addition, important issues regarding database design and the potential for misuse of information contained within these databases are discussed. Access pointers are provided for the major immunological databases and also for a number of other immunological resources accessible over the World Wide Web (WWW).

Allergy and Immunology↗

MaGe: a microbial genome annotation system supported by synteny results.

Magnifying Genomes (MaGe) is a microbial genome annotation system based on a relational database containing information on bacterial genomes, as well as a web interface to achieve genome annotation projects. Our system allows one to initiate the annotation of a genome at the early stage of the finishing phase. MaGe's main features are (i) integration of annotation data from bacterial genomes enhanced by a gene coding re-annotation process using accurate gene models, (ii) integration of results obtained with a wide range of bioinformatics methods, among which exploration of gene context by searching for conserved synteny and reconstruction of metabolic pathways, (iii) an advanced web interface allowing multiple users to refine the automatic assignment of gene product functions. MaGe is also linked to numerous well-known biological databases and systems. Our system has been thoroughly tested during the annotation of complete bacterial genomes (Acinetobacter baylyi ADP1, Pseudoalteromonas haloplanktis, Frankia alni) and is currently used in the context of several new microbial genome annotation projects. In addition, MaGe allows for annotation curation and exploration of already published genomes from various genera (e.g. Yersinia, Bacillus and Neisseria). MaGe can be accessed at http://www.genoscope.cns.fr/agc/mage.

Computational Biology↗

PathAligner: metabolic pathway retrieval and alignment.

MOTIVATION: Analysis of metabolic pathways is a central topic in understanding the relationship between genotype and phenotype. The rapid accumulation of biological data provides the possibility of studying metabolic pathways at both the genomic and the metabolic levels. Retrieving metabolic pathways from current biological data sources, reconstructing metabolic pathways from rudimentary pathway components, and aligning metabolic pathways with each other are major tasks. Our motivation was to develop a conceptual framework and computational system that allows the retrieval of metabolic pathway information and the processing of alignments to reveal the similarities between metabolic pathways. RESULTS: PathAligner extracts metabolic information from biological databases via the Internet and builds metabolic pathways with data sources of genes, sequences, enzymes, metabolites etc. It provides an easy-to-use interface to retrieve, display and manipulate metabolic information. PathAligner also provides an alignment method to compare the similarity between metabolic pathways. AVAILABILITY: PathAligner is available at http://bibiserv.techfak.uni-bielefeld.de/pathaligner.

Algorithms↗

PROMPT: a protein mapping and comparison tool.

BACKGROUND: Comparison of large protein datasets has become a standard task in bioinformatics. Typically researchers wish to know whether one group of proteins is significantly enriched in certain annotation attributes or sequence properties compared to another group, and whether this enrichment is statistically significant. In order to conduct such comparisons it is often required to integrate molecular sequence data and experimental information from disparate incompatible sources. While many specialized programs exist for comparisons of this kind in individual problem domains, such as expression data analysis, no generic software solution capable of addressing a wide spectrum of routine tasks in comparative proteomics is currently available. RESULTS: PROMPT is a comprehensive bioinformatics software environment which enables the user to compare arbitrary protein sequence sets, revealing statistically significant differences in their annotation features. It allows automatic retrieval and integration of data from a multitude of molecular biological databases as well as from a custom XML format. Similarity-based mapping of sequence IDs makes it possible to link experimental information obtained from different sources despite discrepancies in gene identifiers and minor sequence variation. PROMPT provides a full set of statistical procedures to address the following four use cases: i) comparison of the frequencies of categorical annotations between two sets, ii) enrichment of nominal features in one set with respect to another one, iii) comparison of numeric distributions, and iv) correlation of numeric variables. Analysis results can be visualized in the form of plots and spreadsheets and exported in various formats, including Microsoft Excel. CONCLUSION: PROMPT is a versatile, platform-independent, easily expandable, stand-alone application designed to be a practical workhorse in analysing and mining protein sequences and associated annotation. The availability of the Java Application Programming Interface and scripting capabilities on one hand, and the intuitive Graphical User Interface with context-sensitive help system on the other, make it equally accessible to professional bioinformaticians and biologically-oriented users. PROMPT is freely available for academic users from http://webclu.bio.wzw.tum.de/prompt/.

Computational Biology↗

The Eukaryotic Promoter Database EPD: the impact of in silico primer extension.

The Eukaryotic Promoter Database (EPD) is an annotated non-redundant collection of eukaryotic POL II promoters, experimentally defined by a transcription start site (TSS). There may be multiple promoter entries for a single gene. The underlying experimental evidence comes from journal articles and, starting from release 73, from 5' ESTs of full-length cDNA clones used for so-called in silico primer extension. Access to promoter sequences is provided by pointers to TSS positions in nucleotide sequence entries. The annotation part of an EPD entry includes a description of the type and source of the initiation site mapping data, links to other biological databases and bibliographic references. EPD is structured in a way that facilitates dynamic extraction of biologically meaningful promoter subsets for comparative sequence analysis. Web-based interfaces have been developed that enable the user to view EPD entries in different formats, to select and extract promoter sequences according to a variety of criteria and to navigate to related databases exploiting different cross-references. Tools for analysing sequence motifs around TSSs defined in EPD are provided by the signal search analysis server. EPD can be accessed at http://www.epd. isb-sib.ch.

Animals↗

Systems-ADME/Tox: resources and network approaches.

The increasing cost of drug development is partially due to our failure to identify undesirable compounds at an early enough stage of development. The application of higher throughput screening methods have resulted in the generation of very large datasets from cells in vitro or from in vivo experiments following the treatment with drugs or known toxins. In recent years the development of systems biology, databases and pathway software has enabled the analysis of the high-throughput data in the context of the whole cell. One of the latest technology paradigms to be applied alongside the existing in vitro and computational models for absorption, distribution, metabolism, excretion and toxicology (ADME/Tox) involves the integration of complex multidimensional datasets, termed toxicogenomics. The goal is to provide a more complete understanding of the effects a molecule might have on the entire biological system. However, due to the sheer complexity of this data it may be necessary to apply one or more different types of computational approaches that have as yet not been fully utilized in this field. The present review describes the data generated currently and introduces computational approaches as a component of ADME/Tox. These methods include network algorithms and manually curated databases of interactions that have been separately classified under systems biology methods. The integration of these disparate tools will result in systems-ADME/Tox and it is important to understand exactly what data resources and technologies are available and applicable. Examples of networks derived with important drug transporters and drug metabolizing enzymes are provided to demonstrate the network technologies.

ATP Binding Cassette Transporter 1↗

PACRAT: a database and analysis system for archaeal and bacterial intergenic sequence features.

Analysis of intergenic sequences for purposes such as the investigation of transcriptional signals or the identification of small RNA genes is frequently complicated by traditional biological database structures. Genome data is commonly treated as chromosome-length sequence records, detailed by gene calls demarcating subsequences of the chromosomes. Given this model, the determination of non-called subsequences between any gene and its nearest neighbors requires an exhaustive search of all gene calls associated with the chromosome. Further compounding the issue, the location of intergenic regions for many called genes cannot be resolved unambiguously due to uncertainties in gene boundaries, as well as the presence of other conflicting gene calls. To address these difficulties we have constructed the PACRAT (http://www.biosci.ohio-state.edu/~pacrat/) database system. PACRAT preprocesses GenBank genome submissions, evaluates for every gene the character of its relationship to those genes nearest to it, and produces a relationally linked model of the gene ordering for the genome. Using this information, the interface allows the researcher to query gene data as well as intergenic sequence data based on a number of criteria. These include the ability to filter searches based on the status of start and stop positions, or upstream/downstream sequences as conflicting with called genes and automated extension of upstream or downstream searches to find probable operon promoters or terminators. The database is also indexed by KEGG classification, allowing, for example, functionally-related groups of high-quality promoter-containing regions to be easily retrieved as a group.

DNA, Archaeal↗