PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Multi-class protein fold classification using a new ensemble machine learning approach.

Protein structure classification represents an important process in understanding the associations between sequence and structure as well as possible functional and evolutionary relationships. Recent structural genomics initiatives and other high-throughput experiments have populated the biological databases at a rapid pace. The amount of structural data has made traditional methods such as manual inspection of the protein structure become impossible. Machine learning has been widely applied to bioinformatics and has gained a lot of success in this research area. This work proposes a novel ensemble machine learning method that improves the coverage of the classifiers under the multi-class imbalanced sample sets by integrating knowledge induced from different base classifiers, and we illustrate this idea in classifying multi-class SCOP protein fold data. We have compared our approach with PART and show that our method improves the sensitivity of the classifier in protein fold classification. Furthermore, we have extended this method to learning over multiple data types, preserving the independence of their corresponding data sources, and show that our new approach performs at least as well as the traditional technique over a single joined data source. These experimental results are encouraging, and can be applied to other bioinformatics problems similarly characterised by multi-class imbalanced data sets held in multiple data sources.

Amino Acid Sequence↗

The BioImage Database Project: organizing multidimensional biological images in an object-relational database.

The BioImage Database Project collects and structures multidimensional data sets recorded by various microscopic techniques relevant to modern life sciences. It provides, as precisely as possible, the circumstances in which the sample was prepared and the data were recorded. It grants access to the actual data and maintains links between related data sets. In order to promote the interdisciplinary approach of modern science, it offers a large set of key words, which covers essentially all aspects of microscopy. Nonspecialists can, therefore, access and retrieve significant information recorded and submitted by specialists in other areas. A key issue of the undertaking is to exploit the available technology and to provide a well-defined yet flexible structure for dealing with data. Its pivotal element is, therefore, a modern object relational database that structures the metadata and ameliorates the provision of a complete service. The BioImage database can be accessed through the Internet.

Copyright↗

Functional inferences from reconstructed evolutionary biology involving rectified databases--an evolutionarily grounded approach to functional genomics.

If bioinformatics tools are constructed to reproduce the natural, evolutionary history of the biosphere, they offer powerful approaches to some of the most difficult tasks in genomics, including the organization and retrieval of sequence data, the updating of massive genomic databases, the detection of database error, the assignment of introns, the prediction of protein conformation from protein sequences, the detection of distant homologs, the assignment of function to open reading frames, the identification of biochemical pathways from genomic data, and the construction of a comprehensive model correlating the history of biomolecules with the history of planet Earth.

Amino Acid Sequence↗

Biological ontologies in rice databases. An introduction to the activities in Gramene and Oryzabase.

An enormous amount of information and materials in the field of biology has been accumulating, such as nucleotide and amino acid sequences, gene and protein functions, mutants and their phenotypes, and literature references, produced by the rapid development in this field. Effective use of the information may strongly promote biological studies, and may lead to many important findings. It is, however, time-consuming and laborious for individual researchers to collect information from individual original sites and to rearrange it for their own purpose. A concept, ontology, has been introduced in biology to support and encourage researchers to share and reuse information among biological databases. Ontology has a glossary, named dynamic controlled vocabulary, in which relationships between terms are defined. Since each term is strictly defined and identified with an ID number, a set of data represented in biological ontology is easily accessible to automated information processing, even if the data sets are across several databases and/or different organisms. In this mini-review, we introduce activities in Gramene and Oryzabase, which provide biological ontologies for Oryza sativa (rice).

Databases, Genetic↗

Improving interoperability between microbial information and sequence databases.

BACKGROUND: Biological resources are essential tools for biomedical research. Their availability is promoted through on-line catalogues. Common Access to Biological Resources and Information (CABRI) is a service for distribution of biological resources and related data collected by 28 European culture collections. Linking this information to bioinformatics databanks can make the collections' holdings more visible after a search in molecular biology databanks and vice-versa. Identification of links to sequence databases can be useful, but annotation and indexing problems, together with compilation errors, immediately arise. In this paper, we present our efforts for the identification of cross-references between CABRI catalogues and the EMBL Data Library and related results. RESULTS: An SRS site with both EMBL and CABRI catalogues has been set up. Ad-hoc changes in indexing scripts allowed to achieve homogeneous index keys and SRS link features have been used to identify links between databases. After manual checking and comparison with an alternative procedure, about 67,500 valid cross-references were identified, added to the EMBL Data Library and are now distributed with it. HTML links can be established from EMBL to CABRI network service. Procedures can be executed whenever needed. CONCLUSION: Links between EMBL and CABRI catalogues constitute an improved access to micro-organisms of certified quality and can produce positive effects on biomedical research. Further links between CABRI catalogues and other bioinformatics databases can now easily be defined by using these cross-references. Linking genetic information onto natural resources information may stand model for the integration of other databases containing empirical data on these materials.

Base Sequence↗

Seamless integration of biological applications within a database framework.

There are more than two hundred biological data repositories available for public access, and a vast number of applications to process and interpret biological data. A major challenge for bioinformaticians is to extract and process data from multiple data sources using a variety of query interfaces and analytical tools. In this paper, we describe tools that respond to this challenge by providing support for cross-database queries and for integrating analytical tools in a query processing environment. In particular, we describe two alternative methods for integrating biological data processing within traditional database queries: (a) "light-weight" application integration based on Application Specific Data Types (ASDTs) and (b) "heavy-duty" integration of analytical tools based on mediators and wrappers. These methods are supported by the Object-Protocol Model (OPM) suite of tools for managing biological databases.

Computational Biology↗

Patome: a database server for biological sequence annotation and analysis in issued patents and published patent applications.

With the advent of automated and high-throughput techniques, the number of patent applications containing biological sequences has been increasing rapidly. However, they have attracted relatively little attention compared to other sequence resources. We have built a database server called Patome, which contains biological sequence data disclosed in patents and published applications, as well as their analysis information. The analysis is divided into two steps. The first is an annotation step in which the disclosed sequences were annotated with RefSeq database. The second is an association step where the sequences were linked to Entrez Gene, OMIM and GO databases, and their results were saved as a gene-patent table. From the analysis, we found that 55% of human genes were associated with patenting. The gene-patent table can be used to identify whether a particular gene or disease is related to patenting. Patome is available at http://www.patome.org/; the information is updated bimonthly.

Amino Acid Sequence↗

Available pathways database (APD): an essential resource for combinatorial biology.

A relational database, the Available Pathways Database (APD), has been constructed of microbial natural products, their producing strains, and their biosynthetic pathways. The database allows the ready selection of donor strains for combinatorial biology experiments. It provides the same type of resource for combinatorial biology as the Available Chemicals Directory (ACD) does for combinatorial chemical library generation. Its cataloging ability can also provide insight into novel aspects of biosynthetic routes. In particular, no 10-unit Type I polyketides were found in the compilation of this edition of the APD (Version I).

Bacteria↗

Database techniques for biological materials & methods.

The Biological sciences produce an enormous research literature every year. Research papers are highly structured documents whose content is not captured using the traditional techniques of information retrieval: keywords and flat text. This is especially true of the Materials & Methods section of experimental papers. A great deal of highly structured information is packed into this section. It involves logical and temporal sequences of operations that combine and operate on materials using various instruments and depending on many parameters. We are designing and implementing databases that will allow this complex knowledge to be represented, stored in object-oriented databases and retrieved. We are developing an application of this technology called the Laboratory Notebook. This application is a software system that will contain personal laboratory information as well as have access to databases of Materials & Methods sections drawn from the literature.

Artificial Intelligence↗

A service-oriented information sources database for the biological sciences.

Researchers in the biological sciences require access to a variety of information sources located in various places on different computer networks. In order to satisfy the information needs of a researcher, appropriate information sources must be selected and access to these information sources and the computing services supporting them must be provided in a way that does not distract the researcher from problems of real interest. At the University of Missouri-Columbia a service-oriented information sources database is being developed as a key component of a layered-model design of an intelligent system which will provide a research environment appropriate to the needs of researchers in the biological sciences.

Artificial Intelligence↗

A CORBA server for the Radiation Hybrid DataBase.

Modern biology depends on a wide range of software interacting with a large number of data sources, varying both in size, complexity and structure. The range of important databases in molecular biology and genetics makes it crucial to overcome the problems which this multiplicity presents. At EMBL-EBI we have started to use CORBA technology to support interoperability between a variety of databases, as well as to facilitate the integration of tools that access these databases. Within the Radiation Hybrid DataBase project we are confronted daily with the interoperation and linking issues. In this paper we present a CORBA infrastructure implemented to access the Radiation Hybrid DataBase.

Animals↗

Databases in molecular biology: a CODATA task group at work.

A certain concern exists that the exponential growth of nucleic acid and protein sequence data will saturate the channels of data acquisition, distribution and utilization on the one hand and, on the other hand, that even the actual resources are still not fully and easily accessible to any bench scientist. Despite the stake of the scientific community at large in the fundamental data collected in this field, there has been in past years only a modest effort to discuss the common problems at an international level. Three international meetings were organized in 1987 on this subject: the annual meeting of CODATA Task Group on Coordination of Protein Sequence Data Banks (Nice, France, January 1987), the EMBL/NIH Workshop concerned primarily with nucleic acid databases (Heidelberg, FRG, February 1987) and the CODATA Workshop on Nucleic Acid and Protein Sequencing Data (Gaithersburg, USA, May 1987).

Amino Acid Sequence↗

Construction of a database of benzene biological monitoring.

Biological monitoring of occupational exposure to benzene has been conducted in the petroleum, steel and chemical industries. The urinary benzene-specific biomarker, S-phenylmercapturic acid (PMA), was quantified in post-shift samples using a sensitive enzyme-linked immunosorbent assay (ELISA) and expressed as a function of urinary creatinine concentration. The assay, based on a PMA-specific antiserum, is sufficiently sensitive to measure PMA levels in non-occupationally exposed control subjects. The assay delivers batch results in a timely manner which may be as short as 3 h. Samples were analysed from groups of workers engaged in coke oven combustion processes, petroleum refining and decontamination of a benzene land spill. The construction of a database of results provides an index of benzene uptake as a consequence of the respective work processes and tasks and readily enables benchmarking exercises aimed at comparing degrees of exposure across segments of industry.

Acetylcysteine↗

Design and realization of an on-line database for multidimensional microscopic images of biological specimens.

The BioImage database is a new scientific database for multidimensional microscopic images of biological specimens, which is available through the World Wide Web (WWW). The development of this database has followed an iterative approach, in which requirements and functionality have been revised and extended. The complexity and innovative use of the data meant that technical and biological expertise has been crucial in the initial design of the data model. A controlled vocabulary was introduced to ensure data consistency. Pointers are used to reference information stored in other databases. The data model was built using InfoModeler as a database design tool. The database management system is the Informix Dynamic Server with Universal Data Option. This object-relational system allows the handling of complex data using features such as collection types, inheritance, and user-defined data types. Informix datablades are used to provide additional functionality: the Web Integration Option enables WWW access to the database; the Video Foundation Blade provides functionality for video handling.

Animals↗

Issues in incorporation semantic integrity in molecular biological object-oriented databases.

Issues critical to ensuring semantic integrity in molecular biological data collections have been identified and include complexity, exceptions, missing data, changing models, holism and integration, delocalized data, interoperability and nomenclature. This combination is peculiar to biology and presents some interesting problems as a result. Little is known about semantic checking in object-oriented databases in general, but because such technology appears highly suitable for modeling biological data, it is appropriate to examine the ways in which object-oriented technology can support this functionality. It is concluded that object-oriented technology will support semantic checking even in a complex domain like biology. We propose 10 guidelines for future work including ways of treating exceptional cases and 'positioning' of constraints in a schema.

Biotechnology↗