PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Financing a future for public biological data.

MOTIVATION: The public web-based biological database infrastructure is a source of both wonder and worry. Users delight in the ever increasing amounts of information available; database administrators and curators worry about long-term financial support. An earlier study of 153 biological databases (Ellis and Kalumbi, Nature Biotechnol., 16, 1323-1324, 1998) determined that near future (1-5 year) funding for over two-thirds of them was uncertain. More detailed data are required to determine the magnitude of the problem and offer possible solutions. METHODS: This study examines the finances and use statistics of a few of these organizations in more depth, and reviews several economic models that may help sustain them. RESULTS: Six organizations were studied. Their administrative overhead is fairly low; non-administrative personnel and computer-related costs account for 77% of expenses. One smaller, more specialized US database, in 1997, had 60% of total access from US domains; a majority (56%) of its US accesses came from commercial domains, although only 2% of the 153 databases originally studied received any industrial support. The most popular model used to gain industrial support is asymmetric pricing: preferentially charging the commercial users of a database. At least five biological databases have recently begun using this model. Advertising is another model which may be useful for the more general, more heavily used sites. Microcommerce has promise, especially for databases that do not attract advertisers, but needs further testing. The least income reported for any of the databases studied was $50,000/year; applying this rate to 400 biological databases (a lower limit of the number of such databases, many of which require far larger resources) would mean annual support need of at least $20 million. To obtain this level of support is challenging, yet failure to accept the challenge could be catastrophic. CONTACT: lynda@tc.umn. edu

Computational Biology↗

A strategy for database interoperation.

To realize the full potential of biological databases (DBs) requires more than the interactive, hypertext flavor of database interoperation that is now so popular in the bioinformatics community. Interoperation based on declarative queries to multiple network-accessible databases will support analyses and investigations that are orders of magnitude faster and more powerful than what can be accomplished through interactive navigation. I present a vision of the capabilities that a query-based interoperation infrastructure should provide, and identify assumptions underlying, and requirements of, this vision. I then propose an architecture for query-based interoperation that includes a number of novel components of an information infrastructure for molecular biology. These components include a knowledge base that describes relationships among the conceptualizations used in different biological databases, a module that can determine the DBs that are relevant to a particular query, a module that can translate a query and its results from one conceptualization to another, a collection of DB drivers that provide uniform physical access to different database management systems, a suite of translators that can interconvert among different database schema languages, and a database that describes the network location and access methods for biological databases. A number of the components are translators that bridge the heterogeneities that exist between biological DBs at several different levels, including the conceptual level, the data model, the query language, and data formats.

Artificial Intelligence↗

Reassembly and interfacing neural models registered on biological model databases.

The importance of modeling and simulation of biological process is growing for further understanding of living systems at all scales from molecular to cellular, organic, and individuals. In the field of neuroscience, there are so called platform simulators, the de-facto standard neural simulators. More than a hundred neural models are registered on the model database. These models are executable in corresponding simulation environments. But usability of the registered models is not sufficient. In order to make use of the model, the users have to identify the input, output and internal state variables and parameters of the models. The roles and units of each variable and parameter are not explicitly defined in the model files. These are suggested implicitly in the papers where the simulation results are demonstrated. In this study, we propose a novel method of reassembly and interfacing models registered on biological model database. The method was applied to the neural models registered on one of the typical biological model database, ModelDB. The results are described in detail with the hippocampal pyramidal neuron model. The model is executable in NEURON simulator environment, which demonstrates that somatic EPSP amplitude is independent of synapse location. Input and output parameters and variables were identified successfully, and the results of the simulation were recorded in the organized form with annotations.

Computational Biology↗

Database of biologically active peptide sequences.

Proteins are sources of many peptides with diverse biological activity. Such peptides are considered as valuable components of foods with desired and designed biological activity. Two strategies are currently recommended for research in the area of biological activity of food protein fragments. The first strategy covers investigations on products of enzymic hydrolysis of proteins. The second one is synthesis of peptides identical with protein fragments and investigations using these peptides. It is possible to predict biological activity of protein fragments using sequence alignments between proteins and biologically active peptides from database. Our database contains currently 527 sequences of bioactive peptides with antihypertensive, opioid, immunomodulating and other activities. The sequence alignments can give information about localization of biologically active fragments in protein chain, but not about possibilities of enzymic release of such fragments. The information is thus equivalent with this obtained using synthetic peptides identical with protein fragments. Possibilities offered by the database are discussed using wheat alpha/beta-gliadin, bovine beta-lactoglobulin and bovine beta-casein (including influence of genetic polymorphism and genetic engineering on amino acid sequences) as examples.

Animals↗

Storing biological sequence databases in relational form.

SUMMARY: We have created a set of applications using Perl and Java in combination with XML technology to install biological sequence databases into an Oracle RDBMS. An easy-to-use interface using Java has been created for database query and other tools developed to integrate with our in-house bioinformatics applications. AVAILIBILITY: The database schema, DTD file, and source codes are available from the authors via email. CONTACT: guochun_ xie@merck. com

Amino Acid Sequence↗

Current databases on biological variation: pros, cons and progress.

A database with reliable information to derive definitive analytical quality specifications for a large number of clinical laboratory tests was prepared in this work. This was achieved by comparing and correlating descriptive data and relevant observations with the biological variation information, an approach that had not been used in the previous efforts of this type. The material compiled in the database was obtained from published articles referenced in BIOS, CURRENT CONTENTS, EMBASE and MEDLINE using "biological variation & laboratory medicine" as key words, as well as books and doctoral theses provided by their authors. The database covers 316 quantities and reviews 191 articles, fewer than 10 of which had to be rejected. The within- and between-subject coefficients of variation and the subsequent desirable quality specifications for precision, bias and total error for all the quantities accepted are presented. Sex-related stratification of results was justified for only four quantities and, in these cases, quality specifications were derived from the group with lower within-subject variation. For certain quantities, biological variation in pathological states was higher than in the healthy state. In these cases, quality specifications were derived only from the healthy population (most stringent). Several quantities (particularly hormones) have been treated in very few articles and the results found are highly discrepant. Therefore, professionals in laboratory medicine should be strongly encouraged to study the quantities for which results are discrepant, the 90 quantities described in only one paper and the numerous quantities that have not been the subject of study.

Clinical Laboratory Techniques↗

Genomic pathways database and biological data management.

In this paper, we discuss the properties of biological data and challenges it poses for data management, and argue that, in order to meet the data management requirements for 'digital biology', careful integration of the existing technologies and the development of new data management techniques for biological data are needed. Based on this premise, we present PathCase: Case Pathways Database System. PathCase is an integrated set of software tools for modelling, storing, analysing, visualizing and querying biological pathways data at different levels of genetic, molecular, biochemical and organismal detail. The novel features of the system include: (i) genomic information integrated with other biological data and presented starting from pathways; (ii) design for biologists who are possibly unfamiliar with genomics, but whose research is essential for annotating gene and genome sequences with biological functions; (iii) database design, implementation and graphical tools which enable users to visualize pathways data in multiple abstraction levels and to pose exploratory queries; (iv) a wide range of different types of queries including, 'path' and 'neighbourhood queries' and graphical visualization of query outputs; and (v) an implementation that allows for web (XML)-based dissemination of query outputs (i.e. pathways data in BIOPAX format) to researchers in the community, giving them control on the use of pathways data.

Computational Biology↗

An overview of computer software developed to search biological sequence databases.

The scientific community has established a number of databases to receive and maintain the abundant biological sequence information being generated by research investigators worldwide. For researchers in an increasing number of biological disciplines, the information stored in these databases has become an invaluable tool in their daily research endeavors. This article reviews the organization of the largest nucleic acid and protein sequence databases and discusses some of the commonly available computer software that has been developed for searching them.

Amino Acid Sequence↗

Clinical bioinformatics.

Clinical bioinformatics provides biological and medical information to allow for individualized healthcare. In this review, we describe the uses of clinical bioinformatics. After the analysis of the complete human genome sequences, clinical bioinformatics enables researchers to search online biological databases and use the biological information in their medical practices. The data obtained from using microarray is extremely complicated. In clinical bioinformatics, selecting appropriate software to analyze the microarray data for medical decision making is crucial. Proteomics strategy tools usually focus on similarity searches, structure prediction, and protein modeling. In clinical bioinformatics, the proteomic data only have meaning if they are integrated with clinical data. In pharmacogenomics, clinical bioinformatics includes elaborate studies of bioinformatics tools and various facets of proteomics related to drug target identification and clinical validation. Using clinical bioinformatics, researchers apply computational and high-throughput experimental techniques to cancer research and systems biology. Meanwhile, researchers of bioinformatics and medical information have incorporated clinical bioinformatics to improve health care, using biological and medical information. Using the high volume of biological information from clinical bioinformatics will contribute to changes in practice standards in the healthcare system. We believe that clinical bioinformatics provides benefits of improving healthcare, disease prevention and health maintenance as we move toward the era of personalized medicine.

Computational Biology↗

Evaluation of human-readable annotation in biomolecular sequence databases with biological rule libraries.

MOTIVATION: Computer-based selection of entries from sequence databases with respect to a related functional description, e.g. with respect to a common cellular localization or contributing to the same phenotypic function, is a difficult task. Automatic semantic analysis of annotations is not only hampered by incomplete functional assignments. A major problem is that annotations are written in a rich, non-formalized language and are meant for reading by a human expert. This person can extract from the text considerably more information than is immediately apparent due to his extended biological background knowledge and logical reasoning. APPROACH: A technique of automated annotation evaluation based on a combination of lexical analysis and the usage of biological rule libraries has been developed. The proposed algorithm generates new functional descriptors from the annotation of a given entry using the semantic units of the annotation as prepositions for implications executed in accordance with the rule library. RESULTS: The prototype of a software system, the Meta_A(nnotator) program, is described and the results of its application to sequence attribute assignment and sequence selection problems, such as cellular localization and sequence domain annotation of SWISS-PROT entries, are presented. The current software version assigns useful subcellular localization qualifiers to approximately 88% of all SWISS-PROT entries. As shown by demonstrative examples, the combination of sequence and annotation analysis is a powerful approach for the detection of mutual annotation/sequence inconsistencies. AVAILABILITY: Results for the cellular localization assignment can be viewed at the URL http://www.bork. embl-heidelberg.de/CELL_LOC/CELL_LOC.html.

Algorithms↗

Beta-blocker supplementation of standard drug treatment for schizophrenia.

BACKGROUND: Many people with schizophrenia or similar severe mental disorders do not achieve a satisfactory treatment response with ordinary antipsychotic drug treatment. In these cases, various add-on medications are used, among them beta-adrenergic receptor antagonists (beta-blockers). OBJECTIVES: To evaluate the clinical effectiveness of beta-blockers as an adjunct to antipsychotic medication in schizophrenia or similar severe mental disorders. SEARCH STRATEGY: Publications in all languages were searched from the following databases: Biological Abstracts, CENTRAL of The Cochrane Library, Cochrane Schizophrenia Group's Specialised Register, EMBASE, LILACS, MEDLINE, and PsycLIT. The reference section of papers included were screened. SELECTION CRITERIA: All randomised controlled trials comparing beta-blockers with placebo as an adjunct to conventional antipsychotic medication for those with schizophrenia. DATA COLLECTION AND ANALYSIS: Studies were selected and then data extracted, independently, by at least two reviewers. Odds ratios and 95% confidence intervals of homogeneous dichotomous data were calculated with the Peto method. A random effects model was used for heterogeneous dichotomous data. Weighted mean differences were calculated for continuous data. MAIN RESULTS: Currently the review includes five studies but data are poorly presented and do not evidence any effect of beta-blockers as an adjunct to conventional antipsychotic medication. REVIEWER'S CONCLUSIONS: At present beta-blockers cannot be recommended in the treatment of schizophrenia. Any possible benefit of adjunctive beta-blockers is obscured by the poor reporting of the included studies. Existing data on beta-blockers as adjunctive medication to antipsychotics for those with schizophrenia should be collected and re-analysed in order to allow confident conclusions about the effect of this treatment or the need for further trials.

Adrenergic beta-Antagonists↗

[Bioinformatics and GenEnv database in biological risk management].

Identification and molecular typing of environmental isolates by molecular techniques requires knowledge of the genetic characteristics of the microbe species being examined. The introduction of automated sequences has greatly speeded up the entire sequencing process as well as improved the accuracy of the collected information. Bioinformatics tools have become indispensable not only for setting up research studies, but also for storing, organizing and managing enormous quantities of sequencing data. Despite its great advantages, the use of bioinformatics is hindered by difficulties in learning how to use its software tools. The GenEnv database was developed to provide operators involved in biological risk management with a user-friendly tool for sequence analysis. Presently, there are over 20.000 sequence records, and over 9000 bacterial species represented in the database. The initial gene set comprises rDNA16S, rpoB, gyrB. The system allows sequence-driven microbe identification as well as the development of study protocols for research on specific microbe species. Nucleotide sequences are represented graphically. The GenEnv database was designed as a tool for public health operators but also offers wide prospects for scientific research.

Computational Biology↗

LIGAND: chemical database for enzyme reactions.

MOTIVATION: The existing molecular biology databases focus on the sequence and structural aspects of biological macromolecules, i.e. DNAs, RNAs and proteins. However, in order to understand the functional aspects, it is essential to computerize the interaction of these molecules. Furthermore, living cells contain additional molecules, such as metabolic compounds and metal ions, that may also be considered as parts of the basic building blocks of life, but are not well organized in public databases. LIGAND chemical database is our attempt to solve these problems, at least for enzymatic reactions. RESULTS: LIGAND consists of two sections: ENZYME and COMPOUND. The ENZYME section is an extension of previous studies (Suyama et al. , Comput. Applic. Biosci., 9, 9-15, 1993), and it is a flat-file representation of 3303 enzymes and 2976 enzymatic reactions in the chemical equation format that can be parsed by machine. The COMPOUND section has been newly constructed for information on the nomenclature and chemical structures of compounds. It contains 5383 chemical compounds. Both ENZYME and COMPOUND entries contain rich cross-reference information, most of which is automatically generated by the DBGET/LinkDB system, thus providing the linkage between chemical and biological databases. LIGAND is updated daily, tightly coupled with the KEGG metabolic pathway database, and forms the basis for reconstruction and computation of pathways. AVAILABILITY: LIGAND can be accessed through the DBGET/LinkDB and KEGG systems in the Japanese GenomeNet database service via http://www.genome.ad.jp/. The flat-file format of the LIGAND database can be downloaded by anonymous FTP via ftp://kegg. genome.adjp/molecules/ligand/. CONTACT: goto@kuicr.kyoto-u.ac.jp; nishioka@scl.kyoto-u.ac.jp; kanehisa@kuicr.kyoto-u.ac.jp

Computational Biology↗

Integrated access to genomic and other bioinformation: an essential ingredient of the drug discovery process.

Due to the high rate of data production and the need of researchers to have rapid access to new data, public databases have become the major medium through which genome mapping and sequencing data as well as macromolecular structural data are published. There are now more than 250 databases of biomolecular, structural, genetic, or phenotypic data, many of which are doubling in size annually. These databases, many of which were created and are maintained by experimentalists for their own research use, provide valuable collections of organized, validated data. However, the very number and diversity of databases now make efficient data resource discovery as important as effective data resource use. Existing autonomous biological databases contain related data which are more valuable when interconnected than when isolated. Political and scientific realities dictate that these databases will be built by different teams, in different locations, for different purposes, and using different data models and supporting DBMSs. As a consequence, connecting the related data they contain is not straightforward. Experience with existing biological databases indicates that it is possible to form useful queries across these databases, but that doing so usually requires expertise in the semantic structure of each source database. Advancing to the next level of integration among biological information resources poses significant technical and sociological challenges.

Computational Biology↗

A summarization approach for Affymetrix GeneChip data using a reference training set from a large, biologically diverse database.

BACKGROUND: Many of the most popular pre-processing methods for Affymetrix expression arrays, such as RMA, gcRMA, and PLIER, simultaneously analyze data across a set of predetermined arrays to improve precision of the final measures of expression. One problem associated with these algorithms is that expression measurements for a particular sample are highly dependent on the set of samples used for normalization and results obtained by normalization with a different set may not be comparable. A related problem is that an organization producing and/or storing large amounts of data in a sequential fashion will need to either re-run the pre-processing algorithm every time an array is added or store them in batches that are pre-processed together. Furthermore, pre-processing of large numbers of arrays requires loading all the feature-level data into memory which is a difficult task even with modern computers. We utilize a scheme that produces all the information necessary for pre-processing using a very large training set that can be used for summarization of samples outside of the training set. All subsequent pre-processing tasks can be done on an individual array basis. We demonstrate the utility of this approach by defining a new version of the Robust Multi-chip Averaging (RMA) algorithm which we refer to as refRMA. RESULTS: We assess performance based on multiple sets of samples processed over HG U133A Affymetrix GeneChip arrays. We show that the refRMA workflow, when used in conjunction with a large, biologically diverse training set, results in the same general characteristics as that of RMA in its classic form when comparing overall data structure, sample-to-sample correlation, and variation. Further, we demonstrate that the refRMA workflow and reference set can be robustly applied to naïve organ types and to benchmark data where its performance indicates respectable results. CONCLUSION: Our results indicate that a biologically diverse reference database can be used to train a model for estimating probe set intensities of exclusive test sets, while retaining the overall characteristics of the base algorithm. Although the results we present are specific for RMA, similar versions of other multi-array normalization and summarization schemes can be developed.

Algorithms↗

Computational immunology: The coming of age.

The explosive growth in biotechnology combined with major advances in information technology has the potential to radically transform immunology in the postgenomics era. Not only do we now have ready access to vast quantities of existing data, but new data with relevance to immunology are being accumulated at an exponential rate. Resources for computational immunology include biological databases and methods for data extraction, comparison, analysis and interpretation. Publicly accessible biological databases of relevance to immunologists number in the hundreds and are growing daily. The ability to efficiently extract and analyse information from these databases is vital for efficient immunology research. Most importantly, a new generation of computational immunology tools enables modelling of peptide transport by the transporter associated with antigen processing (TAP), modelling of antibody binding sites, identification of allergenic motifs and modelling of T-cell receptor serial triggering.

Antigens↗

Automated structure extraction and XML conversion of life science database flat files.

In the light of the increasing number of biological databases, their integration is a fundamental prerequisite for answering complex biological questions. Database integration, therefore, is an important area of research in bioinformatics. Since most of the publicly available life science databases are still exclusively exchanged by means of proprietary flat files, database integration requires parsers for very different flat file formats. Unfortunately, the development and maintenance of database specific flat file parsers is a nontrivial and time-consuming task, which takes considerable effort in large-scale integration scenarios. This paper introduces heuristically based concepts for automatic structure extraction from life science database flat files. On the basis of these concepts the FlatEx prototype is developed for the automatic conversion of flat files into XML representations.

Algorithms↗