PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Model-driven user interfaces for bioinformatics data resources: regenerating the wheel as an alternative to reinventing it.

BACKGROUND: The proliferation of data repositories in bioinformatics has resulted in the development of numerous interfaces that allow scientists to browse, search and analyse the data that they contain. Interfaces typically support repository access by means of web pages, but other means are also used, such as desktop applications and command line tools. Interfaces often duplicate functionality amongst each other, and this implies that associated development activities are repeated in different laboratories. Interfaces developed by public laboratories are often created with limited developer resources. In such environments, reducing the time spent on creating user interfaces allows for a better deployment of resources for specialised tasks, such as data integration or analysis. Laboratories maintaining data resources are challenged to reconcile requirements for software that is reliable, functional and flexible with limitations on software development resources. RESULTS: This paper proposes a model-driven approach for the partial generation of user interfaces for searching and browsing bioinformatics data repositories. Inspired by the Model Driven Architecture (MDA) of the Object Management Group (OMG), we have developed a system that generates interfaces designed for use with bioinformatics resources. This approach helps laboratory domain experts decrease the amount of time they have to spend dealing with the repetitive aspects of user interface development. As a result, the amount of time they can spend on gathering requirements and helping develop specialised features increases. The resulting system is known as Pierre, and has been validated through its application to use cases in the life sciences, including the PEDRoDB proteomics database and the e-Fungi data warehouse. CONCLUSION: MDAs focus on generating software from models that describe aspects of service capabilities, and can be applied to support rapid development of repository interfaces in bioinformatics. The Pierre MDA is capable of supporting common database access requirements with a variety of auto-generated interfaces and across a variety of repositories. With Pierre, four kinds of interfaces are generated: web, stand-alone application, text-menu, and command line. The kinds of repositories with which Pierre interfaces have been used are relational, XML and object databases.

Computational Biology↗

Pathway databases.

Network representations of biological pathways offer a functional view of molecular biology that is different from and complementary to sequence, expression, and structure databases. There is currently available a wide range of digital collections of pathway data, differing in organisms included, functional area covered (e.g., metabolism vs. signaling), detail of modeling, and support for dynamic pathway construction. While it is currently impossible for these databases to communicate with each other, there are several efforts at standardizing a data exchange language for pathway data. Databases that represent pathway data at the level of individual interactions make it possible to combine data from different predefined pathways and to query by network connectivity. Computable representations of pathways provide a basis for various analyses, including detection of broad network patterns, comparison with mRNA or protein abundance, and simulation.

Computational Biology↗

PAX of mind for pathway researchers.

Scientists seeking to understand the inner workings of cells have access to a multitude of pathway data resources. However, the representations of pathway data within these resources are not consistent or interchangeable. To facilitate easy information retrieval from a wide variety of pathway resources, such as signal transduction, gene regulation, molecular interaction and metabolic pathway databases, a broad effort in the biopathways community called BioPAX was formed. New biological pathway software applications built using the BioPAX standard will be able to integrate knowledge from multiple sources in a coherent and reliable way. This article reports the progress that the BioPAX work-group has made towards building and deploying the BioPAX data-exchange format for biological pathway data.

Computational Biology↗

Assessing the impact of alternative splicing on domain interactions in the human proteome.

We have constructed a database of alternatively spliced protein forms (ASP), consisting of 13,384 protein isoform sequences of 4422 human genes (www.bioinformatics.ucla.edu/ASP). We identified fifty protein domain types that were selectively removed by alternative splicing at much higher frequencies than average (p-value < 0.01). These include many well-known protein-interaction domains (e.g., KRAB; ankyrin repeats; Kelch) including some that have been previously shown to be regulated functionally by alternative splicing (e.g., collagen domain). We present a number of novel examples (Kruppel transcription factors; Pbx2; Enc1) from the ASP database, illustrating how this pattern of alternative splicing changes the structure of a biological pathway, by redirecting protein interaction networks at key switch points. Our bioinformatics analysis indicates that a major impact of alternative splicing is removal of protein-protein interaction domains that mediate key linkages in protein interaction networks. ASP expands the available dataset of human alternatively spliced protein forms from 1989 human genes (SwissProt release 42) to 5413 (nonredundant set, ASP + SwissProt), a nearly 3-fold increase. ASP will enhance the existing pool of protein sequences that are searched by mass spectroscopy software during the identification of peptide fragments.

Alternative Splicing↗

Identifying native-like protein structures using physics-based potentials.

As the field of structural genomics matures, new methods will be required that can accurately and rapidly distinguish reliable structure predictions from those that are more dubious. We present a method based on the CHARMM gas phase implicit hydrogen force field in conjunction with a generalized Born implicit solvation term that allows one to make such discrimination. We begin by analyzing pairs of threaded structures from the EMBL database, and find that it is possible to identify the misfolded structures with over 90% accuracy. Further, we find that misfolded states are generally favored by the solvation term due to the mispairing of favorable intramolecular ionic contacts. We also examine 29 sets of 29 misfolded globin sequences from Levitt's "Decoys 'R' Us" database generated using a sequence homology-based method. Again, we find that discrimination is possible with approximately 90% accuracy. Also, even in these less distorted structures, mispairing of ionic contacts results in a more favorable solvation energy for misfolded states. This is also found to be the case for collapsed, partially folded conformations of CspA and protein G taken from folding free energy calculations. We also find that the inclusion of the generalized Born solvation term, in postprocess energy evaluation, improves the correlation between structural similarity and energy in the globin database. This significantly improves the reliability of the hypothesis that more energetically favorable structures are also more similar to the native conformation. Additionally, we examine seven extensive collections of misfolded structures created by Park and Levitt using a four-state reduced model also contained in the "Decoys 'R' Us" database. Results from these large databases confirm those obtained in the EMBL and misfolded globin databases concerning predictive accuracy, the energetic advantage of misfolded proteins regarding the solvation component, and the improved correlation between energy and structural similarity due to implicit solvation. Z-scores computed for these databases are improved by including the generalized Born implicit solvation term, and are found to be comparable to trained and knowledge-based scoring functions. Finally, we briefly explore the dynamic behavior of a misfolded protein relative to properly folded conformations. We demonstrate that the misfolded conformation diverges quickly from its initial structure while the properly folded states remain stable. Proteins in this study are shown to be more stable than their misfolded counterparts and readily identified based on energetic as well as dynamic criteria. In summary, we demonstrate the utility of physics-based force fields in identifying native-like conformations in a variety of preconstructed structural databases. The details of this discrimination are shown to be dependent on the construction of the structural database.

Algorithms↗

Protein fold similarity estimated by a probabilistic approach based on C(alpha)-C(alpha) distance comparison.

The distribution of the C(alpha)-C(alpha) distances between residues separated by three to 30 amino acid residues is highly characteristic of protein folds and makes it possible to identify them from a straightforward comparison of the distance histograms. The comparison is carried out by contingency table analysis and yields a probability of identity (PRIDE score), with values between zero and 1. For closely related structures, PRIDE is highly correlated with the root-mean-square distance between C(alpha) atoms, but it provides a correct classification even for unrelated structures for which a structural alignment is not meaningful. For example, an analysis of the CATH database of fold structures showed that 98.8% of the folds fall into the correct CATH homologous superfamily category, based on the highest PRIDE score obtained. Structural alignment and secondary-structure assignment are not necessary for the calculation of PRIDE, which is fast enough to allow the scanning of large databases.

Animals↗

Proteome analysis of bacterial pathogens.

Combining two-dimensional electrophoresis with mass spectrometry resulted in a powerful technology ideally suited to recognize and identify proteins of pathogenic microorganisms. This classical proteome analysis is now complemented by capillary chromatography/mass spectrometry combinations, miniaturization by chip technology and protein interaction investigations. Comparative proteomics is used to reveal vaccine candidates and pathogenicity factors. Immunoproteomics identifies specific and nonspecific antigens. For the management of the huge data amounts, bioinformatics is a valuable instrument for the construction of complex protein databases.

Animals↗

Links between kinetic data and sequences in the alpha/beta-hydrolases fold database.

While the number of sequenced genes is increasing dramatically, the number of different protein structural families is expected to be more limited. Changes in enzymatic activity or protein interactions can dramatically modify the role of homologous proteins in different organisms or mutants. However, experimental data associated with sequences or mutations stored in databases are often limited to a short description of the enzymatic pathway, molecular interaction or phenotype associated with the changes in amino acid sequence. In the alpha/beta-hydrolases fold database ESTHER, we are experimenting with links between experimental kinetic data and sequences, mutations and protein structures. This effort will lead to the integration of pharmacological data with genome-wide databases.

Animals↗

In silico criterion for prediction of effects of p53 gene missense mutations on p53-Mdm2 feedback loop.

The Informational Spectrum Method (ISM) is the tool for the in silico analysis of proteins which interprets protein sequence linear information using signal analyses methods. In this paper the ISM was employed to characterize the products of genetic variants of tumor suppressor gene p53 and its natural binding regulator protein Mdm2. Based on this we propose the criterion for identification of missense mutations that have impact on the p53-Mdm2 feedback loop. The efficiency of the proposed criterion was confirmed by the ISM analyses of p53 mutants reported in: (i) healthy individuals, (ii) germline mutations database and (iii) somatic mutations database.

Amino Acid Sequence↗

Update in bioinformatics. Toward a digital database of plant cell signalling networks: advantages, limitations and predictive aspects of the digital model.

The process of signal integration, which contributes to the regulation of multiple cellular activities, can be described in a digital language by a set of connected digital operations. In this article we delineate the basic concepts of cell signalling in the context of a logical description of information processing. Newly described instances of signal integration in plants are given as examples. The different advantages, limitations and predictive aspects of the digital modeling of signal transduction networks, as well as the minimal architecture of a computer database for plant signalling networks are discussed.

Computational Biology↗

A generalized affine gap model significantly improves protein sequence alignment accuracy.

Sequence alignment underpins common tasks in molecular biology, including genome annotation, molecular phylogenetics, and homology modeling. Fundamental to sequence alignment is the placement of gaps, which represent character insertions or deletions. We assessed the ability of a generalized affine gap cost model to reliably detect remote protein homology and to produce high-quality alignments. Generalized affine gap alignment with optimal gap parameters performed as well as the traditional affine gap model in remote homology detection. Evaluation of alignment quality showed that the generalized affine model aligns fewer residue pairs than the traditional affine model but achieves significantly higher per-residue accuracy. We conclude that generalized affine gap costs should be used when alignment accuracy carries more importance than aligned sequence length.

Algorithms↗

Discovering reliable protein interactions from high-throughput experimental data using network topology.

OBJECTIVE: Current protein-protein interaction (PPI) detection via high-throughput experimental methods, such as yeast-two-hybrid has been reported to be highly erroneous, leading to potentially costly spurious discoveries. This work introduces a novel measure called IRAP, i.e. "interaction reliability by alternative path", for assessing the reliability of protein interactions based on the underlying topology of the PPI network. METHODS AND MATERIALS: A candidate PPI is considered to be reliable if it is involved in a closed loop in which the alternative path of interactions between the two interacting proteins is strong. We devise an algorithm called AlternativePathFinder to compute the IRAP value for each interaction in a complex PPI network. Validation of the IRAP as a measure for assessing the reliability of PPIs is performed with extensive experiments on yeast PPI data. All the data used in our experiments can be downloaded from our supplementary data web site at . RESULTS: Results show consistently that IRAP measure is an effective way for discovering reliable PPIs in large datasets of error-prone experimentally-derived PPIs. Results also indicate that IRAP is better than IG2, and markedly better than the more simplistic IG1 measure. CONCLUSION: Experimental results demonstrate that a global, system-wide approach-such as IRAP that considers the entire interaction network instead of merely local neighbors-is a much more promising approach for assessing the reliability of PPIs.

Algorithms↗

Chemical biology on PINs and NeeDLes.

Systematic studies of the organization of biochemical networks that make up the living cell can be defined by studying the organization and dynamics of protein interaction networks (PINs). Here, we describe recent conceptual and experimental advances that can achieve this aim and how chemical perturbations of interactions can be used to define the organization of biochemical networks. Resulting perturbation profiles and subcellular locations of interactions allow us to 'place' each gene product at its relevant point in a network. We discuss how experimental strategies can be used in conjunction with other genome-wide analyses of physical and genetic protein interactions and gene transcription profiles to determine network dynamic linkage (NDL) in the living cell. It is through such dynamic studies that the intricate networks that make up the chemical machinery of the cell will be revealed.

Computational Biology↗