PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Biological databases”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Using bioinformatics to identify kidney genes.

Online knowledge of genes involved in mouse kidney development is available from the literature (PubMed) and text-based (Kidney Development and Mouse Gene Expression-GXD) databases. Further information is in gene databases of other organisms having tissues homologous to the metanephros (e.g. Drosophila Malpighian tubules; tools are available for identifying mouse homologues of non-mouse genes) and from tissues homologous to those in the developing kidney (e.g. mesothelium that, like nephrons, arises from a mesenchyme to epithelial transition). Future databases will also include graphical data, and this knowledge will provide a further level of insight. These databases and tools are discussed here.

Animals↗

MolTalk--a programming library for protein structures and structure analysis.

BACKGROUND: Two of the mostly unsolved but increasingly urgent problems for modern biologists are a) to quickly and easily analyse protein structures and b) to comprehensively mine the wealth of information, which is distributed along with the 3D co-ordinates by the Protein Data Bank (PDB). Tools which address this issue need to be highly flexible and powerful but at the same time must be freely available and easy to learn. RESULTS: We present MolTalk, an elaborate programming language, which consists of the programming library libmoltalk implemented in Objective-C and the Smalltalk-based interpreter MolTalk. MolTalk combines the advantages of an easy to learn and programmable procedural scripting with the flexibility and power of a full programming language. An overview of currently available applications of MolTalk is given and with PDBChainSaw one such application is described in more detail. PDBChainSaw is a MolTalk-based parser and information extraction utility of PDB files. Weekly updates of the PDB are synchronised with PDBChainSaw and are available for free download from the MolTalk project page http://www.moltalk.org following the link to PDBChainSaw. For each chain in a protein structure, PDBChainSaw extracts the sequence from its co-ordinates and provides additional information from the PDB-file header section, such as scientific organism, compound name, and EC code. CONCLUSION: MolTalk provides a rich set of methods to analyse and even modify experimentally determined or modelled protein structures. These methods vary in complexity and are thus suitable for beginners and advanced programmers alike. We envision MolTalk to be most valuable in the following applications:1) To analyse protein structures repetitively in large-scale, i.e. to benchmark protein structure prediction methods or to evaluate structural models. The quality of the resulting 3D-models can be assessed by e.g. calculating a Ramachandran-Sasisekharan plot.2) To quickly retrieve information for (a limited number of) macro-molecular structures, i.e. H-bonds, salt bridges, contacts between amino acids and ligands or at the interface between two chains.3) To programme more complex structural bioinformatics software and to implement demanding algorithms through its portability to Objective-C, e.g. iMolTalk.4) To be used as a front end to databases, e.g. PDBChainSaw.

Artificial Intelligence↗

A semiautomated approach to gene discovery through expressed sequence tag data mining: discovery of new human transporter genes.

Identification and functional characterization of the genes in the human genome remain a major challenge. A principal source of publicly available information used for this purpose is the National Center for Biotechnology Information database of expressed sequence tags (dbEST), which contains over 4 million human ESTs. To extract the information buried in this data more effectively, we have developed a semiautomated method to mine dbEST for uncharacterized human genes. Starting with a single protein input sequence, a family of related proteins from all species is compiled. This entire family is then used to mine the human EST database for new gene candidates. Evaluation of putative new gene candidates in the context of a family of characterized proteins provides a framework for inference of the structure and function of the new genes. When applied to a test data set of 28 families within the major facilitator superfamily (MFS) of membrane transporters, our protocol found 73 previously characterized human MFS genes and 43 new MFS gene candidates. Development of this approach provided insights into the problems and pitfalls of automated data mining using public databases.

Biological Transport, Active↗

Peptide mass fingerprinting: identification of proteins by MALDI-TOF.

MALDI-TOF peptide mass fingerprinting (PMF) is the fastest and cheapest method of protein identification; the studied genome is sequenced and annotated, and the protein is amenable to separation and detection in 2D gel electrophoresis. In plant proteomics there are two main difficulties: few plant genomes are sequenced, and major contaminants are non-plant specific. This chapter describes the classical "bottom-up" method (i.e., from peptide to protein identification) of gel cutting, in-gel digestion, peptide recovery and purification, MALDI-TOF mass spectrometry, and critical survey of protein database queries.

Acrylic Resins↗

A bioinformatics approach to investigating developmental pathways in the kidney and other tissues.

Over the past few years, large amounts of data linking gene-expression (GE) patterns and other genetic data with the development of the mouse kidney have been published, and the next task will be to integrate these data with the molecular networks responsible for the emergence of the kidney phenotype. This paper discusses how a start to this task can be made by using the kidney database and its associated search tools, and shows how the data generated by such an approach can be used as a guide to future experimentation. Many of the events taking place as the kidney develops do, of course, also take place in other tissues and organisms and it will soon be possible to incorporate relevant information from these systems into analyses of kidney data as well as the new information from microarray technology. The key to success here will be the ability to access over the internet data from the textual and graphical databases for the mouse and other organisms now being established. In order to do this, informatic tools will be needed that will allow a user working with one database to query another. This paper also considers both the types of tools that will be necessary and the databases on which they will operate.

Animals↗

[Methods of statistical genetics and use of database for genome information].

Knowledge and technology of bioinformatics have become inevitable for gene and genome research. Education and research in this field of science are not sufficient in Japan. There are two different approaches to trait mapping, the way by which traits are mapped on the genome. Thus, the knowledge-based approach uses functions of molecules while the statistics-based approach uses polymorphisms. Statistics-based approach uses two different methods, linkage analysis and analysis based on linkage disequilibrium. Various phenotypes are efficiently mapped on the genome using such methods. Recently, bioinformatic data base search is mostly performed using internet. Anyone can perform sequence-search, homology-search and SNP-search. Since such data bases change quickly, readers should access the databases themselves and be used to the procedures for them.

Computational Biology↗

The Portable Dictionary of the Mouse Genome: a personal database for gene mapping and molecular biology.

The Portable Dictionary of the Mouse Genome is a database for personal computers that contains information on approximately 10,000 loci in the mouse, along with data on homologs in several other mammalian species, including human, rat, cat, cow, and pig. Key features of the dictionary are its compact size, its network independence, and the ability to convert the entire dictionary to a wide variety of common application programs. Another significant feature is the integration of DNA sequence accession data. Loci in the dictionary can be rapidly resorted by chromosomal position, by type, by human homology, or by gene effect. The dictionary provides an accessible, easily manipulated set of data that has many uses--from a quick review of loci and gene nomenclature to the design of experiments and analysis of results. The Portable Dictionary is available in several formats suitable for conversion to different programs and computer systems. It can be obtained on disk or from Internet Gopher servers (mickey.utmen.edu or anat4.utmen.edu), an anonymous FTP site (nb.utmem.edu in the directory pub/genedict), and a World Wide Web server (http://mickey.utmem.edu/front.html).

Animals↗

N-terminal N-myristoylation of proteins: prediction of substrate proteins from amino acid sequence.

Myristoylation by the myristoyl-CoA:protein N-myristoyltransferase (NMT) is an important lipid anchor modification of eukaryotic and viral proteins. Automated prediction of N-terminal N-myristoylation from the substrate protein sequence alone is necessary for large-scale sequence annotation projects but it requires a low rate of false positive hits in addition to a sufficient sensitivity. Our previous analysis of substrate protein sequence variability, NMT sequences and 3D structures has revealed motif properties in addition to the known PROSITE motif that are utilized in a new predictor described here. The composite prediction function (with separate ad hoc parameterization (a) for queries from non-fungal eukaryotes and their viruses and (b) for sequences from fungal species) consists of terms evaluating amino acid type preferences at sequences positions close to the N terminus as well as terms penalizing deviations from the physical property pattern of amino acid side-chains encoded in multi-residue correlation within the motif sequence. The algorithm has been validated with a self-consistency and two jack-knife tests for the learning set as well as with kinetic data for model substrates. The sensitivity in recognizing documented NMT substrates is above 95 % for both taxon-specific versions. The corresponding rate of false positive prediction (for sequences with an N-terminal glycine residue) is close to 0.5 %; thus, the technique is applicable for large-scale automated sequence database annotation. The predictor is available as public WWW-server with the URL http://mendel.imp.univie.ac.at/myristate/. Additionally, we propose a version of the predictor that identifies a number of proteolytic protein processing sites at internal glycine residues and that evaluates possible N-terminal myristoylation of the protein fragments.A scan of public protein databases revealed new potential NMT targets for which the myristoyl modification may be of critical importance for biological function. Among others, the list includes kinases, phosphatases, proteasomal regulatory subunit 4, kinase interacting proteins KIP1/KIP2, protozoan flagellar proteins, homologues of mitochondrial translocase TOM40, of the neuronal calcium sensor NCS-1 and of the cytochrome c-type heme lyase CCHL. Analyses of complete eukaryote genomes indicate that about 0.5 % of all encoded proteins are apparent NMT substrates except for a higher fraction in Arabidopsis thaliana ( approximately 0.8 %).

Acyltransferases↗

Functional proteomics mapping of a human signaling pathway.

Access to the human genome facilitates extensive functional proteomics studies. Here, we present an integrated approach combining large-scale protein interaction mapping, exploration of the interaction network, and cellular functional assays performed on newly identified proteins involved in a human signaling pathway. As a proof of principle, we studied the Smad signaling system, which is regulated by members of the transforming growth factor beta (TGFbeta) superfamily. We used two-hybrid screening to map Smad signaling protein-protein interactions and to establish a network of 755 interactions, involving 591 proteins, 179 of which were poorly or not annotated. The exploration of such complex interaction databases is improved by the use of PIMRider, a dedicated navigation tool accessible through the Web. The biological meaning of this network is illustrated by the presence of 18 known Smad-associated proteins. Functional assays performed in mammalian cells including siRNA knock-down experiments identified eight novel proteins involved in Smad signaling, thus validating this integrated functional proteomics approach.

Adaptor Proteins, Signal Transducing↗

Identification of endothelial cell genes by combined database mining and microarray analysis.

Vascular endothelial cells maintain the interface between the systemic circulation and soft tissues and mediate critical processes such as inflammation in a vascular bed-selective fashion. To expand our understanding of the genetic pathways that underlie these specific functions, we have focused on the identification of novel genes that are differentially expressed in all endothelial cells, as well as restricted groups of this cell type. Virtual subtraction was conducted employing gene expression data deposited in public databases and 384 genes identified. These genes were spotted on custom microarrays, along with 288 genes identified through subtraction cloning from TGF-beta-stimulated endothelial cells. Arrays were evaluated with RNA samples representing endothelial cells cultured from four vascular sources and five non-endothelial cell types. These studies identified 64 pan-endothelial markers that were differentially expressed with at least a threefold difference (range 3- to 55-fold). In addition, differences in gene expression profiles among endothelial cells from different vascular beds were identified. Validation of these findings was performed by RNA blot expression studies, and a number of the novel genes were shown to be expressed under angiogenic conditions in the developing mouse embryo. The combined tools of database mining and transcriptional profiling thus provide expanded knowledge of endothelial cell gene expression and endothelial cell biology.

Adult↗

Pathway mapping tools for analysis of high content data.

The complexity of human biology requires a systems approach that uses computational approaches to integrate different data types. Systems biology encompasses the complete biological system of metabolic and signaling pathways, which can be assessed by measuring global gene expression, protein content, metabolic profiles, and individual genetic, clinical, and phenotypic data. High content screening assays can also be used to generate systems biology knowledge. In this review, we will summarize the pathway databases and describe biological network tools used predominantly with this genomics, proteomics, and metabolomics data but which are equally as applicable for high content screening data analysis. We describe in detail the integrated data-mining tools applicable to building biological networks developed by GeneGo, namely, MetaCore and MetaDrug.

Computational Biology↗

GPCR antitarget modeling: pharmacophore models for biogenic amine binding GPCRs to avoid GPCR-mediated side effects.

G protein-coupled receptors (GPCRs) form a large protein family that plays an important role in many physiological and pathophysiological processes. However, the central role that the biogenic amine binding GPCRs and their ligands play in cell signaling poses a risk in new drug candidates that reveal side affinities towards these receptor sites. These candidates have the potential to interfere with the physiological signaling processes and to cause undesired effects in preclinical or clinical studies. Here, we present 3D cross-chemotype pharmacophore models for three biogenic amine antitargets: the alpha(1A) adrenergic, the 5-HT(2A) serotonin, and the D2 dopamine receptors. These pharmacophores describe the key chemical features present within these biogenic amine antagonists and rationalize the biogenic amine side affinities found for numerous new drug candidates. First applications of the alpha(1A) adrenergic receptor model reveal that these in silico tools can be used to guide the chemical optimization towards development candidates with fewer alpha(1A)-mediated side effects (for example, orthostatic hypotension) and, thus, with an improved clinical safety profile.

Adrenergic Antagonists↗

Rough set-based proteochemometrics modeling of G-protein-coupled receptor-ligand interactions.

G-Protein-coupled receptors (GPCRs) are among the most important drug targets. Because of a shortage of 3D crystal structures, most of the drug design for GPCRs has been ligand-based. We propose a novel, rough set-based proteochemometric approach to the study of receptor and ligand recognition. The approach is validated on three datasets containing GPCRs. In proteochemometrics, properties of receptors and ligands are used in conjunction and modeled to predict binding affinity. The rough set (RS) rule-based models presented herein consist of minimal decision rules that associate properties of receptors and ligands with high or low binding affinity. The information provided by the rules is then used to develop a mechanistic interpretation of interactions between the ligands and receptors included in the datasets. The first two datasets contained descriptors of melanocortin receptors and peptide ligands. The third set contained descriptors of adrenergic receptors and ligands. All the rule models induced from these datasets have a high predictive quality. An example of a decision rule is "If R1_ligand(Ethyl) and TM helix 2 position 27(Methionine) then Binding(High)." The easily interpretable rule sets are able to identify determinative receptor and ligand parts. For instance, all three models suggest that transmembrane helix 2 is determinative for high and low binding affinity. RS models show that it is possible to use rule-based models to predict ligand-binding affinities. The models may be used to gain a deeper biological understanding of the combinatorial nature of receptor-ligand interactions.

Algorithms↗

Computational and experimental analysis identifies many novel human genes.

Because of advances in automation, human genomic sequences are being deposited in public databases at a dramatic rate. However, the process of detecting genes in these sequences is still something of an art. Here we describe the implementation and testing of a relatively straightforward computational approach, the Virtual Transcribed Sequence project, which analyzes their gene content using the gene prediction program GENSCAN (GENSCAN 1.0 1,2) in combination with similarity-based methods. This approach identifies many novel human genes not found even in EST databases.

Amino Acid Sequence↗

Biological information: making it accessible and integrated (and trying to make sense of it).

The availability of the genome sequences of human and mouse, human sequence variation data and other large genetic data sets will lead to a revolution in understanding of the human machine and the treatment of its diseases. The success of the international genome sequencing consortiums shows what can be achieved by well coordinated large scale public domain projects and the benefits of data access to all. It is already clear that the availability of this sequence is having a huge impact on research worldwide. Complete genome sequences provide a framework to pull all biological data together such that each piece has the potential to say something about biology as a whole. Biology is too complex for any organisation to have a monopoly of ideas or data, so the collection, analysis and access to this data can be contributed to by research institutes around the world. However, although it is possible for all this data to be accessible to all through the internet, the more organisations provide data or analysis separately, the harder it becomes for anyone to collect and integrate the results. To address these problems of intergration of data, open standards for biological data exchange, such as the 'Distributed Annotation System' (DAS) are being developed and bioinformatics (Dowell et al., 2001) as a whole is now being strongly driven by the open source software (OSS) model for collaborative software development (Hubbard and Birney, 1999). The leading provider of human genome annotation, the Ensembl project (http://www.ensembl.org), is entirely an OSS project and has been widely adopted by academic and commerical organisations alike (Hubbard et al., 2002). Accurate automatic annotation of features such as genes in vertebrate genomes currently relies on supporting evidence in the form of homologies to mRNAs, ESTs or protein. However, it appears that sufficient high quality experimentally curated annotation now exists to be used as a substrate for machine learning algorithms to create effective models of biological signal sequences (Down and Hubbard, 2002). Is there hope for ab initio prediction methods after all?

Chromosome Mapping↗