PubMed Health⌕ Search

Biomedical subjects

Georg Fuellen

Publications and source records attributed to Georg Fuellen.

12 recordsLinked to original sources

Simplifying gene trees for easier comprehension.

BACKGROUND: In the genomic age, gene trees may contain large amounts of data making them hard to read and understand. Therefore, an automated simplification is important. RESULTS: We present a simplification tool for gene trees called TreeSimplifier. Based on species tree information and HUGO gene names, it summarizes "monophyla". These monophyla correspond to subtrees of the gene tree where the evolution of a gene follows species phylogeny, and they are simplified to single leaves in the gene tree. Such a simplification may fail, for example, due to genes in the gene tree that are misplaced. In this way, misplaced genes can be identified. Optionally, our tool glosses over a limited degree of "paraphyly" in a further simplification step. In both simplification steps, species can be summarized into groups and treated as equivalent. In the present study we used our tool to derive a simplified tree of 397 leaves from a tree of 1138 leaves. Comparing the simplified tree to a "cartoon tree" created manually, we note that both agree to a high degree. CONCLUSION: Our automatic simplification tool for gene trees is fast, accurate, and effective. It yields results of similar quality as manual simplification. It should be valuable in phylogenetic studies of large protein families. The software is available at http://www.uni-muenster.de/Bioinformatics/services/treesim/.

Algorithms↗

IsoSVM--distinguishing isoforms and paralogs on the protein level.

BACKGROUND: Recent progress in cDNA and EST sequencing is yielding a deluge of sequence data. Like database search results and proteome databases, this data gives rise to inferred protein sequences without ready access to the underlying genomic data. Analysis of this information (e.g. for EST clustering or phylogenetic reconstruction from proteome data) is hampered because it is not known if two protein sequences are isoforms (splice variants) or not (i.e. paralogs/orthologs). However, even without knowing the intron/exon structure, visual analysis of the pattern of similarity across the alignment of the two protein sequences is usually helpful since paralogs and orthologs feature substitutions with respect to each other, as opposed to isoforms, which do not. RESULTS: The IsoSVM tool introduces an automated approach to identifying isoforms on the protein level using a support vector machine (SVM) classifier. Based on three specific features used as input of the SVM classifier, it is possible to automatically identify isoforms with little effort and with an accuracy of more than 97%. We show that the SVM is superior to a radial basis function network and to a linear classifier. As an example application we use IsoSVM to estimate that a set of Xenopus laevis EST clusters consists of approximately 81% cases where sequences are each other's paralogs and 19% cases where sequences are each other's isoforms. The number of isoforms and paralogs in this allotetraploid species is of interest in the study of evolution. CONCLUSION: We developed an SVM classifier that can be used to distinguish isoforms from paralogs with high accuracy and without access to the genomic data. It can be used to analyze, for example, EST data and database search results. Our software is freely available on the Web, under the name IsoSVM.

Algorithms↗

Biotool2Web: creating simple Web interfaces for bioinformatics applications.

UNLABELLED: Currently there are many bioinformatics applications being developed, but there is no easy way to publish them on the World Wide Web. We have developed a Perl script, called Biotool2Web, which makes the task of creating web interfaces for simple ('home-made') bioinformatics applications quick and easy. Biotool2Web uses an XML document containing the parameters to run the tool on the Web, and generates the corresponding HTML and common gateway interface (CGI) files ready to be published on a web server. AVAILABILITY: This tool is available for download at URL http://www.uni-muenster.de/Bioinformatics/services/biotool2web/ CONTACT: Georg Fuellen (fuellen@alum.mit.edu).

Computational Biology↗

Correspondence of function and phylogeny of ABC proteins based on an automated analysis of 20 model protein data sets.

Using our BLAST-based procedure RiPE (Retrieval-induced Phylogeny Environment), which automates the evolutionary analysis of a protein family, we assembled a set of 1138 ABC protein components [adenosine triphosphate (ATP)-binding cassette and transmembrane domain] from the protein data sets of 20 model organisms and subjected them to phylogenetic and functional analysis. For maximum speed, we based the alignment directly on a homology search with a profile of all known human ABC proteins and used neighbor-joining tree estimation. All but 11 sequences from Homo sapiens, Arabidopsis thaliana, Drosophila melanogaster, and Saccharomyces cerevisiae were placed into the correct subtree/subfamily, reproducing published classifications of the individual organisms. By following a simple "function transfer rule", our comparative phylogenetic analysis successfully predicted the known function of human ABC proteins in 19 of 22 cases. Three functional predictions did not correspond, and 10 were novel. Predictions based on BLAST alone were inferior in five cases and superior in two. Bacterial sequences were placed close to the root of most subtrees. This placement coincides with domain architecture, suggesting an early diversification of the ABC family before the kingdoms split apart. Our approach can, in principle, be used to annotate any protein family of any organism included in the study.

ATP-Binding Cassette Transporters↗

Cloning, cellular localization, genomic organization, and tissue-specific expression of the TGFbeta1-inducible SMAP-5 gene.

SMAP-5 is a member of the five-pass transmembrane protein family localizing in the Golgi apparatus and the endoplasmic reticulum. These proteins have been implicated in intracellular trafficking, in secretion and in vesicular transport. Phylogenetic analyses revealed that SMAP-5 is a member of a small Rab GTPase interacting factor protein family. The human SMAP-5 gene spans about 12.5 kb and comprises 6 exons on chromosomal locus 5q32. The proximal 5'-flanking region of the gene lacks a TATA box and is highly GC rich. Consistent with this, the SMAP-5 gene is expressed in all tissues. The highest level of expression was found in coronary smooth muscle cells, in which expression of the SMAP-5 gene was induced by transforming growth factor beta1, thus indicating that this protein may play an important role in inflammation.

Alternative Splicing↗

Comparative homology agreement search: an effective combination of homology-search methods.

Many methods have been developed to search for homologous members of a protein family in databases, and the reliability of results and conclusions may be compromised if only one method is used, neglecting the others. Here we introduce a general scheme for combining such methods. Based on this scheme, we implemented a tool called comparative homology agreement search (chase) that integrates different search strategies to obtain a combined "E value." Our results show that a consensus method integrating distinct strategies easily outperforms any of its component algorithms. More specifically, an evaluation based on the Structural Classification of Proteins database reveals that, on average, a coverage of 47% can be obtained in searches for distantly related homologues (i.e., members of the same superfamily but not the same family, which is a very difficult task), accepting only 10 false positives, whereas the individual methods obtain a coverage of 28-38%.

Databases, Protein↗

VisCoSe: visualization and comparison of consensus sequences.

We introduce visualization and comparison of consensus sequences (VisCoSe) as a WWW service and a stand-alone command line Perl script for visualizing and comparing consensus sequences of protein and nucleotide sequences. VisCoSe is the only interface available that simultaneously calculates consensus sequences of multiple data sets and automatically compares these consensus sequences. Furthermore, VisCoSe allows visualization of chemical properties of amino acids.

Algorithms↗

Computational searches for missing orthologs: the case of S100A12 in mice.

The interaction of the Ca2+-binding protein S100A12 with RAGE (receptor of advanced glycation endproducts) has been considered as a novel proinflammatory axis, since blockage of RAGE/S100A12 ligation suppresses chronic cellular activation and tissue injury in mouse models. However, the existence of a murine S100A12 ortholog is unknown. Because experimental approaches failed to identify it, we started an analysis of gene locus evolution. Human S100A12 is localized in the S100 gene cluster between S100A8 and S100A9, which are neighbors in both mouse and human. Confirming identical gene order, we found a DNA region between the murine S100A8 and S100A9 genes that is 60.9% identical to a region of the human S100A12 gene, including the first exon. Instead of the second and third exon, we found homology to a region close to the human S100A9 locus. To exclude a murine S100A12 ortholog elsewhere in the genome, we used human S100A12 as query for TBlastN homology searches. The matches were either too short, or identity was too low, or they could clearly be identified as distinct S100 genes. Obviously, an S100A12 ortholog is neither present in mouse nor rat, indicating that S100A12 has been lost during rodent evolution, probably due to a deletion.

Algorithms↗

BLASTing proteomes, yielding phylogenies.

We develop a procedure called RiPE (Retrieval-induced Phylogeny Environment) that automatically performs an evolutionary analysis of a protein (sub)family, (i) by retrieving the relevant sequences via a homology search, (ii) by using the search report to construct the alignment using only homologous subsequences (taking into account their neighborhood with a low chance of homology), (iii) by realigning, and (iv) by generating phylogenetic trees based on the alignment. In a first implementation of our scheme, we start with the available proteome data of model organisms, perform a PSI-BLAST search, use MView to convert hits into a multiple alignment, and perform realignment and tree building. As a test case, we have investigated the human ABC transporters of the subfamily G, starting with the five known human ABCG transporters. Our method retrieved homologous sequences not previously analyzed, generating a tree that is more plausible and better supported than previously published trees. The RiPE 0.1 prototype is available at the RiPE website, http://ifg-izkf.uni-muenster.de/fuellen/RiPE/ripe.html.

ATP Binding Cassette Transporter, Subfamily G, Mem↗

Paper2sequences: retrieval of sequences listed in a publication.

Our web-based tool simplifies the often laborious procedure of retrieving a set of biosequences in a publication or webpage. As a front-end to the Bioperl toolkit, it accepts as an input a list of identifiers. They are specified in an ASCII table (copy-pasted from the publication's PDF or HTML page) and give rise to queries in multiple databases for the protein/nucleic acid data specified. Currently, GenBank, PIR (Protein Information Resource) and Swiss-Prot are supported. For any sequence accession code listed, the database can be specified and, if retrieval fails, automatic lookup for the same code in other databases can be requested. Sequence length information (if specified) and heuristic rules are used to drive the lookup if multiple protein coding sequences (CDS) are part of a single accession. Warnings are issued in cases of ambiguities and inconsistencies. An advanced option enables the user to format the output in whatever format they wish.

Amino Acid Sequence↗

The Bioperl toolkit: Perl modules for the life sciences.

The Bioperl project is an international open-source collaboration of biologists, bioinformaticians, and computer scientists that has evolved over the past 7 yr into the most comprehensive library of Perl modules available for managing and manipulating life-science information. Bioperl provides an easy-to-use, stable, and consistent programming interface for bioinformatics application programmers. The Bioperl modules have been successfully and repeatedly used to reduce otherwise complex tasks to only a few lines of code. The Bioperl object model has been proven to be flexible enough to support enterprise-level applications such as EnsEMBL, while maintaining an easy learning curve for novice Perl programmers. Bioperl is capable of executing analyses and processing results from programs such as BLAST, ClustalW, or the EMBOSS suite. Interoperation with modules written in Python and Java is supported through the evolving BioCORBA bridge. Bioperl provides access to data stores such as GenBank and SwissProt via a flexible series of sequence input/output modules, and to the emerging common sequence data storage format of the Open Bioinformatics Database Access project. This study describes the overall architecture of the toolkit, the problem domains that it addresses, and gives specific examples of how the toolkit can be used to solve common life-sciences problems. We conclude with a discussion of how the open-source nature of the project has contributed to the development effort.

Algorithms↗