PubMed Health⌕ Search

Biomedical subjects

Edgar Wingender

Publications and source records attributed to Edgar Wingender.

At least 19 recordsLinked to original sources

Beyond microarrays: find key transcription factors controlling signal transduction pathways.

BACKGROUND: Massive gene expression changes in different cellular states measured by microarrays, in fact, reflect just an "echo" of real molecular processes in the cells. Transcription factors constitute a class of the regulatory molecules that typically require posttranscriptional modifications or ligand binding in order to exert their function. Therefore, such important functional changes of transcription factors are not directly visible in the microarray experiments. RESULTS: We developed a novel approach to find key transcription factors that may explain concerted expression changes of specific components of the signal transduction network. The approach aims at revealing evidence of positive feedback loops in the signal transduction circuits through activation of pathway-specific transcription factors. We demonstrate that promoters of genes encoding components of many known signal transduction pathways are enriched by binding sites of those transcription factors that are endpoints of the considered pathways. Application of the approach to the microarray gene expression data on TNF-alpha stimulated primary human endothelial cells helped to reveal novel key transcription factors potentially involved in the regulation of the signal transduction pathways of the cells. CONCLUSION: We developed a novel computational approach for revealing key transcription factors by knowledge-based analysis of gene expression data with the help of databases on gene regulatory networks (TRANSFAC and TRANSPATH. The corresponding software and databases are available at http://www.gene-regulation.com.

Binding Sites↗

TRANSPATH: an information resource for storing and visualizing signaling pathways and their pathological aberrations.

TRANSPATH is a database about signal transduction events. It provides information about signaling molecules, their reactions and the pathways these reactions constitute. The representation of signaling molecules is organized in a number of orthogonal hierarchies reflecting the classification of the molecules, their species-specific or generic features, and their post-translational modifications. Reactions are similarly hierarchically organized in a three-layer architecture, differentiating between reactions that are evidenced by individual publications, generalizations of these reactions to construct species-independent 'reference pathways' and the 'semantic projections' of these pathways. A number of search and browse options allow easy access to the database contents, which can be visualized with the tool PathwayBuildertrade mark. The module PathoSign adds data about pathologically relevant mutations in signaling components, including their genotypes and phenotypes. TRANSPATH and PathoSign can be used as encyclopaedia, in the educational process, for vizualization and modeling of signal transduction networks and for the analysis of gene expression data. TRANSPATH Public 6.0 is freely accessible for users from non-profit organizations under http://www.gene-regulation.com/pub/databases.html.

Computer Graphics↗

TiProD: the Tissue-specific Promoter Database.

TiProD is a database of human promoter sequences for which some functional features are known. It allows a user to query individual promoters and the expression pattern they mediate, gene expression signatures of individual tissues, and to retrieve sets of promoters according to their tissue-specific activity or according to individual Gene Ontology terms the corresponding genes are assigned to. We have defined a measure for tissue-specificity that allows the user to discriminate between ubiquitously and specifically expressed genes. The database is accessible at http://tiprod.cbi.pku.edu.cn:8080/index.html.

Databases, Nucleic Acid↗

EndoNet: an information resource about endocrine networks.

EndoNet is a new database that provides information about the components of endocrine networks and their relations. It focuses on the endocrine cell-to-cell communication and enables the analysis of intercellular regulatory pathways in humans. In the EndoNet data model, two classes of components span a bipartite directed graph. One class represents the hormones (in the broadest sense) secreted by defined donor cells. The other class consists of the acceptor or target cells expressing the corresponding hormone receptors. The identity and anatomical environment of cell types, tissues and organs is defined through references to the CYTOMER ontology. With the EndoNet user interface, it is possible to query the database for hormones, receptors or tissues and to combine several items from different search rounds in one complex result set, from which a network can be reconstructed and visualized. For each entity, a detailed characteristics page is available. Some well-established endocrine pathways are offered as showcases in the form of predefined result sets. These sets can be used as a starting point for a more complex query or for obtaining a quick overview. The EndoNet database is accessible at http://endonet.bioinf.med.uni-goettingen.de/.

Cell Communication↗

Evaluating phylogenetic footprinting for human-rodent comparisons.

MOTIVATION: 'Phylogenetic footprinting' is a widely applied approach to identify regulatory regions and potential transcription factor binding sites (TFBSs) using alignments of non-coding orthologous regions from two or more organisms. A systematic evaluation of its validity and usability based on known TFBSs is needed to use phylogenetic footprinting most effectively in the identification of unknown TFBSs. RESULTS: In this paper we use 2678 human, mouse and rat TFBSs from the TRANSFAC database for this evaluation. To ensure the retrieval of correct orthologous sequences, we combine gene annotation and sequence homology searches. Demanding a sequence identity of at least 65% is most effective in discriminating TFBSs from non-functional sequence parts, while different alignment algorithms only have a minor influence on TFBS identification by human-rodent comparisons. With this threshold approximately 72% of the known TFBSs are found conserved, a number which varies significantly between different transcription factors and also depends on the function of the regulated gene. TFBSs for certain transcription factors do not require strict sequence conservation but instead may show a high pattern conservation, limiting somewhat the validity of purely sequence-based phylogenetic footprinting.

Animals↗

Construction of predictive promoter models on the example of antibacterial response of human epithelial cells.

BACKGROUND: Binding of a bacteria to a eukaryotic cell triggers a complex network of interactions in and between both cells. P. aeruginosa is a pathogen that causes acute and chronic lung infections by interacting with the pulmonary epithelial cells. We use this example for examining the ways of triggering the response of the eukaryotic cell(s), leading us to a better understanding of the details of the inflammatory process in general. RESULTS: Considering a set of genes co-expressed during the antibacterial response of human lung epithelial cells, we constructed a promoter model for the search of additional target genes potentially involved in the same cell response. The model construction is based on the consideration of pair-wise combinations of transcription factor binding sites (TFBS). It has been shown that the antibacterial response of human epithelial cells is triggered by at least two distinct pathways. We therefore supposed that there are two subsets of promoters activated by each of them. Optimally, they should be "complementary" in the sense of appearing in complementary subsets of the (+)-training set. We developed the concept of complementary pairs, i.e., two mutually exclusive pairs of TFBS, each of which should be found in one of the two complementary subsets. CONCLUSIONS: We suggest a simple, but exhaustive method for searching for TFBS pairs which characterize the whole (+)-training set, as well as for complementary pairs. Applying this method, we came up with a promoter model of antibacterial response genes that consists of one TFBS pair which should be found in the whole training set and four complementary pairs. We applied this model to screening of 13,000 upstream regions of human genes and identified 430 new target genes which are potentially involved in antibacterial defense mechanisms.

Algorithms↗

Deriving an ontology for human gene expression sources from the CYTOMER database on human organs and cell types.

CYTOMER is a relational database of organs/tissues, cell types, physiological systems and developmental stages that currently focuses on the human system. From this database, we have derived an ontology for anatomical and morphological structures for the human organism which includes all embryonal stages and the cell types constituting these structures. The ontology has been transferred to the OWL format and is freely available for download at http://cytomer/bioinf.med.uni-goettingen.de.

Animals↗

Topology of mammalian transcription networks.

We present a first attempt to evaluate the generic topological principles underlying the mammalian transcriptional regulatory networks. Transcription networks, TN, studied here are represented as graphs where vertices are genes coding for transcription factors and edges are causal links between the genes, each edge combining both gene expression and trans-regulation events. Two transcription networks were retrieved from the TRANSPATH database: The first one, TN_RN, is a 'complete' transcription network referred to as a reference network. The second one, TN_p53, displays a particular transcriptional sub-network centered at p53 gene. We found these networks to be fundamentally non-random and inhomogeneous. Their topology follows a power-law degree distribution and is best described by the scale-free model. Shortest-path-length distribution and the average clustering coefficient indicate a small-world feature of these networks. The networks show the dependence of the clustering coefficient on the degree of a vertex, thereby indicating the presence of hierarchical modularity. Clear positive correlation between the values of betweenness and the degree of vertices has been observed in both networks. The top list of genes displaying high degree and high betweennes, such as p53, c-fos, c-jun and c-myc, is enriched with genes that are known as having tumor-suppressor or proto-oncogene properties, which supports the biological significance of the identified key topological elements.

Animals↗

A novel computational approach for the prediction of networked transcription factors of aryl hydrocarbon-receptor-regulated genes.

A novel computational method based on a genetic algorithm was developed to study composite structure of promoters of coexpressed genes. Our method enabled an identification of combinations of multiple transcription factor binding sites regulating the concerted expression of genes. In this article, we study genes whose expression is regulated by a ligand-activated transcription factor, aryl hydrocarbon receptor (AhR), that mediates responses to a variety of toxins. AhR-mediated change in expression of AhR target genes was measured by oligonucleotide microarrays and by reverse transcription-polymerase chain reaction in human and rat hepatocytes. Promoters and long-distance regulatory regions (>10 kb) of AhR-responsive genes were analyzed by the genetic algorithm and a variety of other computational methods. Rules were established on the local oligonucleotide context in the flanks of the AhR binding sites, on the occurrence of clusters of AhR recognition elements, and on the presence in the promoters of specific combinations of multiple binding sites for the transcription factors cooperating in the AhR regulatory network. Our rules were applied to search for yet unknown Ah-receptor target genes. Experimental evidence is presented to demonstrate high fidelity of this novel in silico approach.

Algorithms↗

TRANSFAC, TRANSPATH and CYTOMER as starting points for an ontology of regulatory networks.

Building an ontology of a defined knowledge domain can help to model an appropriate database structure for the relevant contents. On the other hand, having a comprehensive overview of the knowledge of a certain domain as it may be provided by corresponding databases facilitates building an appropriate ontology, or ontologies with different granularities, which may then provide many additional benefits in handling the stored and in retrieving additional information from heterogenous sources. In this communication, the first steps are reported how we may derive an ontology for the domain of "molecular regulation" from our databases TRANSFAC (transcriptional regulation), TRANSPATH (signal transduction) and CYTOMER (cellular locations).

Cells↗

Consistent re-modeling of signaling pathways and its implementation in the TRANSPATH database.

The data model of the signaling pathways database TRANSPATH has been re-engineered to a three-layer model comprising experimental evidences and summarized pathway information, both in a mechanistically detailed manner, and a "semantic" projection for the abstract overview. Each molecule is described in the context of a certain reaction in the multidimensional space of posttranslational modification, molecular family relationships, and the biological species of its origin. The new model makes the data better suitable for reconstructing signaling pathways and networks and mapping expression data, for instance from microarray experiments, onto regulatory networks.

Algorithms↗

Systematic DNA-binding domain classification of transcription factors.

Based on the manual annotation of transcription factors stored in the TRANSFAC database, we developed a library of hidden Markov models (HMM) to represent their DNA-binding domains and used it for a comprehensive classification. The models constructed were applied on the UniProt/Swiss-Prot database, leading to a systematic classification of further DNA-binding protein entries. The HMM library obtained can be used to classify any newly discovered transcription factor according to its DNA-binding domain and, thus, to generate hypotheses about its DNA-binding specificity.

Binding Sites↗

Composition-sensitive analysis of the human genome for regulatory signals.

Known transcription regulatory signals which generally act as transcription factor binding sites (TFs) differ significantly in their base composition. Therefore, their occurrence in a genome largely depends on the local base composition. In an attempt to initiate an all human genome analysis for the occurrence of potential TFs, we systematically analyzed the GC-content of distinct functional regions (e. g., upstream and downstream gene regions, exons, long and short introns, repetitive elements) and correlated the frequencies of potential binding sites of a representative set of TFs in these regions. For these analyses, we used the pattern collection of the TRANSFAC database on transcriptional regulation, the information about functionally relevant combinations of them from the database TRANSCompel, and our new resource, TRANSGenomeTM, which provides an overall annotation of the human genome with emphasis on its regulatory characteristics. We show that the occurrence of sequence patterns with regulatory potential may be supported by, but cannot be fully explained by either the GC content of a whole chromosome or its putative promoter regions, nor by the information content of the patterns. Several patterns, HNF-3, NFAT, and GC box, show a clear overrepresentation in all promoter groups as well as in all chromosomes. Other patterns, like E2F and CRE-BP1, are underrepresented in all promoter groups as well as in all chromosomes in comparison with random sequences. Simultaneously, both patterns are over-represented in promoters in comparison with repetitive elements. We define several structural characteristics of the proximal promoters that differentiate them from other functional genomic regions. Two well-known promoter elements, GC- and TATA-boxes, are statistically enriched in promoters in comparison with random sequences, repetitive elements and exons. Altogether, our findings provide insights into the macroheterogeneity amongst the individual chromosomes, into the microheterogeneity among different functional regions of individual chromosomes, contribute to further understanding of structural organization of gene regulatory regions, and give first hints on the development of regulatory features during evolution.

Animals↗

PRODORIC: prokaryotic database of gene regulation.

The database PRODORIC aims to systematically organize information on prokaryotic gene expression, and to integrate this information into regulatory networks. The present version focuses on pathogenic bacteria such as Pseudomonas aeruginosa. PRODORIC links data on environmental stimuli with trans-acting transcription factors, cis-acting promoter elements and regulon definition. Interactive graphical representations of operon, gene and promoter structures including regulator-binding sites, transcriptional and translational start sites, supplemented with information on regulatory proteins are available at varying levels of detail. The data collection provided is based on exhaustive analyses of scientific literature and computational sequence prediction. Included within PRODORIC are tools to define and predict regulator binding sites. It is accessible at http://prodoric.tu-bs.de.

Bacteria↗

TRANSPATH: an integrated database on signal transduction and a tool for array analysis.

TRANSPATH is a database system about gene regulatory networks that combines encyclopedic information on signal transduction with tools for visualization and analysis. The integration with TRANSFAC, a database about transcription factors and their DNA binding sites, provides the possibility to obtain complete signaling pathways from ligand to target genes and their products, which may themselves be involved in regulatory action. As of July 2002, the TRANSPATH Professional release 3.2 contains about 9800 molecules, >1800 genes and >11 400 reactions collected from approximately 5000 references. With the ArrayAnalyzer, an integrated tool has been developed for evaluation of microarray data. It uses the TRANSPATH data set to identify key regulators in pathways connected with up- or down-regulated genes of the respective array. The key molecules and their surrounding networks can be viewed with the PathwayBuilder, a tool that offers four different modes of visualization. More information on TRANSPATH is available at http://www.biobase.de/pages/products/databases.html.

Animals↗

Prediction of potential C/EBP/NF-kappaB composite elements using matrix-based search methods.

Bacterial infections trigger a wide range of host cell responses. For the interaction of Pseudomonas aeruginosa and epithelial cells it is known that transcription factor NF-kappaB plays a central role, but its effects have to be specified by cooperation with additional factors. NF-B containing composite elements, e. g. with C/EBP, may be appropriate indicators for new antibacterial response genes. We refined matrix-based search methods for C/EBP, which was necessary because of weak consensi of the previously existing C/EBP matrices, established a model for C/EBP/ NF-kappaB composite element, used it for scanning all known human 5'-flanking sequences and identified 135 new candidate genes. The newly constructed C/EBP binding patterns will be available with one of the next releases of the TRANSFAC database (http://www.gene-regulation.de).

Base Sequence↗

TRANSCompel: a database on composite regulatory elements in eukaryotic genes.

Originating from COMPEL, the TRANSCompel database emphasizes the key role of specific interactions between transcription factors binding to their target sites providing specific features of gene regulation in a particular cellular content. Composite regulatory elements contain two closely situated binding sites for distinct transcription factors and represent minimal functional units providing combinatorial transcriptional regulation. Both specific factor--DNA and factor--factor interactions contribute to the function of composite elements (CEs). Information about the structure of known CEs and specific gene regulation achieved through such CEs appears to be extremely useful for promoter prediction, for gene function prediction and for applied gene engineering as well. Each database entry corresponds to an individual CE within a particular gene and contains information about two binding sites, two corresponding transcription factors and experiments confirming cooperative action between transcription factors. The COMPEL database, equipped with the search and browse tools, is available at http://www.gene-regulation.com/pub/databases.html#transcompel. Moreover, we have developed the program CATCH for searching potential CEs in DNA sequences. It is freely available as CompelPatternSearch at http://compel.bionet.nsc.ru/FunSite/CompelPatternSearch.html.

Animals↗