PubMed Health⌕ Search

Biomedical subjects

A Kel

Publications and source records attributed to A Kel.

7 recordsLinked to original sources

Composite Module Analyst: identification of transcription factor binding site combinations using genetic algorithm.

Composite Module Analyst (CMA) is a novel software tool aiming to identify promoter-enhancer models based on the composition of transcription factor (TF) binding sites and their pairs. CMA is closely interconnected with the TRANSFAC database. In particular, CMA uses the positional weight matrix (PWM) library collected in TRANSFAC and therefore provides the possibility to search for a large variety of different TF binding sites. We model the structure of the long gene regulatory regions by a Boolean function that joins several local modules, each consisting of co-localized TF binding sites. Having as an input a set of co-regulated genes, CMA builds the promoter model and optimizes the parameters of the model automatically by applying a genetic-regression algorithm. We use a multicomponent fitness function of the algorithm which includes several statistical criteria in a weighted linear function. We show examples of successful application of CMA to a microarray data on transcription profiling of TNF-alpha stimulated primary human endothelial cells. The CMA web server is freely accessible at http://www.gene-regulation.com/pub/programs/cma/CMA.html. An advanced version of CMA is also a part of the commercial system ExPlaintrade mark (www.biobase.de) designed for causal analysis of gene expression data.

Algorithms↗

Composite Module Analyst: a fitness-based tool for identification of transcription factor binding site combinations.

MOTIVATION: Functionally related genes involved in the same molecular-genetic, biochemical or physiological process are often regulated coordinately. Such regulation is provided by precisely organized binding of a multiplicity of special proteins [transcription factors (TFs)] to their target sites (cis-elements) in regulatory regions of genes. Cis-element combinations provide a structural basis for the generation of unique patterns of gene expression. RESULTS: Here we present a new approach for defining promoter models based on the composition of TF binding sites and their pairs. We utilize a multicomponent fitness function for selection of the promoter model that fits best to the observed gene expression profile. We demonstrate examples of successful application of the fitness function with the help of a genetic algorithm for the analysis of functionally related or co-expressed genes as well as testing on simulated and permutated data. AVAILABILITY: The CMA program is freely available for non-commercial users. URL http://www.gene-regulation.com/pub/programs.html#CMAnalyst. It is also a part of the commercial system ExPlain (www.biobase.de) designed for causal analysis of gene expression data..

Algorithms↗

Recognition of multiple patterns in unaligned sets of sequences: comparison of kernel clustering method with other methods.

MOTIVATION: Transcription factor binding sites often differ significantly in their primary sequence and can hardly be aligned. Often one set of sites can contain several subsets of sequences that follow not just one but several different patterns. There is a need for sensitive methods to reveal multiple patterns in unaligned sets of sequences. RESULTS: We developed a novel method for analysis of unaligned sets of sequences based on kernel estimation. The method is able to reveal 'multiple local patterns'-a set of weight matrices. Every weight matrix characterizes a pattern that can be found in a significant subset of sequences under analysis. The method developed has been compared with several other methods of pattern discovery such as Gibbs sampling, MEME, CONSENSUS, MULTIPROFILER and PROJECTION. The kernel method showed the best performance in terms of how close the revealed weight matrices are to the original ones. We applied the kernel method to analyze three samples of promoters (cell-cycle, T-cells and muscle-specific). We compared the multiple patterns revealed with the TRANSFAC library of weight matrices and found a strong similarity to several weight matrices for transcription factors known to be involved in the mentioned specific gene regulation. AVAILABILITY: The program is available for on-line use at: http://www.biobase.de/cgi-bin/biobase/cbs2/bin/template.cgi?template=cbscall.html

Algorithms↗

Whole genome human/mouse phylogenetic footprinting of potential transcription regulatory signals.

UNLABELLED: Phylogenetic footprinting is an efficient approach for revealing potential transcription factor binding sites in promoter sequences. The idea is based on an assumption that functional sites in promoters should evolve much slower then other regions that do not bear any conservative function. Therefore, potential transcription factor (TF) binding sites that are found in the evolutionally conservative regions of promoters have more chances to be considered as "real" sites. The most difficult step of the phylogenetic footprinting is alignment of promoter sequences between different organisms (fe. human and mouse). The conventional alignment methods often can not align promoters due to the high level of sequence variability. We have developed a new alignment method that takes into account similarity in distribution of potential binding sites (motif-based alignment). This method has been used effectively for promoter alignment and for revealing new potential binding sites for various transcription factors. We made a systematic phylogenetic footprinting of human/mouse conserved non-coding sequences (CNS). 60 thousand potential binding sites were revealed in human and mouse genomes. We have developed a database of the predicted potential TF binding sites. AVAILABILITY: http://compel.bionet.nsc.ru/FunSite/footprint/; www.gene-regulation.com/.

Algorithms↗

Recognition of NFATp/AP-1 composite elements within genes induced upon the activation of immune cells.

Composite elements are regulatory modules of promoters or enhancers that consist of binding sites of two different but synergizing transcription factors. A well-studied example is nuclear factors of activated T-cell (NFAT) sites which are composite elements of a NFATp/c and an activating protein 1 (AP-1) binding site. We have developed a computational approach to identify potential NFAT target genes which (a) comprises an improved method to scan for individual NFAT composite elements; (b) considers positional effects relative to transcription start sites; and (c) involves cluster analysis of potential NFAT composite elements. All three steps progressively helpX?ed to discriminate T-cell-specific promoter sequences against other functional regions (coding and intronic sequences) of the same genes, against promoters of muscle-specific genes or against random sequences. Using this approach, we identified potential NFAT composite elements in promoters of cytokine genes and their receptors as well as in promoters of genes for AP-1 family members, Ca2+-binding proteins and some other components of the regulatory network operating in activated T-cells and other immune cells. The method developed can be adapted to characterize and identify other composite elements as well. The program for recognition NFAT composite elements is available through the World Wide Web (http://compel.bionet.nsc.ru/FunSite/CompelScan. html and http://transfac.gbf.de/dbsearch/funsitep/s _comp.html).

Animals↗

A genetic algorithm for designing gene family-specific oligonucleotide sets used for hybridization: the G protein-coupled receptor protein superfamily.

MOTIVATION: Massive oligonucleotide hybridization is one of the most promising technologies of functional genome analysis. The critical point is to design appropriate sets of oligonucleotides that can be used effectively in identification by hybridization. RESULTS: Using a genetic algorithm approach, we have attempted to design sets of oligo probes capable of identifying new genes belonging to a defined gene family within a cDNA or genomic library. It is not limited by oligonucleotide length and admits the letter 'N' in the structure of the oligonucleotides selected. One of the major advantages of this approach is the low homology required to identify functional families of sequences with little homology. We have designed the oligonucleotide sets that are most selective for the cDNA clones of transmembrane G protein-coupled receptors (GPCRs), a large family of proteins that form part of a modular system of extracellular signal transduction to the intracellular second messenger pathways. The accuracy of identification has been checked on the EST library containing 713 870 cDNA sequences. A set of 15 oligos between 7 and 14 bases in length has correctly identified 70% of the GPCR cDNA collection sequences with 0.02% false positives. AVAILABILITY: The developed software is available by ftp://ftp.bionet.nsc. ru/pub/biology/ and on the Web page http://www.bionet.nsc. ru/SRCG/Oligoselector/. CONTACT: kel@.bionet.nsc.ru; sebastian. meier-ewert@gpc-ag.com

Algorithms↗