PubMed Health⌕ Search

Biomedical subjects

Kathleen Marchal

Publications and source records attributed to Kathleen Marchal.

10 recordsLinked to original sources

INCLUSive: A web portal and service registry for microarray and regulatory sequence analysis.

INCLUSive is a suite of algorithms and tools for the analysis of gene expression data and the discovery of cis-regulatory sequence elements. The tools allow normalization, filtering and clustering of microarray data, functional scoring of gene clusters, sequence retrieval, and detection of known and unknown regulatory elements using probabilistic sequence models and Gibbs sampling. All tools are available via different web pages and as web services. The web pages are connected and integrated to reflect a methodology and facilitate complex analysis using different tools. The web services can be invoked using standard SOAP messaging. Example clients are available for download to invoke the services from a remote computer or to be integrated with other applications. All services are catalogued and described in a web service registry. The INCLUSive web portal is available for academic purposes at http://www.esat.kuleuven.ac.be/inclusive.

Algorithms↗

MARAN: normalizing micro-array data.

SUMMARY: MARAN is a web-based application for normalizing microarray data. MARAN comprises a generic ANOVA model, an option for Loess fitting prior to ANOVA analysis, and a module for selecting genes with significantly changing expression. AVAILABILITY: http://www.esat.kuleuven.ac.be/maran/.

Algorithms↗

Genome-specific higher-order background models to improve motif detection.

Motif detection based on Gibbs sampling is a common procedure used to retrieve regulatory motifs in silico. Using a species-specific background model was previously shown to increase the robustness of the algorithm. Here, we demonstrate that selecting a non-species-adapted background model can have an adverse effect on the results of motif detection. The large differences in the average nucleotide composition of prokaryotic sequences exacerbate the problem of exchanging background models. Therefore, we have developed complex background models for all prokaryotic species with available genome sequences.

Algorithms↗

Computational approaches to identify promoters and cis-regulatory elements in plant genomes.

The identification of promoters and their regulatory elements is one of the major challenges in bioinformatics and integrates comparative, structural, and functional genomics. Many different approaches have been developed to detect conserved motifs in a set of genes that are either coregulated or orthologous. However, although recent approaches seem promising, in general, unambiguous identification of regulatory elements is not straightforward. The delineation of promoters is even harder, due to its complex nature, and in silico promoter prediction is still in its infancy. Here, we review the different approaches that have been developed for identifying promoters and their regulatory elements. We discuss the detection of cis-acting regulatory elements using word-counting or probabilistic methods (so-called "search by signal" methods) and the delineation of promoters by considering both sequence content and structural features ("search by content" methods). As an example of search by content, we explored in greater detail the association of promoters with CpG islands. However, due to differences in sequence content, the parameters used to detect CpG islands in humans and other vertebrates cannot be used for plants. Therefore, a preliminary attempt was made to define parameters that could possibly define CpG and CpNpG islands in Arabidopsis, by exploring the compositional landscape around the transcriptional start site. To this end, a data set of more than 5,000 gene sequences was built, including the promoter region, the 5'-untranslated region, and the first introns and coding exons. Preliminary analysis shows that promoter location based on the detection of potential CpG/CpNpG islands in the Arabidopsis genome is not straightforward. Nevertheless, because the landscape of CpG/CpNpG islands differs considerably between promoters and introns on the one side and exons (whether coding or not) on the other, more sophisticated approaches can probably be developed for the successful detection of "putative" CpG and CpNpG islands in plants.

Computational Biology↗

Prediction and overview of the RpoN-regulon in closely related species of the Rhizobiales.

BACKGROUND: In the rhizobia, a group of symbiotic Gram-negative soil bacteria, RpoN (sigma54, sigmaN, NtrA) is best known as the sigma factor enabling transcription of the nitrogen fixation genes. Recent reports, however, demonstrate the involvement of RpoN in other symbiotic functions, although no large-scale effort has yet been undertaken to unravel the RpoN-regulon in rhizobia. We screened two complete rhizobial genomes (Mesorhizobium loti, Sinorhizobium meliloti) and four symbiotic regions (Rhizobium etli, Rhizobium sp. NGR234, Bradyrhizobium japonicum, M. loti) for the presence of the highly conserved RpoN-binding sites. A comparison was also made with two closely related non-symbiotic members of the Rhizobiales (Agrobacterium tumefaciens, Brucella melitensis). RESULTS: A highly specific weight-matrix-based screening method was applied to predict members of the RpoN-regulon, which were stored in a highly annotated and manually curated dataset. Possible enhancer-binding proteins (EBPs) controlling the expression of RpoN-dependent genes were predicted with a profile hidden Markov model. CONCLUSIONS: The methodology used to predict RpoN-binding sites proved highly effective as nearly all known RpoN-controlled genes were identified. In addition, many new RpoN-dependent functions were found. The dependency of several of these diverse functions on RpoN seems species-specific. Around 30% of the identified genes are hypothetical. Rhizobia appear to have recruited RpoN for symbiotic processes, whereas the role of RpoN in A. tumefaciens and B. melitensis remains largely to be elucidated. All species screened possess at least one uncharacterized EBP as well as the usual ones. Lastly, RpoN could significantly broaden its working range by direct interfering with the binding of regulatory proteins to the promoter DNA.

Bacterial Proteins↗

Sensitivity function-based model reduction: A bacterial gene expression case study.

Mathematical models used to predict the behavior of genetically modified organisms require 1). a (rather) large number of state variables, and 2). complicated kinetic expressions containing a large number of parameters. Since these models are hardly identifiable and of limited use in model-based optimization and control strategies, a generic methodology based on sensitivity function analysis is presented to reduce the model complexity at the level of the kinetics, while maintaining high prediction power. As a case study to illustrate the method and results obtained, the influence of the dissolved oxygen concentration on the cytN gene expression in the bacterium Azospirillum brasilense Sp7 is modeled. As a first modeling approach, available mechanistic knowledge is incorporated into a mass balance equation model with 3 states and 14 parameters. The large differences in order of magnitude of the model parameters identified on the available experimental data indicate 1). possible structural problems in the kinetic model and, associated with this, 2). a possibly too high number of model parameters. A careful sensitivity function analysis reveals that a reduced model with only seven parameters is almost as accurate as the original model.

Azospirillum brasilense↗

PlantCARE, a database of plant cis-acting regulatory elements and a portal to tools for in silico analysis of promoter sequences.

PlantCARE is a database of plant cis-acting regulatory elements, enhancers and repressors. Regulatory elements are represented by positional matrices, consensus sequences and individual sites on particular promoter sequences. Links to the EMBL, TRANSFAC and MEDLINE databases are provided when available. Data about the transcription sites are extracted mainly from the literature, supplemented with an increasing number of in silico predicted data. Apart from a general description for specific transcription factor sites, levels of confidence for the experimental evidence, functional information and the position on the promoter are given as well. New features have been implemented to search for plant cis-acting regulatory elements in a query sequence. Furthermore, links are now provided to a new clustering and motif search method to investigate clusters of co-expressed genes. New regulatory elements can be sent automatically and will be added to the database after curation. The PlantCARE relational database is available via the World Wide Web at http://sphinx.rug.ac.be:8080/PlantCARE/.

Consensus Sequence↗

A Gibbs sampling method to detect overrepresented motifs in the upstream regions of coexpressed genes.

Microarray experiments can reveal important information about transcriptional regulation. In our case, we look for potential promoter regulatory elements in the upstream region of coexpressed genes. Here we present two modifications of the original Gibbs sampling algorithm for motif finding (Lawrence et al., 1993). First, we introduce the use of a probability distribution to estimate the number of copies of the motif in a sequence. Second, we describe the technical aspects of the incorporation of a higher-order background model whose application we discussed in Thijs et al. (2001). Our implementation is referred to as the Motif Sampler. We successfully validate our algorithm on several data sets. First, we show results for three sets of upstream sequences containing known motifs: 1) the G-box light-response element in plants, 2) elements involved in methionine response in Saccharomyces cerevisiae, and 3) the FNR O(2)-responsive element in bacteria. We use these data sets to explain the influence of the parameters on the performance of our algorithm. Second, we show results for upstream sequences from four clusters of coexpressed genes identified in a microarray experiment on wounding in Arabidopsis thaliana. Several motifs could be matched to regulatory elements from plant defence pathways in our database of plant cis-acting regulatory elements (PlantCARE). Some other strong motifs do not have corresponding motifs in PlantCARE but are promising candidates for further analysis.

Algorithms↗

INCLUSive: integrated clustering, upstream sequence retrieval and motif sampling.

INCLUSive allows automatic multistep analysis of microarray data (clustering and motif finding). The clustering algorithm (adaptive quality-based clustering) groups together genes with highly similar expression profiles. The upstream sequences of the genes belonging to a cluster are automatically retrieved from GenBank and can be fed directly into Motif Sampler, a Gibbs sampling algorithm that retrieves statistically over-represented motifs in sets of sequences, in this case upstream regions of co-expressed genes.

Algorithms↗

Adaptive quality-based clustering of gene expression profiles.

MOTIVATION: Microarray experiments generate a considerable amount of data, which analyzed properly help us gain a huge amount of biologically relevant information about the global cellular behaviour. Clustering (grouping genes with similar expression profiles) is one of the first steps in data analysis of high-throughput expression measurements. A number of clustering algorithms have proved useful to make sense of such data. These classical algorithms, though useful, suffer from several drawbacks (e.g. they require the predefinition of arbitrary parameters like the number of clusters; they force every gene into a cluster despite a low correlation with other cluster members). In the following we describe a novel adaptive quality-based clustering algorithm that tackles some of these drawbacks. RESULTS: We propose a heuristic iterative two-step algorithm: First, we find in the high-dimensional representation of the data a sphere where the "density" of expression profiles is locally maximal (based on a preliminary estimate of the radius of the cluster-quality-based approach). In a second step, we derive an optimal radius of the cluster (adaptive approach) so that only the significantly coexpressed genes are included in the cluster. This estimation is achieved by fitting a model to the data using an EM-algorithm. By inferring the radius from the data itself, the biologist is freed from finding an optimal value for this radius by trial-and-error. The computational complexity of this method is approximately linear in the number of gene expression profiles in the data set. Finally, our method is successfully validated using existing data sets. AVAILABILITY: http://www.esat.kuleuven.ac.be/~thijs/Work/Clustering.html

Algorithms↗