PubMed Health⌕ Search

Biomedical subjects

J van Helden

Publications and source records attributed to J van Helden.

10 recordsLinked to original sources

From molecular activities and processes to biological function.

This paper describes how biological function can be represented in terms of molecular activities and processes. It presents several key features of a data model that is based on a conceptual description of the network of interactions between molecular entities within the cell and between cells. This model is implemented in the aMAZE database that presently deals with information on metabolic pathways, gene regulation, sub- or supracellular locations, and transport. It is shown that this model constitutes a useful generalisation of data representations currently implemented in metabolic pathway databases, and that it can furthermore include multiple schemes for categorising and classifying molecular entities, activities, processes and localisations. In particular, we highlight the flexibility offered by our system in representing multiple molecular activities and their control, in viewing biological function at different levels of resolution and in updating this view as our knowledge evolves.

Animals↗

Discovering regulatory elements in non-coding sequences by analysis of spaced dyads.

The application of microarray and related technologies is currently generating a systematic catalog of the transcriptional response of any single gene to a multiplicity of experimental conditions. Clustering genes according to the similarity of their transcriptional response provides a direct hint to the regulons of the different transcription factors, many of which have still not been characterized. We have developed a new method for deciphering the mechanism underlying the common transcriptional response of a set of genes, i.e. discovering cis -acting regulatory elements from a set of unaligned upstream sequences. This method, called dyad analysis, is based on the observation that many regulatory sites consist of a pair of highly conserved trinucleotides, spaced by a non-conserved region of fixed width. The approach is to count the number of occurrences of each possible spaced pair of trinucleotides, and to assess its statistical significance. The method is highly efficient in the detection of sites bound by C(6)Zn(2)binuclear cluster proteins, as well as other transcription factors. In addition, we show that the dyad and single-word analyses are efficient for the detection of regulatory patterns in gene clusters from DNA chip experiments. In combination, these programs should provide a fast and efficient way to discover new regulatory sites for as yet unknown transcription factors.

Base Sequence↗

Statistical analysis of yeast genomic downstream sequences reveals putative polyadenylation signals.

The study of a few genes has permitted the identification of three elements that constitute a yeast polyadenyl-ation signal: the efficiency element (EE), the positioning element and the actual site for cleavage and poly-adenyl-ation. In this paper we perform an analysis of oligonucleotide composition on the sequences located downstream of the stop codon of all yeast genes. Several oligonucleotide families appear over-represented with a high significance (referred to herein as 'words'). The family with the highest over-representation includes the oligonucleotides shown experimentally to play a role as EEs. The word with the highest score is TATATA, followed, among others, by a series of single-nucleotide variants (TATGTA, TACATA, TAAATA.) and one-letter shifts (ATATAT). A position analysis reveals that those words have a high preference to be in 3' flanks of yeast genes and there they have a very uneven distribution, with a marked peak around 35 bp after the stop codon. Of the predicted ORFs, 85% show one or more of those sequences. Similar results were obtained using a data set of EST sequences. Other clusters of over-represented words are also detected, namely T- and A-rich signals. Using these results and previously known data we propose a general model for the 3' trailers of yeast mRNAs.

Base Sequence↗

A web site for the computational analysis of yeast regulatory sequences.

A series of computer programs were developed for the analysis of regulatory sequences, with a special focus on yeast. These tools are publicly available on the web (http://copan.cifn.unam. mx/Computational_Biology/yeast-tools or http://www.ucmb.ulb.ac. be/bioinformatics/rsa-tools/). Basically, three classical problems can be addressed: (a) search for known regulatory patterns in the upstream regions of known genes; (b) discovery of unknown regulatory patterns within a set of upstream regions known to be co-regulated; (c) search for unknown genes potentially regulated by a known transcription factor. Each of these tasks can be performed on basis of a simple (string) or more refined (matrix) description of the regulatory patterns. A feature-map program automatically generates visual representations of the positions at which patterns were found. The site also provides a series of general utilities, such as generation of random sequence, automatic drawing of XY graphs, interconversions between sequence formats, etc. Several tools are linked together to allow their sequential utilization (piping), but each one can also be used independently by filling the web form with external data. This widens the scope of the site to the analysis of non-regulatory and/or non-yeast sequences.

Computational Biology↗

Interactive visualization and exploration of relationships between biological objects.

Genome sequencing and microarray technology produce ever-increasing amounts of complex data that need analysis. Visualization is an effective analytical technique that exploits the ability of the human brain to process large amounts of data. Here, we review traditional visualization methods based on clustering and tree representation, and also describe an alternative approach that involves projecting objects onto a Euclidean space in a way that reflects their structural or functional distances. Data are visualized without preclustering and can be dynamically explored by the user using 'virtual-reality'. We illustrate this approach with two case studies from protein topology and gene expression.

Biometry↗

RegulonDB (version 2.0): a database on transcriptional regulation in Escherichia coli.

RegulonDB version 2.0, a database on transcriptional regulation and operon organization in Escherichia coli, is now available on the web at the following URL: http://www.cifn.unam. mx/Computational_Biology/regulondb/. In this paper we describe the main computational changes to the database, which include migrating the database to Sybase, providing graphical descriptions of the internal organization of operons and regulons, and direct links to MEDLINE references. The web interface offers searching either by mechanisms of regulation or by operon organization. The results of a search (operon organization, or site collection) are displayed as hypertext, and can also be displayed graphically. In terms of its contents, RegulonDB contains a large number of operons, as well as the absolute position in the completed genome sequence of sites, promoters, and individual genes of E.coli.

Databases, Factual↗

Extracting regulatory sites from the upstream region of yeast genes by computational analysis of oligonucleotide frequencies.

We present here a simple and fast method allowing the isolation of DNA binding sites for transcription factors from families of coregulated genes, with results illustrated in Saccharomyces cerevisiae. Although conceptually simple, the algorithm proved efficient for extracting, from most of the yeast regulatory families analyzed, the upstream regulatory sequences which had been previously found by experimental analysis. Furthermore, putative new regulatory sites are predicted within upstream regions of several regulons. The method is based on the detection of over-represented oligonucleotides. A specificity of this approach is to define the statistical significance of a site based on tables of oligonucleotide frequencies observed in all non-coding sequences from the yeast genome. In contrast with heuristic methods, this oligonucleotide analysis is rigorous and exhaustive. Its range of detection is however limited to relatively simple patterns: short motifs with a highly conserved core. These features seem to be shared by a good number of regulatory sites in yeast. This, and similar methods, should be increasingly required to identify unknown regulatory elements within the numerous new coregulated families resulting from measurements of gene expression levels at the genomic scale. All tools described here are available on the web at the site http://copan.cifn.unam.mx/Computational_Biology/ yeast-tools

Algorithms↗

The iroquois complex controls the somatotopy of Drosophila notum mechanosensory projections.

Sensory neurons can establish topologically ordered projections in the central nervous system, thereby building an internal representation of the external world. We analyze how this ordering is genetically controlled in Drosophila, using as a model system the neurons that innervate the mechanosensory bristles on the back of the fly (the notum). Sensory neurons innervating the medially located bristles send an axonal branch that crosses the central nervous system midline, defining a 'medial' identity, while the ones that innervate the lateral bristles send no such branch, defining a 'lateral' identity. We analyze the role of the proneural genes achaete and scute, which are involved in the formation of the medial and lateral bristles, and we show that they have no effect on the 'medial' and 'lateral' identities of the neurons. We also analyze the role of the prepattern genes araucan and caupolican, two members of the iroquois gene complex which are required for the expression of achaete and scute in the lateral region of the notum, and we show that their expression is responsible for the 'lateral' identity of the projection.

Animals↗

Representing and analysing molecular and cellular function using the computer.

Determining the biological function of a myriad of genes, and understanding how they interact to yield a living cell, is the major challenge of the post genome-sequencing era. The complexity of biological systems is such that this cannot be envisaged without the help of powerful computer systems capable of representing and analysing the intricate networks of physical and functional interactions between the different cellular components. In this review we try to provide the reader with an appreciation of where we stand in this regard. We discuss some of the inherent problems in describing the different facets of biological function, give an overview of how information on function is currently represented in the major biological databases, and describe different systems for organising and categorising the functions of gene products. In a second part, we present a new general data model, currently under development, which describes information on molecular function and cellular processes in a rigorous manner. The model is capable of representing a large variety of biochemical processes, including metabolic pathways, regulation of gene expression and signal transduction. It also incorporates taxonomies for categorising molecular entities, interactions and processes, and it offers means of viewing the information at different levels of resolution, and dealing with incomplete knowledge. The data model has been implemented in the database on protein function and cellular processes 'aMAZE' (http://www.ebi.ac.uk/research/pfbp/), which presently covers metabolic pathways and their regulation. Several tools for querying, displaying, and performing analyses on such pathways are briefly described in order to illustrate the practical applications enabled by the model.

Cell Physiological Phenomena↗