PubMed Health⌕ Search

Biomedical subjects

F Lisacek

Publications and source records attributed to F Lisacek.

7 recordsLinked to original sources

A multi-agent system simulating human splice site recognition.

The present paper describes a method detecting splice sites automatically on the basis of sequence data and models of site/signal recognition supported by experimental evidences. The method is designed to simulate splicing and while doing so, track prediction failures, missing information and possibly test correcting hypotheses. Correlations between nucleotides in the splice site regions and the various elements of the acceptor region are evaluated and combined to assess compensating interactions between elements of the splicing machinery. A scanning model of the acceptor region and a model of interaction between the splicing complexes (exon definition model) are also incorporated in the detection process. Subsets of sites presenting deficiencies of several splice site elements could be identified. Further examination of these sites helps to determine lacking elements and refine models.

Computer Simulation↗

Codon usage and gene function are related in sequences of Arabidopsis thaliana.

In this paper, the relationship between codon usage and the physiological pattern of expression of a gene is investigated while considering a dataset of 815 nuclear genes of Arabidopsis thaliana. Factorial Correspondence Analysis, a commonly used multivariate statistical approach in codon usage analysis, was used in order to analyse codon usage bias gene by gene. The analysis reveals a single major trend in codon usage among genes in Arabidopsis. At one end of the trend lie genes with a highly G/C biased codon usage. This group contains mainly photosynthetic and housekeeping genes which are known to encode the most abundant proteins of the vegetal cell. At the other extreme lie genes with a weaker A/T-biased codon usage. This group contain genes with various functions which exhibits most of the time a strong tissue-specific pattern of expression in relation, for example, to stress conditions. These observations were confirmed by the detailed analysis of codon usage in the multigene family of tubulins and appear to be general in plant species, even as distant from Arabidopsis thaliana as a monocotyledonous plant such as maize.

Arabidopsis↗

Global analysis of genomic texts: the distribution of AGCT tetranucleotides in the Escherichia coli and Bacillus subtilis genomes predicts translational frameshifting and ribosomal hopping in several genes.

Present availability of the genomic text of bacteria allows assignment of biological known functions to many genes (typically, half of the genome's gene content). It is now time to try and predict new unexpected functions, using inductive procedures that allow correlating the content of the genomic text to possible biological functions. We show here that analysis of the genomes of Escherichia coli and Bacillus subtilis for the distribution of AGCT motifs predicts that genes exist for which the mRNA molecule can be translated as several different proteins synthesized after ribosomal frameshifting or hopping. Among these genes we found that several coded for the same function in E. coli and B. subtilis. We analyzed in depth the situation of the infB gene (experimentally known to specify synthesis of several proteins differing in their translation starts), the aceF/pdhC gene, the eno gene, and the rplI gene. In addition, genes specific to E. coli were also studied: ompA, ompFand tolA (predicting epigenetic variation that could help escape infection by phages or colicins).

Acetyltransferases↗

A Multi-Agent System for Exon Prediction in Human Sequences.

Given the problem of identifying exons in new genomic DNA, the sketch of a resolution process was drawn using sequence data and models of site/signal recognition. A multi-agent architecture is used to validate these models and test hypotheses on the chronology of events involved in gene splicing. Information is channelled through a hierarchy of agents. Each type of agent is the result of a successful step in the resolution process. The system does not rely on the compositional bias of coding sequences which is a key feature of current computer methods.

Journal Article↗

Very fast identification of RNA motifs in genomic DNA. Application to tRNA search in the yeast genome.

A common strategy characterises the various methods independently defined to identify almost unambiguously different types of RNA molecules in DNA fragments. So far, the good quality of detection of RNA motif has been the prior motivation and effectively delayed the optimisation of programs. As an illustration of possible improvements, a modified version of tRNAscan is described. The previous algorithm was altered to run 500 times faster and to lower both rates of false positives and false negatives. The newly sequenced genome of Saccharomyces cerevisiae is scanned both ways in less than three minutes and results match annotations found in databanks with three exceptions, two of which being arguably not real tRNAs.

Algorithms↗

Exon prediction in eucaryotic genomes.

Two independent computer systems, NetPlantGene and AMELIE, dedicated to the identification of splice sites in plant and human genomes, respectively, are introduced here. Both methods were designed in relation to experimental work; they rely on automatically generated rules involving the nucleotide content of sequences regardless of the coding properties of exons. The specificity of plant sequences as considered in NetPlantGene is shown to enhance the quality of detection as opposed to general methods such as GRAIL. A scanning model of the acceptor site recognition is being simulated by AMELIE leading to a relatively accurate selection process of sites.

Arabidopsis↗

Automatic identification of group I intron cores in genomic DNA sequences.

Automatic identification of the ribozyme core of group I catalytic introns in genomic sequences is shown to be feasible in spite of the scarcity of strictly conserved features in the sequence and secondary structure of group I introns. An algorithm is described that successfully identified 132 out of the 143 currently reported group I cores with a false positive rate of only 10(-6) per nucleotide. The recognition process consists in generating and rating large sets of potential local solutions which are gradually combined into more complex structures until an entire core (six to seven pairings, six connecting segments, three terminal loops) has been assembled. The extent to which successful recognition may be prevented by sequencing errors is assessed. Also discussed are (1) possible relationships between scores allocated by the program and ability to self-splice in vitro and (2) the potential for objectively assessing the degree of relatedness to group I of structures claimed to resemble group I introns.

Algorithms↗