PubMed Health⌕ Search

Biomedical subjects

I Sauvaget

Publications and source records attributed to I Sauvaget.

7 recordsLinked to original sources

WOBB.C: a portable software package for defining and searching ambiguous sequence patterns.

WOBB.C is a set of C-written programs designed to build and manipulate ambiguous sequence patterns and to locate them within collections of protein or nucleotide sequences. The search module involves the perceptron algorithm introduced by Stormo et al. (1982) [Nucleic Acids Res 10:2997-3011] in the context of biosequence analysis. The originality of WOBB.C resides in its portability and in a flexible interface, allowing the definition of patterns in three different ways: automatically, from a multi-alignment; interactively, from a character string, or from an explicit text file script.

Amino Acid Sequence↗

Identification of four conserved motifs among the RNA-dependent polymerase encoding elements.

Four consensus sequences are conserved with the same linear arrangement in RNA-dependent DNA polymerases encoded by retroid elements and in RNA-dependent RNA polymerases encoded by plus-, minus- and double-strand RNA viruses. One of these motifs corresponds to the YGDD span previously described by Kamer and Argos (1984). These consensus sequences altogether lead to 4 strictly and 18 conservatively maintained amino acids embedded in a large domain of 120 to 210 amino acids. As judged from secondary structure predictions, each of the 4 motifs, which may cooperate to form a well-ordered domain, places one invariant amino acid in or proximal to turn structures that may be crucial for their correct positioning in a catalytic process. We suggest that this domain may constitute a prerequisite 'polymerase module' implicated in template seating and polymerase activity. At the evolutionary level, the sequence similarities, gap distribution and distances between each motif strongly suggest that the ancestral polymerase module was encoded by an individual genetic element which was most closely related to the plus-strand RNA viruses and the non-viral retroposons. This polymerase module gene may have subsequently propagated in the viral kingdom by distinct gene set recombination events leading to the wide viral variety observed today.

Amino Acid Sequence↗

Objective comparison of exon and intron sequences by means of 2-dimensional data analysis methods.

Here we advocate the use of 2-dimensional data representation in the context of the informational approach of sequence analysis (Claverie & Bougueleret (1986) Nucleic Acids Research 14, 179-196) by applying these methods to the problem of intron/exon discrimination. Two main findings are reported: i) oligonucleotide patterns complementary to the Ul small nuclear RNA are specifically avoided in exon sequences, ii) vertebrate intron sequences, to the exclusion of other eukaryotic phyla, are characterized by a peculiar distribution of CpG containing patterns.

Algorithms↗

Computer generation and statistical analysis of a data bank of protein sequences translated from GenBank.

We describe PGtrans, a new and freely available protein sequence databank (2625 sequences, 554198 amino-acids). This data bank is routinely produced by automatic computer translation of the nucleotide sequence library GenBank. The information needed for the translation process (transcriptional orientation, location of coding regions, splice sites and pertinent genetic code) is gathered by the translation program through an "intelligent" scanning of the documentary field of each GenBank entry. Inconsistencies resulting in unexpected termination codons are detected and reported thus allowing the correction of data bank errors. PGtrans is intended as a tool for protein similarity searches. Its reasonable overall size (2 Moctets) makes it suitable for micro-computer environments. Up to date amino-acid composition data and relative abundances of di-, tri-, and tetra-peptides in proteins of known sequences are presented and discussed.

Amino Acid Sequence↗

Assessing the biological significance of primary structure consensus patterns using sequence databanks. I. Heat-shock and glucocorticoid control elements in eukaryotic promoters.

We describe FORTRAN 77 software allowing for convenient searching of any segmented and ambiguous pattern in the currently available protein or nucleotide sequence databanks. For proteins, this software can be instrumental in defining conserved functional domains among non-homologous overall primary structures. For nucleic acids, it is used in detecting complex and/or low consensus structural or regulatory patterns. As first applications we have studied the distribution of short consensus sequences believed to characterize heat-shock and glucocorticoid regulated promoters. This analysis allowed an evaluation of the specificity, probable role and thus biological significance of various regions of these consensus. In addition, the expression of several known genes are predicted to be heat-shock or glucocorticoid sensitive.

Algorithms↗