PubMed Health⌕ Search

Biomedical subjects

Odile Lecompte

Publications and source records attributed to Odile Lecompte.

6 recordsLinked to original sources

ICDS database: interrupted CoDing sequences in prokaryotic genomes.

Unrecognized frameshifts, in-frame stop codons and sequencing errors lead to Interrupted CoDing Sequence (ICDS) that can seriously affect all subsequent steps of functional characterization, from in silico analysis to high-throughput proteomic projects. Here, we describe the Interrupted CoDing Sequence database containing ICDS detected by a similarity-based approach in 80 complete prokaryotic genomes. ICDS can be retrieved by species browsing or similarity searches via a web interface (http://www-bio3d-igbmc.u-strasbg.fr/ICDS/). The definition of each interrupted gene is provided as well as the ICDS genomic localization with the surrounding sequence. Furthermore, to facilitate the experimental characterization of ICDS, we propose optimized primers for re-sequencing purposes. The database will be regularly updated with additional data from ongoing sequenced genomes. Our strategy has been validated by three independent tests: (i) ICDS prediction on a benchmark of artificially created frameshifts, (ii) comparison of predicted ICDS and results obtained from the comparison of the two genomic sequences of Bacillus licheniformis strain ATCC 14580 and (iii) re-sequencing of 25 predicted ICDS of the recently sequenced genome of Mycobacterium smegmatis. This allows us to estimate the specificity and sensitivity (95 and 82%, respectively) of our program and the efficiency of primer determination.

Bacillus↗

Cloning, purification and crystallization of a Walker-type Pyrococcus abyssi ATPase family member.

Several ATPase proteins play essential roles in the initiation of chromosomal DNA replication in archaea. Walker-type ATPases are defined by their conserved Walker A and B motifs, which are associated with nucleotide binding and ATP hydrolysis. A family of 28 ATPase proteins with non-canonical Walker A sequences has been identified by a bioinformatics study of comparative genomics in Pyrococcus genomes. A high-throughput structural study on P. abyssi has been started in order to establish the structure of these proteins. 16 genes have been cloned and characterized. Six out of the seven soluble constructs were purified in Escherichia coli and one of them, PABY2304, has been crystallized. X-ray diffraction data were collected from selenomethionine-derivative crystals using synchrotron radiation. The crystals belong to the orthorhombic space group C2, with unit-cell parameters a = 79.41, b = 48.63, c = 108.77 A, and diffract to beyond 2.6 A resolution.

Adenosine Triphosphatases↗

vALId: validation of protein sequence quality based on multiple alignment data.

The validation of sequences is essential to perform accurate phylogeny and structure/function analysis. However among the thousands of protein sequences available in the public databases, most have been predicted in silico and have not systematically undergone a quality verification. It has recently become evident that they often contain sequence errors. To address the problem of automatic protein quality control, we have developed vALId, an interactive web interfaced software. Taking advantage of high quality multiple alignments of complete protein sequences (MACS), vALId first warns about the presence of suspicious insertions, deletions (indels) and divergent segments, and second, proposes corrections based on transcripts and genome contigs. In a first evaluation test, hundreds of indels and divergent segments were randomly generated in a manually refined MACS. The sensitivity (Sn) and specificity (Sp) of indel detection were excellent (0.96) while the mean Sn(0.49) and Sp(0.56) of divergent segment delineation depended on the percent identity between sequence neighbors. In a second test, 6195 sequences in 100 MACS corresponding to different functional and structural protein families were analyzed. 65% of the sequences were in silico predictions and 44% of eukaryote predicted proteins were partially incorrect with at least one suspicious indel or divergent segment.

Algorithms↗

PipeAlign: A new toolkit for protein family analysis.

PipeAlign is a protein family analysis tool integrating a five step process ranging from the search for sequence homologues in protein and 3D structure databases to the definition of the hierarchical relationships within and between subfamilies. The complete, automatic pipeline takes a single sequence or a set of sequences as input and constructs a high-quality, validated MACS (multiple alignment of complete sequences) in which sequences are clustered into potential functional subgroups. For the more experienced user, the PipeAlign server also provides numerous options to run only a part of the analysis, with the possibility to modify the default parameters of each software module. For example, the user can choose to enter an existing multiple sequence alignment for refinement, validation and subsequent clustering of the sequences. The aim is to provide an interactive workbench for the validation, integration and presentation of a protein family, not only at the sequence level, but also at the structural and functional levels. PipeAlign is available at http://igbmc.u-strasbg.fr/PipeAlign/.

Internet↗

An integrated analysis of the genome of the hyperthermophilic archaeon Pyrococcus abyssi.

The hyperthermophilic euryarchaeon Pyrococcus abyssi and the related species Pyrococcus furiosus and Pyrococcus horikoshii, whose genomes have been completely sequenced, are presently used as model organisms in different laboratories to study archaeal DNA replication and gene expression and to develop genetic tools for hyperthermophiles. We have performed an extensive re-annotation of the genome of P. abyssi to obtain an integrated view of its phylogeny, molecular biology and physiology. Many new functions are predicted for both informational and operational proteins. Moreover, several candidate genes have been identified that might encode missing links in key metabolic pathways, some of which have unique biochemical features. The great majority of Pyrococcus proteins are typical archaeal proteins and their phylogenetic pattern agrees with its position near the root of the archaeal tree. However, proteins probably from bacterial origin, including some from mesophilic bacteria, are also present in the P. abyssi genome.

Adaptation, Physiological↗

Comparative analysis of ribosomal proteins in complete genomes: an example of reductive evolution at the domain scale.

A comprehensive investigation of ribosomal genes in complete genomes from 66 different species allows us to address the distribution of r-proteins between and within the three primary domains. Thirty-four r-protein families are represented in all domains but 33 families are specific to Archaea and Eucarya, providing evidence for specialisation at an early stage of evolution between the bacterial lineage and the lineage leading to Archaea and Eukaryotes. With only one specific r-protein, the archaeal ribosome appears to be a small-scale model of the eukaryotic one in terms of protein composition. However, the mechanism of evolution of the protein component of the ribosome appears dramatically different in Archaea. In Bacteria and Eucarya, a restricted number of ribosomal genes can be lost with a bias toward losses in intracellular pathogens. In Archaea, losses implicate 15% of the ribosomal genes revealing an unexpected plasticity of the translation apparatus and the pattern of gene losses indicates a progressive elimination of ribosomal genes in the course of archaeal evolution. This first documented case of reductive evolution at the domain scale provides a new framework for discussing the shape of the universal tree of life and the selective forces directing the evolution of prokaryotes.

Animals↗