PubMed Health⌕ Search

Biomedical subjects

Edouard Yeramian

Publications and source records attributed to Edouard Yeramian.

7 recordsLinked to original sources

Evolution of proteomes: fundamental signatures and global trends in amino acid compositions.

BACKGROUND: The evolutionary characterization of species and lifestyles at global levels is nowadays a subject of considerable interest, particularly with the availability of many complete genomes. Are there specific properties associated with lifestyles and phylogenies? What are the underlying evolutionary trends? One of the simplest analyses to address such questions concerns characterization of proteomes at the amino acids composition level. RESULTS: In this work, amino acid compositions of a large set of 208 proteomes, with significant number of representatives from the three phylogenetic domains and different lifestyles are analyzed, resorting to an appropriate multidimensional method: Correspondence analysis. The analysis reveals striking discrimination between eukaryotes, prokaryotic mesophiles and hyperthemophiles-themophiles, following amino acid usage. In sharp contrast, no similar discrimination is observed for psychrophiles. The observed distributional properties are compared with various inferred chronologies for the recruitment of amino acids into the genetic code. Such comparisons reveal correlations between the observed segregations of species following amino acid usage, and the separation of amino acids following early or late recruitment. CONCLUSION: A simple description of proteomes according to amino acid compositions reveals striking signatures, with sharp segregations or on the contrary non-discriminations following phylogenies and lifestyles. The distribution of species, following amino acid usage, exhibits a discrimination between [high GC]-[high optimal growth temperatures] and [low GC]-[moderate temperatures] characteristics. This discrimination appears to coincide closely with the separation of amino acids following their inferred early or late recruitment into the genetic code. Taken together the various results provide a consistent picture for the evolution of proteomes, in terms of amino acid usage.

Amino Acids↗

Genome trees from conservation profiles.

The concept of the genome tree depends on the potential evolutionary significance in the clustering of species according to similarities in the gene content of their genomes. In this respect, genome trees have often been identified with species trees. With the rapid expansion of genome sequence data it becomes of increasing importance to develop accurate methods for grasping global trends for the phylogenetic signals that mutually link the various genomes. We therefore derive here the methodological concept of genome trees based on protein conservation profiles in multiple species. The basic idea in this derivation is that the multi-component "presence-absence" protein conservation profiles permit tracking of common evolutionary histories of genes across multiple genomes. We show that a significant reduction in informational redundancy is achieved by considering only the subset of distinct conservation profiles. Beyond these basic ideas, we point out various pitfalls and limitations associated with the data handling, paving the way for further improvements. As an illustration for the methods, we analyze a genome tree based on the above principles, along with a series of other trees derived from the same data and based on pair-wise comparisons (ancestral duplication-conservation and shared orthologs). In all trees we observe a sharp discrimination between the three primary domains of life: Bacteria, Archaea, and Eukarya. The new genome tree, based on conservation profiles, displays a significant correspondence with classically recognized taxonomical groupings, along with a series of departures from such conventional clusterings.

Animals↗

GadE (YhiE): a novel activator involved in the response to acid environment in Escherichia coli.

In several Gram-positive and Gram-negative bacteria glutamate decarboxylases play an important role in the maintenance of cellular homeostasis in acid environments. Here, new insight is brought to the regulation of the acid response in Escherichia coli. Overexpression of yhiE, similarly to overexpression of gadX, a known regulator of glutamate decarboxylase expression, leads to increased resistance of E. coli strains under high acid conditions, suggesting that YhiE is a regulator of gene expression in the acid response. Target genes of both YhiE (renamed GadE) and GadX were identified by a transcriptomic approach. In vitro experiments with GadE purified protein provided evidence that this regulator binds to the promoter region of these target genes. Several of them are clustered together on the chromosome and this chromosomal organization is conserved in many E. coli strains. Detailed structural (in silico) analysis of this chromosomal region suggests that the promoters of the corresponding genes are preferentially denatured. These results, along with the G+C signature of the chromosomal region, support the existence of a fitness island for acid adaptation on the E. coli chromosome.

AraC Transcription Factor↗

GeneFizz: A web tool to compare genetic (coding/non-coding) and physical (helix/coil) segmentations of DNA sequences. Gene discovery and evolutionary perspectives.

The GeneFizz (http://pbga.pasteur.fr/GeneFizz) web tool permits the direct comparison between two types of segmentations for DNA sequences (possibly annotated): the coding/non-coding segmentation associated with genomic annotations (simple genes or exons in split genes) and the physics-based structural segmentation between helix and coil domains (as provided by the classical helix-coil model). There appears to be a varying degree of coincidence for different genomes between the two types of segmentations, from almost perfect to non-relevant. Following these two extremes, GeneFizz can be used for two purposes: ab initio physics-based identification of new genes (as recently shown for Plasmodium falciparum) or the exploration of possible evolutionary signals revealed by the discrepancies observed between the two types of information.

Algorithms↗

The Plasmodium falciparum family of Rab GTPases.

Rab GTPases are key regulators of vesicular traffic in eukaryotic cells. Here we sought a global characterization and description of the Plasmodium falciparum family of Rab GTPases. We used a combination of bioinformatic analyses, experimental testing of predictions, structure modelling and phylogenetics. These analyses led to the identification of seven new parasite Rabs. Accordingly we estimate that the P. falciparum family is made up of 11 genes. We show that ten members of this family are transcribed in infected erythrocytes. Concerning the various members of the family, a series of specific as well as global conclusions can be drawn. Rabs predicted to be compartment-specific show different subcellular distributions. This is demonstrated for PfRab1A and PfRab11A, with the generation of specific antisera. The sequence analyses reveal several peculiarities, with possible functional implications. One of the transcribed genes, Pfrab5b, does not encode a classical C-terminus, suggestive of a novel regulatory role for this GTPase. Another, Pfrab5a, previously identified as a rab gene located on chromosome 2, possesses a 30-amino-acid insertion in its GTP-binding domain. Structural considerations suggest that this insertion could represent a novel interaction interface. We used conserved RabF and RabSF motifs to discriminate between specific parasite Rabs, and followed their predicted change in position on the structure of PfRab6, as GTP is hydrolysed to GDP. This allowed us to propose their involvement in potential interaction surfaces, that we extended to human Rab6 and the motifs known to mediate Rabkinesine-6 binding. Finally, we compared the P. falciparum Rab family to those of Saccharomyces cerevisiae and Schizosaccharomyces pombe and found that parasite Rabs segregate into possible functional clads. Such grouping into clads may give clues to parasite Rab function, and may shed light on P. falciparum secretory/endocytic pathways.

Amino Acid Sequence↗

Amino acid composition of genomes, lifestyles of organisms, and evolutionary trends: a global picture with correspondence analysis.

Can we infer the lifestyle of an organism from the characteristic properties of its genome? More precisely, what are the relations between easily quantifiable properties from genomic sequences, such as amino-acid compositions, and more subtle characteristics concerning for example lifestyles or evolutionary trends? Here, we seek a global picture for such properties, based on a large number (56) of complete genomes, including significant numbers of representatives from the three domains of life. We consider the amino acid compositions of the predicted proteomes, and we use correspondence analysis, as a multivariate method to extract the relevant information from the large-scale data. From these analyses we derive a series of conclusions, concerning lifestyles, as well as physico-chemical and evolutionary trends: (1) correspondence analysis of the amino acid compositions permits discrimination between the three known lifestyles (mesophily/thermophily/hyperthermophily). (2) For various organisms, amino-acid composition properties are essentially driven by GC content, and to a significantly lesser extent by growth temperatures associated with lifestyles. Roughly speaking, the respective contributions of these two components are 57 and 20%. It is notable that these proportions are essentially unchanged with respect to a previous analysis (Nature 393 (1998) 537), which involved only 15 genomes, available at the time. (3) In terms of amino acid compositional biases, two specific 'signatures' for thermophily (in a broad sense, including hyperthermophily) can be detected. First, thermophilic species display a relative abundance in glutamic acid (Glu), concomitantly with the depletion in glutamine. Second, in thermophilic species, the relative abundance in Glu (negative charge) is significantly correlated (Pearson correlation coefficient r=0.83 with P<0.0001), with the increase in the lumped 'pool' lysine+arginine (positive charges). This correlation (absent in mesophiles) could be interpreted on a physico-chemical basis, relevant to the thermostability of proteins. (4) Statistically significant differences are observed between the average lengths of the genes in the surveyed species, which follow their distribution between the three domains of life. Also a significant difference is observed between the average lengths of thermophilic (283.0+/-5.8) versus mesophilic (340+/-9.4) genes. It is thus possible that the 'general' shortening of the primary sequences in thermophilic proteins plays a role in thermostability. (5) Considering various combinations of conservation properties (genes conserved exclusively in eukaryotes, in archaea, in bacteria, in combinations of two domains, etc.) correspondence analysis reveals a trend towards thermophilic-hyperthermophilic profiles for the most conserved subset of genes (ancient genes). (6) When limited to the subset of species-specific genes, correspondence analysis leads to a different picture for the clustering of genomes following amino-acid compositions: for example, the 'core' specific part of a genome can bear lifestyle signatures different from those of the complete genome.Various results are discussed both on methodological and biological grounds. The evolutionary perspectives opened by our analyses are noted.

Amino Acids↗

Physics-based gene identification: proof of concept for Plasmodium falciparum.

The ab initio prediction of new genes in eukaryotic genomes represents a difficult task, notably for the identification of complex split genes. A Physics-Based Gene Identification (PBGI) method was formulated recently (Yeramian, Gene, 255, 139-150, 151-168, 2000a,b) to address this problem, taking as a model the Plasmodium falciparum genome. Here, the predictive power of this method is put under experimental test for this genome. The presented results demonstrate the usefulness of the PBGI as a gene-identification tool for P. falciparum, notably for the discovery of new genes with no homology to known genes. Perspectives opened by this new method for other eukaryotic genomes are also mentioned.

Algorithms↗