PubMed Health⌕ Search

Biomedical subjects

B Labedan

Publications and source records attributed to B Labedan.

At least 19 recordsLinked to original sources

Puzzling over orphan enzymes.

Despite the current availability of several hundreds of thousands of amino acid sequences, more than 39% of the well-defined enzyme activities (EC numbers) are not associated with any sequence in major public databases. This wide gap separating knowledge of biochemical function and sequence information is found in nearly all classes of enzymes. Thus, there is an urgent need to explore the 1525 orphan enzymes (EC numbers without associated sequences), in order to progressively bridge this unwanted gap. Improving genome annotation could unveil a significant proportion of sequenceless enzymes. Peptide mass mapping and further genome mining would be useful to identify proper sequence for enzymes found in species for which genetic tools are missing. Finally, the whole community must help major public databases to begin addressing the problem of missing or incomplete information.

Chromosome Mapping↗

EMGLib: the enhanced microbial genomes library (update 2000).

As the number of complete microbial genomes publicly available is still growing, the problem of annotation quality in these very large sequences remains unsolved. Indeed, the number of annotations associated with complete genomes is usually lower than those of the shorter entries encountered in the repository collections. Moreover, classical sequence database management systems have difficulties in handling entries of such size. In this context, the Enhanced Microbial Genomes Library (EMGLib) was developed to try to alleviate these problems. This library contains all the complete genomes from prokaryotes (bacteria and archaea) already sequenced and the yeast genome in GenBank format. The annotations are improved by the introduction of data on codon usage, gene orientation on the chromosome and gene families. It is possible to access EMGLib through two database systems set up on WWW servers: the PBIL server at http://pbil.univ-lyon1.fr/emglib.html and the MICADO server at http://locus.jouy.inra.fr/micado

Base Sequence↗

The Enhanced Microbial Genomes Library.

Since the obtention of the complete sequence of Haemophilus influenzae Rd in 1995, the number of bacterial genomes entirely sequenced has regularly increased. A problem is that the quality of the annotations of these very large sequences is usually lower than those of the shorter entries encountered in the repository collections. Moreover, classical sequence database management systems have difficulties in handling entries of that size. In this context, we have decided to build the Enhanced Microbial Genomes Library (EMGLib) in which these two problems are alleviated. This library contains all the complete genomes from bacteria already sequenced and the yeast genome in GenBank format. The annotations are improved by the introduction of data on codon usage, gene orientation on the chromosome and gene families. It is possible to access EMGLib through two database systems set up on World Wide Web servers: the PBIL server at http://pbil.univ-lyon1.fr/emglib/emglib. html and the MICADO server at http://locus.jouy.inra.fr/micado

Base Sequence↗

The evolutionary history of carbamoyltransferases: A complex set of paralogous genes was already present in the last universal common ancestor.

Forty-four sequences of ornithine carbamoyltransferases (OTCases) and 33 sequences of aspartate carbamoyltransferases (ATCases) representing the three domains of life were multiply aligned and a phylogenetic tree was inferred from this multiple alignment. The global topology of the composite rooted tree (each enzyme family being used as an outgroup to root the other one) suggests that present-day genes are derived from paralogous ancestral genes which were already of the same size and argues against a mechanism of fusion of independent modules. A closer observation of the detailed topology shows that this tree could not be used to assess the actual order of organismal descent. Indeed, this tree displays a complex topology for many prokaryotic sequences, with polyphyly for Bacteria in both enzyme trees and for the Archaea in the OTCase tree. Moreover, representatives of the two prokaryotic Domains are found to be interspersed in various combinations in both enzyme trees. This complexity may be explained by assuming the occurrence of two subfamilies in the OTCase tree (OTC alpha and OTC beta) and two other ones in the ATCase tree (ATC I and ATC II). These subfamilies could have arisen from duplication and selective losses of some differentiated copies during the successive speciations. We suggest that Archaea and Eukaryotes share a common ancestor in which the ancestral copies giving the present-day ATC II/OTC beta combinations were present, whereas Bacteria comprise two classes: one containing the ATC II/OTC alpha combination and the other harboring the ATC I/OTC beta combination. Moreover, multiple horizontal gene transfers could have occurred rather recently amongst prokaryotes. Whichever the actual history of carbamoyltransferases, our data suggest that the last common ancestor to all extant life possessed differentiated copies of genes coding for both carbamoyltransferases, indicating it as a rather sophisticated organism.

Amino Acid Sequence↗

Isolation of a minD-like gene in the hyperthermophilic archaeon Pyrococcus AL585, and phylogenetic characterization of related proteins in the three domains of life.

The region upstream of the dinF-like gene of the hyperthermophilic archaeon Pyrococcus strain AL585 has been cloned and sequenced. This region contains an open reading frame (ORF) that encodes a polypeptide with a high similarity to MinD proteins and their Mrp paralogues. Transcripts of the dinF-like and the minD-like genes were detected by RT-PCR, indicating that they are both expressed in vivo. The MinD and MinD-like proteins belong to a broad family of ATPases involved in chromosome and plasmid partitioning. MinD-like proteins can be defined by specific amino-acid sequence signatures. A systematic search for proteins sharing these signatures in current databases and newly sequenced genomes show that MinD-like proteins are present in all archaeal genomes sequenced so far, often in several copies. Phylogenetic analysis identifies two groups of MinD-like proteins which are also characterized by more conserved amino-acid motifs. A first group, which includes the Escherichia coli MinD and the Pyrococcus AL585 MinDL protein, contains only procaryotic proteins. This group can be further divided into a subgroup of archaeal proteins and two subgroups of bacterial proteins. A second group includes proteins more related to the E. coli Mrp protein and contains representants of the three domains of life. The conservation of MinD-like proteins in the three domains of life suggests that these proteins play a central role in cellular metabolism.

Adenosine Triphosphatases↗

The evolutionary relationships between the two bacteria Escherichia coli and Haemophilus influenzae and their putative last common ancestor.

We have tried to approach the nature of the last common ancestor to Haemophilus influenzae and Escherichia coli and to determine how each bacterium could have diverged from this putative organism. The approach used was exhaustive analysis of the homologous proteins coded by genes present in these bacteria, using as criteria for sequence relatedness an alignment of at least 80 amino acid residues and a PAM distance (number of accepted point mutations per 100 residues separating two sequences) below 250. Evolutionarily significant similarities were found between 1,345 H. influenzae proteins (85% of the total genome) and 3,058 E. coli. proteins (75% of the total genome), many of them belonging to families of various sizes (from 666 doublets to 35 large groups of more than 10 members). Nearly all the genes found by this approach to be duplicated in both bacteria were already duplicated in their last common ancestor. This was deduced from (1) the comparison of the respective distributions of evolutionary distances between orthologs (genes separated only by speciation events) and paralogs (genes duplicated in the same genome) and (2) the analysis of the phylogenetic trees reconstructed for each family of paralogs containing at least two members belonging to each bacterium. The distributions of the different categories of homologs show a significant loss of paralogous genes in H. influenzae (reduction proportional to the genome size), of many sequences which are still present in one copy in E. coli, and of some entire gene families. Phylogenetic trees also confirmed this recent loss of paralogous genes in H. influenzae. Thus, the genome size of the last common ancestor of these two bacteria would have been close to that of present-day E. coli, and the evolution of H. influenzae toward a parasitic life led to an important decrease in its genome size by some mechanism of streamlining. During this recent evolution, the memory of the gene order present in the last common ancestor has been blurred, but a few short conserved chromosomal fragments can still be detected in present-day E. coli and H. influenzae.

Bacterial Proteins↗

Protein evolution viewed through Escherichia coli protein sequences: introducing the notion of a structural segment of homology, the module.

Paralogous genes are genes which descend from a progenitor gene which has duplicated as an ancestral gene, each copy having diverged prior to speciation. With comprehensive information available on functions of Escherichia coli proteins, analysis of sequence-related E. coli paralogous proteins can give information on the early ancestors of families of proteins now residing in many contemporary organisms, such as the enzymes of metabolism, some kinds of transport mechanisms and some kinds of regulatory mechanisms. In the first step, we have confirmed that E. coli contains a very high proportion of paralogous proteins. Next, we have defined two main classes of paralogous proteins. One class is formed of proteins which contain a unique structural segment homologous to a single set of related proteins. The other class corresponds to proteins which contain more than one structural segment of homology, each segment homologous to unrelated sets of proteins. We define such an independent structural segment of homology as a module. This modular structure (mean length equivalent to 209 amino acids) corresponds often to entire proteins, but there are also proteins that appear to be assembled from two or three independent modules having independent origins. Most multimodular proteins appear to have been formed early in their history, a minority appear to be relatively recent fusions of independent modules. Examining 1404 independent structural segments of homology, composed of both modules and entire proteins, we found that the segments of homology fell into 352 sequence-related groups or families. The majority of these families (ranging from 2 to 62 members) are functionally homogeneous. This strongly suggests that the 1404 present-day modules and proteins derive from a minimal set of 352 ancestral modules, each one being already of the same size and having a function similar to all members of its progeny.

Bacterial Proteins↗

Glutamate dehydrogenase from the hyperthermophilic bacterium Thermotoga maritima: molecular characterization and phylogenetic implications.

The hyperthermophilic bacterium Thermotoga maritima, which grows at up to 90 degrees C, contains an L-glutamate dehydrogenase (GDH). Activity of this enzyme could be detected in T. maritima crude extracts, and appeared to be associated with a 47-kDa protein which cross-reacted with antibodies against purified GDH from the hyperthermophilic archaeon Pyrococcus woesei. The single-copy T. maritima gdh gene was cloned by complementation in a glutamate auxotrophic Escherichia coli strain. The nucleotide sequence of the gdh gene predicts a 416-residue protein with a calculated molecular weight of 45,852. The gdh gene was inserted in an expression vector and expressed in E. coli as an active enzyme. The T. maritima GDH was purified to homogeneity. The NH2-terminal sequence of the purified enzyme was PEKSLYEMAVEQ, which is identical to positions 2-13 of the peptide sequence derived from the gdh gene. The purified native enzyme has a size of 265 kDa and a subunit size of 47kDa, indicating that GDH is a homohexamer. Maximum activity of the enzyme was measured at 75 degrees C and the pH optima are 8.3 and 8.8 for the anabolic and catabolic reaction, respectively. The enzyme was found to be very stable at 80 degrees C, but appeared to lose activity quickly at higher temperatures. The T. maritima GDH shows the highest rate of activity with NADH (Vmax of 172 U/mg protein), but also utilizes NADPH (Vmax of 12 U/mg protein). Sequence comparisons showed that the T. maritima GDH is a member of the family II of hexameric GDHs which includes all the GDHs isolated so far from hyperthermophiles. Remarkably, phylogenetic analysis positions all these hyperthermophilic GDHs in the middle of the GDH family II tree, with the bacterial T. maritima GDH located between that of halophilic and thermophilic euryarchaeota.

Amino Acid Sequence↗

A gyrB-like gene from the hyperthermophilic bacterion Thermotoga maritima.

We have cloned and sequenced two overlapping DNA fragments (3236 bp) containing a gene encoding the ATPase subunit of a type II DNA topoisomerase from the hyperthermophilic bacterion Thermotoga maritima (Tm Top2B). The deduced protein is composed of 636 aa with a calculated molecular mass of 72415 Da. It shares significant similarities with the ATPase subunits of mesophilic bacterial DNA topoisomerases II, either DNA gyrase (GyrB) or DNA topoisomerase IV (ParE). Although the highest similarity scores are obtained with GyrB proteins (55% identity with Bacillus subtilis DNA gyrase), a detailed phylogenetic analysis of all known DNA topoisomerases II does not allow us to determine if Tm Top2B corresponds to a DNA gyrase or a DNA topoisomerase IV. This hyperthermophilic Top2B protein exhibits a larger amount of charged amino acids than its mesophilic homologues, a feature which could be important for its thermostability. No gyrA-like gene has been found near top2B. A gene coding for a transaminase B-like protein was found in the upstream region of top2B.

Amino Acid Sequence↗

The adenylosuccinate synthetase from the hyperthermophilic archaeon Pyrococcus species displays unusual structural features.

The first example of a hyperthermophilic adenylosuccinate synthetase is reported, which is an enzyme that must maintain its folded structure at temperatures as high as 102 degrees C. The amino acid sequence of this key enzyme has been determined after cloning and sequencing the purA-like gene from the archaeal Pyrococcus sp. strain ST700. The corresponding protein displays two unexpected features: (1) it is 21% shorter than the homologous mesophilic enzymes and this shortening corresponds to the loss of two alpha-helices and three beta-strands present in the Escherichia coli enzyme; (2) surprisingly, the archaeal adenylosuccinate synthetase has a significant number of substitutions in residues that are conserved in all other homologous enzymes from bacteria to man. In E. coli, the conserved residues have been described as essential for catalytic activity and/or for maintaining the folded structure of the homodimer. Despite these drastic differences, the purA-like archaeal gene seems to be normally expressed and its product functions in vivo in bacteria, since it complemented an E. coli purA auxotroph. The archaeal adenylosuccinate synthetase appears to be a good example of a bona fide orthologous protein. Reconstruction of phylogenetic trees showed that the archaeal gene is equally distantly related to both eukaryotes and bacteria, independently of the numerous substitutions observed at critical positions.

Adenylosuccinate Synthase↗

The gene encoding the beta-1,4-endoglucanase (CelA) from Myxococcus xanthus: evidence for independent acquisition by horizontal transfer of binding and catalytic domains from actinomycetes.

The celA gene encoding a beta-1,4 endoglucanase (CelA) from Myxococcus xanthus has been cloned in Escherichia coli and sequenced. The C-terminal region of CelA displayed a high level of similarity with the catalytic domain of several Egl belonging to the glycosyl hydrolases family 6 (CenA from Cellulomonas fimi, CelA from Microbispora bispora, E2 from Thermonospora fusca, CasA from Streptomyces KSM9 and CelA1 from Streptomyces halstedii) and less similarity to the cellobiohydrolases of the fungi Trichoderma reesei and Agaricus bisporus. Using PCR amplification we found in another myxobacterium, Stigmatella aurantiaca, a part of a glycosyl hydrolase belonging to the same family. The N-terminal part of CelA displayed significant similarities with the cellulose-binding domain of other cellulases belonging to a rare subset of family II, such as the avicelase I from Streptomyces reticuli, both tandem repeats N1 and N2 of the cellulase CenC from Cellulomonas fimi, and the N-terminal part of the Egl E1 from Thermonospora fusca. Analyses of the multiple alignments and reconstruction of phylogenetic trees strongly suggest that both domains of CelA were acquired by independent horizontal transfers between Gram+ soil bacteria and scavenging myxobacteria followed by domain shuffling.

Actinomyces↗

Gene products of Escherichia coli: sequence comparisons and common ancestries.

Sequences of 1,862 chromosomally encoded Escherichia coli K12 proteins were examined to identify genes likely to have arisen by duplication of genes in an ancestral chromosome. The criteria for sequence relatedness were an alignment of at least 100 amino acid residues and a PAM distance (number of accepted point mutations per 100 residues separating two sequences) below 250. A total of 971 of the 1,862 proteins examined were found in 2,329 sequence-related pairs that met these criteria. Most proteins of the sequence-related pairs were related in cellular function, as judged by biochemical and/or physiological features. Many of the pairs of proteins could be grouped into sequence-related families. If such groupings were generated from ancestral genes by duplication and divergence events, through these sequence comparisons we can identify putative ancestral sequences of the present-day genes of E. coli and other organisms. The results suggest that the 971 paralogous genes could have been derived from only 204 ancestral genes. We have also shown that the process of duplication and divergence is not the exclusive mechanism of evolution of all E. coli genes. Indeed, the relationships among the sequences of multiple (in the sense of redundant) enzymes indicate that nearly half could have arisen either by convergent evolution or by lateral transfer. Therefore, not all functionally related genes need arise by duplication and divergence.

Bacterial Proteins↗

Widespread protein sequence similarities: origins of Escherichia coli genes.

To learn more about the evolutionary origins of Escherichia coli genes, we surveyed systematically for extended sequence similarities among the 1,264 amino acid sequences encoded by chromosomal genes of E. coli K-12 in SwissProt release 26 by using the FASTA program and imposing the following criteria: (i) alignment of segments at least 100 amino acids long and (ii) at least 20% amino acid identity. Altogether, 624 extended alignments meeting the two criteria were identified, corresponding to 577 protein sequences (45.6% of the 1,264 E. coli protein sequences) that had an extended alignment with at least one other E. coli protein sequence. To exclude alignments of questionable biological significance, we imposed a high threshold on the number of gaps allowed in each of the 624 extended alignments, giving us a subset of 464 proteins. The population of 464 alignments has the following characteristics expressed as median values of the group: 254 amino acids in the alignment, representing 86% of the length of the protein, 33% of the amino acids in the alignment being identical, and 1.1 gaps introduced per 100 amino acids of alignment. Where functions are known, nearly all pairs consist of functionally related proteins. This implies that the sequence similarity we detected has biological meaning and did not arise by chance. That a major fraction of E. coli proteins form extended alignments strongly suggests the predominance of duplication and divergence of ancestral genes in the evolution of E. coli genes. The range of degrees of similarity shows that some genes originated more recently than others. There is no evidence of genome doubling in the past, since map distances between genes of sequence-related proteins show no coherent pattern of favored separations.

Algorithms↗

PCR-mediated cloning and sequencing of the gene encoding glutamate dehydrogenase from the archaeon Sulfolobus shibatae: identification of putative amino-acid signatures for extremophilic adaptation.

Highly degenerate oligodeoxyribonucleotides (oligos) were used to PCR amplify the most conserved region of the glutamate dehydrogenase (GDH)-encoding gene from the extreme thermophilic archaeon, Sulfolobus shibatae. The amplified fragment was cloned and sequenced, and then used as a homologous probe to clone a genomic restriction fragment containing the near-complete gdhA gene. The deduced amino acid (aa) sequence shows a very high degree of similarity with the aa sequence previously determined by direct sequencing of the purified enzyme from Sulfolobus solfataricus [Maras et al., Eur. J. Biochem. 203 (1992) 81-87]. The introduction of this new sequence into our GDH phylogenetic trees [Benachenhou-Lahfa et al., J. Mol. Evol. 35 (1993) 335-346] showed that it is a new member of hexameric GDH family II, and did not modify the topology of the trees. Comparison of the primary structures of extremophilic GDH enzymes (halophilic, thermophilic and hyperthermophilic) with those of their non-halophilic and mesophilic counterparts in this family II led us to identify a few aa changes which seem to be specific either to hyperthermophilic or halophilic adaptation.

Adaptation, Physiological↗

Evolution of glutamate dehydrogenase genes: evidence for two paralogous protein families and unusual branching patterns of the archaebacteria in the universal tree of life.

The existence of two families of genes coding for hexameric glutamate dehydrogenases has been deduced from the alignment of 21 primary sequences and the determination of the percentages of similarity between each pair of proteins. Each family could also be characterized by specific motifs. One family (Family I) was composed of gdh genes from six eubacteria and six lower eukaryotes (the primitive protozoan Giardia lamblia, the green alga Chlorella sorokiniana, and several fungi and yeasts). The other one (Family II) was composed of gdh genes from two eubacteria, two archaebacteria, and five higher eukaryotes (vertebrates). Reconstruction of phylogenetic trees using several parsimony and distance methods confirmed the existence of these two families. Therefore, these results reinforced our previously proposed hypothesis that two close but already different gdh genes were present in the last common ancestor to the three Ur-kingdoms (eubacteria, archaebacteria, and eukaryotes). The branching order of the different species of Family I was found to be the same whatever the method of tree reconstruction although it varied slightly according the region analyzed. Similarly, the topological positions of eubacteria and eukaryotes of Family II were independent of the method used. However, the branching of the two archaebacteria in Family II appeared to be unexpected: (1) the thermoacidophilic Sulfolobus solfataricus was found clustered with the two eubacteria of this family both in parsimony and distance trees, a situation not predicted by either one of the contradictory trees recently proposed; and (2) the branching of the halophilic Halobacterium salinarium varied according to the method of tree construction: it was closer to the eubacteria in the maximum parsimony tree and to eukaryotes in distance trees. Therefore, whatever the actual position of the halophilic species, archaebacteria did not appear to be monophyletic in these gdh gene trees. This result questions the firmness of the presently accepted interpretation of previous protein trees which were supposed to root unambiguously the universal tree of life and place the archaebacteria in this tree.

Algorithms↗

The nature of the last universal ancestor and the root of the tree of life, still open questions.

The nature of the last universal ancestor to all extent cellular organisms and the rooting of the universal tree of life are fundamental questions which can now be addressed by molecular evolutionists. Several scenarios have been proposed during the last years, based on the phylogenies of ribosomal RNA and of duplicated proteins, which suggest that the last universal ancestor was either an RNA progenote or an hyperthermophilic prokaryote. We discuss these hypotheses in the light of new data on the evolution of DNA metabolizing enzymes and of contradictions between different protein phylogenies. We conclude that the last universal ancestor was a member of the DNA world already containing several DNA polymerases and DNA topoisomerases. Furthermore, we criticize current data which suggest that the rooting of the universal tree of life is located in the eubacterial branch and we conclude that both rooting the universal tree and the nature of the last universal ancestor are still open questions.

Archaea↗

Influence of DNA supercoiling on the loss of culturability of Escherichia coli cells incubated in seawater.

The relationship between the loss of culturability of Escherichia coli cells in seawater and the DNA supercoiling level of a reporter plasmid (pUC8) have been studied under different experimental conditions. Transfer to seawater of cells grown at low osmolarity decreased their ability to grow without apparent modification of the plasmid supercoiling. We found that E. coli cells could be protected against seawater-induced loss of culturability by increasing their DNA-negative supercoiling in response to environmental factors: either a growth at high osmolarity before the transfer to seawater, or addition of organic matter (50-mg/l peptone) in seawater. We further found conditions where a DNA-induced relaxation was accompanied by an increase in seawater sensitivity. Indeed, inactivation of either one of the subunits A and B of DNA gyrase, which leads to important DNA relaxation, was accompanied in both cases by an increased loss of culturability of conditional mutants after transfer to seawater which could not be explained uniquely by the increase in the temperature required to inactivate the gyrase. Similarly, a strain harbouring a mutation in topoisomerase I, compensated by another mutation in subunit B of the gyrase, was more sensitive to seawater than the isogenic wild-type cell and this greater sensitivity was correlated to a relaxation of plasmid DNA. Again, in these different cases, a previous growth at high osmolarity protected against this seawater sensitivity. We thus propose that the ability of E. coli cells to survive in seawater and maintain their ability to grow on culture media could be linked, at least in part, to the topological state of their DNA.(ABSTRACT TRUNCATED AT 250 WORDS)

Culture Media↗