PubMed Health⌕ Search

Biomedical subjects

M Y Galperin

Publications and source records attributed to M Y Galperin.

At least 19 recordsLinked to original sources

Acetyl-CoA synthetase from the amitochondriate eukaryote Giardia lamblia belongs to the newly recognized superfamily of acyl-CoA synthetases (Nucleoside diphosphate-forming).

The gene coding for the acetyl-CoA synthetase (ADP-forming) from the amitochondriate eukaryote Giardia lamblia has been expressed in Escherichia coli. The recombinant enzyme exhibited the same substrate specificity as the native enzyme, utilizing acetyl-CoA and adenine nucleotides as preferred substrates and less efficiently, propionyl- and succinyl-CoA. N- and C-terminal parts of the G. lamblia acetyl-CoA synthetase sequence were found to be homologous to the alpha- and beta-subunits, respectively, of succinyl-CoA synthetase. Sequence analysis of homologous enzymes from various bacteria, archaea, and the eukaryote, Plasmodium falciparum, identified conserved features in their organization, which allowed us to delineate a new superfamily of acyl-CoA synthetases (nucleoside diphosphate-forming) and its signature motifs. The representatives of this new superfamily of thiokinases vary in their domain arrangement, some consisting of separate alpha- and beta-subunits and others comprising fusion proteins in alpha-beta or beta-alpha orientation. The presence of homologs of acetyl-CoA synthetase (ADP-forming) in such human pathogens as G. lamblia, Yersinia pestis, Bordetella pertussis, Pseudomonas aeruginosa, Vibrio cholerae, Salmonella typhi, Porphyromonas gingivalis, and the malaria agent P. falciparum suggests that they might be used as potential drug targets.

Acetate-CoA Ligase↗

Aldolases of the DhnA family: a possible solution to the problem of pentose and hexose biosynthesis in archaea.

Sequence analysis of the recently identified class I aldolase of Escherichia coli (dhnA gene product) helped to identify its homologs in Chlamydia trachomatis, Chlamydiophyla pneumoniae and in each of the completely sequenced archaeal genomes. Iterative database searches revealed sequence similarities between the DhnA-family enzymes, deoxyribose phosphate aldolases and bacterial (class II) fructose bisphosphate aldolases and allowed prediction of similar three-dimensional structures (TIM-barrel fold) in all these enzymes. The Schiff base-forming lysyl residues of DhnA and deoxyribose phosphate aldolase are conserved in all members of the DhnA and deoxyribose phosphate aldolase families, indicating that these enzymes share common features with both class I and class II aldolases. The DhnA-family enzymes are predicted to possess an aldolase activity and to play a critical role in sugar biosynthesis in archaea.

Amino Acid Sequence↗

The COG database: a tool for genome-scale analysis of protein functions and evolution.

Rational classification of proteins encoded in sequenced genomes is critical for making the genome sequences maximally useful for functional and evolutionary studies. The database of Clusters of Orthologous Groups of proteins (COGs) is an attempt on a phylogenetic classification of the proteins encoded in 21 complete genomes of bacteria, archaea and eukaryotes (http://www. ncbi.nlm. nih.gov/COG). The COGs were constructed by applying the criterion of consistency of genome-specific best hits to the results of an exhaustive comparison of all protein sequences from these genomes. The database comprises 2091 COGs that include 56-83% of the gene products from each of the complete bacterial and archaeal genomes and approximately 35% of those from the yeast Saccharomyces cerevisiae genome. The COG database is accompanied by the COGNITOR program that is used to fit new proteins into the COGs and can be applied to functional and phylogenetic annotation of newly sequenced genomes.

Database Management Systems↗

Who's your neighbor? New computational approaches for functional genomics.

Several recently developed computational approaches in comparative genomics go beyond sequence comparison. By analyzing phylogenetic profiles of protein families, domain fusions, gene adjacency in genomes, and expression patterns, these methods predict many functional interactions between proteins and help deduce specific functions for numerous proteins. Although some of the resultant predictions may not be highly specific, these developments herald a new era in genomics in which the benefits of comparative analysis of the rapidly growing collection of complete genomes will become increasingly obvious.

Algorithms↗

Searching for drug targets in microbial genomes.

Comparative analysis of the complete genome sequences of 10 bacterial pathogens available in the public databases offers the first insights into the drug discovery approaches of the near future. Genes that are conserved in different genomes often turn out to be essential, which makes them attractive targets for new broad-spectrum antibiotics. Subtractive genome analysis reveals the genes that are conserved in all or most of the pathogenic bacteria but not in eukaryotes; these are the most obvious candidates for drug targets. Species-specific genes, on the other hand, may offer the possibility to design drugs against a particular, narrow group of pathogens.

Anti-Bacterial Agents↗

Functional genomics and enzyme evolution. Homologous and analogous enzymes encoded in microbial genomes.

Computational analysis of complete genomes, followed by experimental testing of emerging hypotheses--the area of research often referred to as 'functional genomics'--aims at deciphering the wealth of information contained in genome sequences and at using it to improve our understanding of the mechanisms of cell function. This review centers on the recent progress in the genome analysis with special emphasis on the new insights in enzyme evolution. Standard methods of predicting functions for new proteins are listed and the common errors in their application are discussed. A new method of improving the functional predictions is introduced, based on a phylogenetic approach to functional prediction, as implemented in the recently constructed Clusters of Orthologous Groups (COG) database (available at http:@www.ncbi.nlm.nih.gov/COG). This approach provides a convenient way to characterize the protein families (and metabolic pathways) that are present or absent in any given organism. Comparative analysis of microbial genomes based on this approach shows that metabolic diversity generally correlates with the genome size-parasitic bacteria code for fewer enzymes and lesser number of metabolic pathways than their free-living relatives. Comparison of different genomes reveals another evolutionary trend, the non-orthologous gene displacement of some enzymes by unrelated proteins with the same cellular function. An examination of the phylogenetic distribution of such cases provides new clues to the problems of biochemical evolution, including evolution of glycolysis and the TCA cycle.

Databases, Factual↗

Comparative genomics of the Archaea (Euryarchaeota): evolution of conserved protein families, the stable core, and the variable shell.

Comparative analysis of the protein sequences encoded in the four euryarchaeal species whose genomes have been sequenced completely (Methanococcus jannaschii, Methanobacterium thermoautotrophicum, Archaeoglobus fulgidus, and Pyrococcus horikoshii) revealed 1326 orthologous sets, of which 543 are represented in all four species. The proteins that belong to these conserved euryarchaeal families comprise 31%-35% of the gene complement and may be considered the evolutionarily stable core of the archaeal genomes. The core gene set includes the great majority of genes coding for proteins involved in genome replication and expression, but only a relatively small subset of metabolic functions. For many gene families that are conserved in all euryarchaea, previously undetected orthologs in bacteria and eukaryotes were identified. A number of euryarchaeal synapomorphies (unique shared characters) were identified; these are protein families that possess sequence signatures or domain architectures that are conserved in all euryarchaea but are not found in bacteria or eukaryotes. In addition, euryarchaea-specific expansions of several protein and domain families were detected. In terms of their apparent phylogenetic affinities, the archaeal protein families split into bacterial and eukaryotic families. The majority of the proteins that have only eukaryotic orthologs or show the greatest similarity to their eukaryotic counterparts belong to the core set. The families of euryarchaeal genes that are conserved in only two or three species constitute a relatively mobile component of the genomes whose evolution should have involved multiple events of lineage-specific gene loss and horizontal gene transfer. Frequently these proteins have detectable orthologs only in bacteria or show the greatest similarity to the bacterial homologs, which might suggest a significant role of horizontal gene transfer from bacteria in the evolution of the euryarchaeota.

Amino Acid Sequence↗

Purification, cloning, and expression of an apyrase from the bed bug Cimex lectularius. A new type of nucleotide-binding enzyme.

An enzyme that hydrolyzes the phosphodiester bonds of nucleoside tri- and diphosphates, but not monophosphates, thus displaying apyrase (EC 3.6.1.5) activity, was purified from salivary glands of the bed bug, Cimex lectularius. The purified C. lectularius apyrase was an acidic protein with a pI of 5.1 and molecular mass of approximately 40 kDa that inhibited ADP-induced platelet aggregation and hydrolyzed platelet agonist ADP with specific activity of 379 units/mg protein. Amplification of C. lectularius cDNA corresponding to the N-terminal sequence of purified apyrase produced a probe that allowed identification of a 1.3 kilobase pair cDNA clone coding for a protein of 364 amino acid residues, the first 35 of which constituted the signal peptide. The processed form of the protein was predicted to have a molecular mass of 37.5 kDa and pI of 4.95. The identity of the product of the cDNA clone with native C. lectularius apyrase was proved by immunological testing and by expressing the gene in a heterologous host. Immune serum made against a synthetic peptide with sequence corresponding to the C-terminal region of the predicted cDNA clone recognized both C. lectularius apyrase fractions eluted from a molecular sieving high pressure liquid chromatography and the apyrase active band from chromatofocusing gels. Furthermore, transfected COS-7 cells secreted a Ca2+-dependent apyrase with a pI of 5.1 and immunoreactive material detected by the anti-apyrase serum. C. lectularius apyrase has no significant sequence similarity to any other known apyrases, but homologous sequences have been found in the genome of the nematode C. elegans and in mouse and human expressed sequence tags from fetal and tumor EST libraries.

Amino Acid Sequence↗

A superfamily of metalloenzymes unifies phosphopentomutase and cofactor-independent phosphoglycerate mutase with alkaline phosphatases and sulfatases.

Sequence analysis of the probable archaeal phosphoglycerate mutase resulted in the identification of a superfamily of metalloenzymes with similar metal-binding sites and predicted conserved structural fold. This superfamily unites alkaline phosphatase, N-acetylgalactosamine-4-sulfatase, and cerebroside sulfatase, enzymes with known three-dimensional structures, with phosphopentomutase, 2,3-bisphosphoglycerate-independent phosphoglycerate mutase, phosphoglycerol transferase, phosphonate monoesterase, streptomycin-6-phosphate phosphatase, alkaline phosphodiesterase/nucleotide pyrophosphatase PC-1, and several closely related sulfatases. In addition to the metal-binding motifs, all these enzymes contain a set of conserved amino acid residues that are likely to be required for the enzymatic activity. Mutational changes in the vicinity of these residues in several sulfatases cause mucopolysaccharidosis (Hunter, Maroteaux-Lamy, Morquio, and Sanfilippo syndromes) and metachromatic leucodystrophy.

Alkaline Phosphatase↗

Beyond complete genomes: from sequence to structure and function.

Computer analysis of complete prokaryotic genomes shows that microbial proteins are in general highly conserved--approximately 70% of them contain ancient conserved regions. This allows us to delineate families of orthologs across a wide phylogenetic range and, in many cases, predict protein functions with considerable precision. Sequence database searches using newly developed, sensitive algorithms result in the unification of such orthologous families into larger superfamilies sharing common sequence motifs. For many of these superfamilies, prediction of the structural fold and specific amino acid residues involved in enzymatic catalysis is possible. Taken together, sequence and structure comparisons provide a powerful methodology that can successfully complement traditional experimental approaches.

Animals↗

Analogous enzymes: independent inventions in enzyme evolution.

It is known that the same reaction may be catalyzed by structurally unrelated enzymes. We performed a systematic search for such analogous (as opposed to homologous) enzymes by evaluating sequence conservation among enzymes with the same enzyme classification (EC) number using sensitive, iterative sequence database search methods. Enzymes without detectable sequence similarity to each other were found for 105 EC numbers (a total of 243 distinct proteins). In 34 cases, independent evolutionary origin of the suspected analogous enzymes was corroborated by showing that they possess different structural folds. Analogous enzymes were found in each class of enzymes, but their overall distribution on the map of biochemical pathways is patchy, suggesting multiple events of gene transfer and selective loss in evolution, rather than acquisition of entire pathways catalyzed by a set of unrelated enzymes. Recruitment of enzymes that catalyze a similar but distinct reaction seems to be a major scenario for the evolution of analogous enzymes, which should be taken into account for functional annotation of genomes. For many analogous enzymes, the bacterial form of the enzyme is different from the eukaryotic one; such enzymes may be promising targets for the development of new antibacterial drugs.

Amino Acid Sequence↗

A diverse superfamily of enzymes with ATP-dependent carboxylate-amine/thiol ligase activity.

The recently developed PSI-BLAST method for sequence database search and methods for motif analysis were used to define and expand a superfamily of enzymes with an unusual nucleotide-binding fold, referred to as palmate, or ATP-grasp fold. In addition to D-alanine-D-alanine ligase, glutathione synthetase, biotin carboxylase, and carbamoyl phosphate synthetase, enzymes with known three-dimensional structures, the ATP-grasp domain is predicted in the ribosomal protein S6 modification enzyme (RimK), urea amidolyase, tubulin-tyrosine ligase, and three enzymes of purine biosynthesis. All these enzymes possess ATP-dependent carboxylate-amine ligase activity, and their catalytic mechanisms are likely to include acylphosphate intermediates. The ATP-grasp superfamily also includes succinate-CoA ligase (both ADP-forming and GDP-forming variants), malate-CoA ligase, and ATP-citrate lyase, enzymes with a carboxylate-thiol ligase activity, and several uncharacterized proteins. These findings significantly extend the variety of the substrates of ATP-grasp enzymes and the range of biochemical pathways in which they are involved, and demonstrate the complementarity between structural comparison and powerful methods for sequence analysis.

Adenosine Triphosphate↗

Prokaryotic genomes: the emerging paradigm of genome-based microbiology.

Comparative analysis of the complete sequences of seven bacterial and three archaeal genomes leads to the first generalizations of emerging genome-based microbiology. Protein sequences are, generally, highly conserved, with -70% of the gene products in bacteria and archaea containing ancient conserved regions. In contrast, there is little conservation of genome organization, except for a few essential operons. The most striking conclusions derived by comparison of multiple genomes from phylogenetically distant species are that the number of universally conserved gene families is very small and that multiple events of horizontal gene transfer and genome fusion are major forces in evolution.

Archaea↗

Comparison of archaeal and bacterial genomes: computer analysis of protein sequences predicts novel functions and suggests a chimeric origin for the archaea.

Protein sequences encoded in three complete bacterial genomes, those of Haemophilus influenzae, Mycoplasma genitalium and Synechocystis sp., and the first available archaeal genome sequence, that of Methanococcus jannaschii, were analysed using the BLAST2 algorithm and methods for amino acid motif detection. Between 75% and 90% of the predicted proteins encoded in each of the bacterial genomes and 73% of the M. jannaschii proteins showed significant sequence similarity to proteins from other species. The fraction of bacterial and archaeal proteins containing regions conserved over long phylogenetic distances is nearly the same and close to 70%. Functions of 70-85% of the bacterial proteins and about 70% of the archaeal proteins were predicted with varying precision. This contrasts with the previous report that more than half of the archaeal proteins have no homologues and shows that, with more sensitive methods and detailed analysis of conserved motifs, archaeal genomes become as amenable to meaningful interpretation by computer as bacterial genomes. The analysis of conserved motifs resulted in the prediction of a number of previously undetected functions of bacterial and archaeal proteins and in the identification of novel protein families. In spite of the generally high conservation of protein sequences, orthologues of 25% or less of the M. jannaschii genes were detected in each individual completely sequenced genome, supporting the uniqueness of archaea as a distinct domain of life. About 53% of the M. jannaschii proteins belong to families of paralogues, a fraction similar to that in bacteria with larger genomes, such as Synechocystis sp. and Escherichia coli, but higher than that in H. influenzae, which has approximately the same number of genes as M. jannaschii. Certain groups of proteins, e.g. molecular chaperones and DNA repair enzymes, thought to be ubiquitous and represented in the minimal gene set derived by bacterial genome comparison, are missing in M. jannaschii, indicating massive non-orthologous displacement of genes responsible for essential functions. An unexpectedly large fraction of the M. jannaschii gene products, 44%, shows significantly higher similarity to bacterial than to eukaryotic proteins, compared with 13% that have eukaryotic proteins as their closest homologues (the rest of the proteins show approximately the same level of similarity to bacterial and eukaryotic homologues or have no homologues). Proteins involved in translation, transcription, replication and protein secretion are most closely related to eukaryotic proteins, whereas metabolic enzymes, metabolite uptake systems, enzymes for cell wall biosynthesis and many uncharacterized proteins appear to be 'bacterial'. A similar prevalence of proteins of apparent bacterial origin was observed among the currently available sequences from the distantly related archaeal genus, Sulfolobus. It is likely that the evolution of archaea included at least one major merger between ancestral cells from the bacterial lineage and the lineage leading to the eukaryotic nucleocytoplasm.

Algorithms↗