PubMed Health⌕ Search

Biomedical subjects

N V Grishin

Publications and source records attributed to N V Grishin.

At least 19 recordsLinked to original sources

Homology between O-linked GlcNAc transferases and proteins of the glycogen phosphorylase superfamily.

The O-linked GlcNAc transferases (OGTs) are a recently characterized group of largely eukaryotic enzymes that add a single beta-N-acetylglucosamine moiety to specific serine or threonine hydroxyls. In humans, this process may be part of a sugar regulation mechanism or cellular signaling pathway that is involved in many important diseases, such as diabetes, cancer, and neurodegeneration. However, no structural information about the human OGT exists, except for the identification of tetratricopeptide repeats (TPR) at the N terminus. The locations of substrate binding sites are unknown and the structural basis for this enzyme's function is not clear. Here, remote homology is reported between the OGTs and a large group of diverse sugar processing enzymes, including proteins with known structure such as glycogen phosphorylase, UDP-GlcNAc 2-epimerase, and the glycosyl transferase MurG. This relationship, in conjunction with amino acid similarity spanning the entire length of the sequence, implies that the fold of the human OGT consists of two Rossmann-like domains C-terminal to the TPR region. A conserved motif in the second Rossmann domain points to the UDP-GlcNAc donor binding site. This conclusion is supported by a combination of statistically significant PSI-BLAST hits, consensus secondary structure predictions, and a fold recognition hit to MurG. Additionally, iterative PSI-BLAST database searches reveal that proteins homologous to the OGTs form a large and diverse superfamily that is termed GPGTF (glycogen phosphorylase/glycosyl transferase). Up to one-third of the 51 functional families in the CAZY database, a glycosyl transferase classification scheme based on catalytic residue and sequence homology considerations, can be unified through this common predicted fold. GPGTF homologs constitute a substantial fraction of known proteins: 0.4% of all non-redundant sequences and about 1% of proteins in the Escherichia coli genome are found to belong to the GPGTF superfamily.

Amino Acid Sequence↗

Genome trees constructed using five different approaches suggest new major bacterial clades.

BACKGROUND: The availability of multiple complete genome sequences from diverse taxa prompts the development of new phylogenetic approaches, which attempt to incorporate information derived from comparative analysis of complete gene sets or large subsets thereof. Such attempts are particularly relevant because of the major role of horizontal gene transfer and lineage-specific gene loss, at least in the evolution of prokaryotes. RESULTS: Five largely independent approaches were employed to construct trees for completely sequenced bacterial and archaeal genomes: i) presence-absence of genomes in clusters of orthologous genes; ii) conservation of local gene order (gene pairs) among prokaryotic genomes; iii) parameters of identity distribution for probable orthologs; iv) analysis of concatenated alignments of ribosomal proteins; v) comparison of trees constructed for multiple protein families. All constructed trees support the separation of the two primary prokaryotic domains, bacteria and archaea, as well as some terminal bifurcations within the bacterial and archaeal domains. Beyond these obvious groupings, the trees made with different methods appeared to differ substantially in terms of the relative contributions of phylogenetic relationships and similarities in gene repertoires caused by similar life styles and horizontal gene transfer to the tree topology. The trees based on presence-absence of genomes in orthologous clusters and the trees based on conserved gene pairs appear to be strongly affected by gene loss and horizontal gene transfer. The trees based on identity distributions for orthologs and particularly the tree made of concatenated ribosomal protein sequences seemed to carry a stronger phylogenetic signal. The latter tree supported three potential high-level bacterial clades,: i) Chlamydia-Spirochetes, ii) Thermotogales-Aquificales (bacterial hyperthermophiles), and ii) Actinomycetes-Deinococcales-Cyanobacteria. The latter group also appeared to join the low-GC Gram-positive bacteria at a deeper tree node. These new groupings of bacteria were supported by the analysis of alternative topologies in the concatenated ribosomal protein tree using the Kishino-Hasegawa test and by a census of the topologies of 132 individual groups of orthologous proteins. Additionally, the results of this analysis put into question the sister-group relationship between the two major archaeal groups, Euryarchaeota and Crenarchaeota, and suggest instead that Euryarchaeota might be a paraphyletic group with respect to Crenarchaeota. CONCLUSIONS: We conclude that, the extensive horizontal gene flow and lineage-specific gene loss notwithstanding, extension of phylogenetic analysis to the genome scale has the potential of uncovering deep evolutionary relationships between prokaryotic lineages.

Bacteria↗

Structure prediction and active site analysis of the metal binding determinants in gamma -glutamylcysteine synthetase.

gamma-Glultamylcysteine synthetase (gamma-GCS) catalyzes the first step in the de novo biosynthesis of glutathione. In trypanosomes, glutathione is conjugated to spermidine to form a unique cofactor termed trypanothione, an essential cofactor for the maintenance of redox balance in the cell. Using extensive similarity searches and sequence motif analysis we detected homology between gamma-GCS and glutamine synthetase (GS), allowing these proteins to be unified into a superfamily of carboxylate-amine/ammonia ligases. The structure of gamma-GCS, which was previously poorly understood, was modeled using the known structure of GS. Two metal-binding sites, each ligated by three conserved active site residues (n1: Glu-55, Glu-93, Glu-100; and n2: Glu-53, Gln-321, and Glu-489), are predicted to form the catalytic center of the active site, where the n1 site is expected to bind free metal and the n2 site to interact with MgATP. To elucidate the roles of the metals and their ligands in catalysis, these six residues were mutated to alanine in the Trypanosoma brucei enzyme. All mutations caused a substantial loss of activity. Most notably, E93A was able to catalyze the l-Glu-dependent ATP hydrolysis but not the peptide bond ligation, suggesting that the n1 metal plays an important role in positioning l-Glu for the reaction chemistry. The apparent K(m) values for ATP were increased for both the E489A and Q321A mutant enzymes, consistent with a role for the n2 metal in ATP binding and phosphoryl transfer. Furthermore, the apparent K(d) values for activation of E489A and Q321A by free Mg(2+) increased. Finally, substitution of Mn(2+) for Mg(2+) in the reaction rescued the catalytic deficits caused by both mutations, demonstrating that the nature of the metal ligands plays an important role in metal specificity.

Amino Acid Sequence↗

Treble clef finger--a functionally diverse zinc-binding structural motif.

Detection of similarity is particularly difficult for small proteins and thus connections between many of them remain unnoticed. Structure and sequence analysis of several metal-binding proteins reveals unexpected similarities in structural domains classified as different protein folds in SCOP and suggests unification of seven folds that belong to two protein classes. The common motif, termed treble clef finger in this study, forms the protein structural core and is 25-45 residues long. The treble clef motif is assembled around the central zinc ion and consists of a zinc knuckle, loop, beta-hairpin and an alpha-helix. The knuckle and the first turn of the helix each incorporate two zinc ligands. Treble clef domains constitute the core of many structures such as ribosomal proteins L24E and S14, RING fingers, protein kinase cysteine-rich domains, nuclear receptor-like fingers, LIM domains, phosphatidylinositol-3-phosphate-binding domains and His-Me finger endonucleases. The treble clef finger is a uniquely versatile motif adaptable for various functions. This small domain with a 25 residue structural core can accommodate eight different metal-binding sites and can have many types of functions from binding of nucleic acids, proteins and small molecules, to catalysis of phosphodiester bond hydrolysis. Treble clef motifs are frequently incorporated in larger structures or occur in doublets. Present analysis suggests that the treble clef motif defines a distinct structural fold found in proteins with diverse functional properties and forms one of the major zinc finger groups.

ADP-Ribosylation Factors↗

Mh1 domain of Smad is a degraded homing endonuclease.

Smad proteins are eukarytic transcription regulators in the TGF-beta signaling cascade. Using a combination of sequence and structure-based analyses, we argue that MH1 domain of Smad is homologous to the diverse His-Me finger endonuclease family enzymes. The similarity is particularly extensive with the I-PpoI endonuclease. In addition to the global fold similarities, both proteins possess a conserved motif of three cysteine residues and one histidine residue which form a zinc-binding site in I-PpoI. Sequence and structure conservation in the motif region strongly suggest that MH1 domain may also incorporate a metal ion in its structural core. MH1 of Smad3 and I-PpoI exhibit similar nucleic acid binding mode and interact with DNA major groove through an antiparallel beta-sheet. MH1 is an example of transcription regulator derived from the ancient enzymatic domain that lost its catalytic activity but retained DNA-binding sites.

Amino Acid Sequence↗

GGDEF domain is homologous to adenylyl cyclase.

The GGDEF domain is detected in many prokaryotic proteins, most of which are of unknown function. Several bacteria carry 12-22 different GGDEF homologues in their genomes. Conducting extensive profile-based searches, we detect statistically supported sequence similarity between GGDEF domain and adenylyl cyclase catalytic domain. From this homology, we deduce that the prokaryotic GGDEF domain is a regulatory enzyme involved in nucleotide cyclization, with the fold similar to that of the eukaryotic cyclase catalytic domain. This prediction correlates with the functional information available on two GGDEF-containing proteins, namely diguanylate cyclase and phosphodiesterase A of Acetobacter xylinum, both of which regulate the turnover of cyclic diguanosine monophosphate. Domain architecture analysis shows that GGDEF is typically present in multidomain proteins containing regulatory domains of signaling pathways or protein-protein interaction modules. Evolutionary tree analysis indicates that GGDEF/cyclase superfamily forms a large diversified cluster of orthologous proteins present in bacteria, archaea, and eukaryotes. Proteins 2001;42:210-216.

Acetobacter↗

KH domain: one motif, two folds.

The K homology (KH) module is a widespread RNA-binding motif that has been detected by sequence similarity searches in such proteins as heterogeneous nuclear ribonucleoprotein K (hnRNP K) and ribosomal protein S3. Analysis of spatial structures of KH domains in hnRNP K and S3 reveals that they are topologically dissimilar and thus belong to different protein folds. Thus KH motif proteins provide a rare example of protein domains that share significant sequence similarity in the motif regions but possess globally distinct structures. The two distinct topologies might have arisen from an ancestral KH motif protein by N- and C-terminal extensions, or one of the existing topologies may have evolved from the other by extension, displacement and deletion. C-terminal extension (deletion) requires ss-sheet rearrangement through the insertion (removal) of a ss-strand in a manner similar to that observed in serine protease inhibitors serpins. Current analysis offers a new look on how proteins can change fold in the course of evolution.

Amino Acid Motifs↗

The C-terminal domain of HPII catalase is a member of the type I glutamine amidotransferase superfamily.

Discovering distant evolutionary relationships between proteins requires detecting subtle similarities. Here we use a combination of sequence and structure analysis to show that the C-terminal domain of Escherichia coli HPII catalase with available spatial structure is a divergent member of the type I glutamine amidotransferase (GAT) superfamily. GAT-containing proteins include many biosynthetic enzymes such as E. coli carbamoyl phosphate synthetase and anthranilate synthase. Typical GAT domains have Rossmann fold-like topology and possess a catalytic triad similar to that of proteases. The C-terminal domain of HPII catalase has the GAT Rossmann fold but lacks the triad and therefore loses enzymatic activity. In addition, we detect significant sequence similarity between thiJ domains, some of which are known to have protease activity, and typical GAT proteins. Evolutionary tree analysis of the entire GAT superfamily indicates that the HPII catalase is more closely related to thiJ domains than to classical GAT domains and is likely to have evolved from a thiJ-like protein. This work illustrates the strength of sequence-based profile analysis techniques coupled with structural superpositions in developing an evolutionarily relevant classification of protein structures. Proteins 2001;42:230-236.

Amino Acid Motifs↗

Type II CAAX prenyl endopeptidases belong to a novel superfamily of putative membrane-bound metalloproteases.

In this article, a novel, large and diverse superfamily of putative membrane-bound proteins that includes the type II CAAX prenyl endopeptidases is described. The majority of the members of this superfamily are hypothetical proteins from bacteria and plants. Analysis of the conserved motifs, combined with available experimental data, suggests that these proteins are putative metal-dependent proteases that are potentially involved in protein and/or peptide modification and secretion.

Amino Acid Motifs↗

AL2CO: calculation of positional conservation in a protein sequence alignment.

MOTIVATION: Amino acid sequence alignments are widely used in the analysis of protein structure, function and evolutionary relationships. Proteins within a superfamily usually share the same fold and possess related functions. These structural and functional constraints are reflected in the alignment conservation patterns. Positions of functional and/or structural importance tend to be more conserved. Conserved positions are usually clustered in distinct motifs surrounded by sequence segments of low conservation. Poorly conserved regions might also arise from the imperfections in multiple alignment algorithms and thus indicate possible alignment errors. Quantification of conservation by attributing a conservation index to each aligned position makes motif detection more convenient. Mapping these conservation indices onto a protein spatial structure helps to visualize spatial conservation features of the molecule and to predict functionally and/or structurally important sites. Analysis of conservation indices could be a useful tool in detection of potentially misaligned regions and will aid in improvement of multiple alignments. RESULTS: We developed a program to calculate a conservation index at each position in a multiple sequence alignment using several methods. Namely, amino acid frequencies at each position are estimated and the conservation index is calculated from these frequencies. We utilize both unweighted frequencies and frequencies weighted using two different strategies. Three conceptually different approaches (entropy-based, variance-based and matrix score-based) are implemented in the algorithm to define the conservation index. Calculating conservation indices for 35522 positions in 284 alignments from SMART database we demonstrate that different methods result in highly correlated (correlation coefficient more than 0.85) conservation indices. Conservation indices show statistically significant correlation between sequentially adjacent positions i and i + j, where j < 13, and averaging of the indices over the window of three positions is optimal for motif detection. Positions with gaps display substantially lower conservation properties. We compare conservation properties of the SMART alignments or FSSP structural alignments to those of the ClustalW alignments. The results suggest that conservation indices should be a valuable tool of alignment quality assessment and might be used as an objective function for refinement of multiple alignments. AVAILABILITY: The C code of the AL2CO program and its pre-compiled versions for several platforms as well as the details of the analysis are freely available at ftp://iole.swmed.edu/pub/al2co/.

Algorithms↗

Structure and mechanism of homoserine kinase: prototype for the GHMP kinase superfamily.

BACKGROUND: Homoserine kinase (HSK) catalyzes an important step in the threonine biosynthesis pathway. It belongs to a large yet unique class of small metabolite kinases, the GHMP kinase superfamily. Members in the GHMP superfamily participate in several essential metabolic pathways, such as amino acid biosynthesis, galactose metabolism, and the mevalonate pathway. RESULTS: The crystal structure of HSK and its complex with ADP reveal a novel nucleotide binding fold. The N-terminal domain contains an unusual left-handed betaalphabeta unit, while the C-terminal domain has a central alpha-beta plait fold with an insertion of four helices. The phosphate binding loop in HSK is distinct from the classical P loops found in many ATP/GTP binding proteins. The bound ADP molecule adopts a rare syn conformation and is in the opposite orientation from those bound to the P loop-containing proteins. Inspection of the substrate binding cavity indicates several amino acid residues that are likely to be involved in substrate binding and catalysis. CONCLUSIONS: The crystal structure of HSK is the first representative in the GHMP superfamily to have determined structure. It provides insight into the structure and nucleotide binding mechanism of not only the HSK family but also a variety of enzymes in the GHMP superfamily. Such enzymes include galactokinases, mevalonate kinases, phosphomevalonate kinases, mevalonate pyrophosphate decarboxylases, and several proteins of yet unknown functions.

Adenosine Diphosphate↗

Accumulation of dietary cholesterol in sitosterolemia caused by mutations in adjacent ABC transporters.

In healthy individuals, acute changes in cholesterol intake produce modest changes in plasma cholesterol levels. A striking exception occurs in sitosterolemia, an autosomal recessive disorder characterized by increased intestinal absorption and decreased biliary excretion of dietary sterols, hypercholesterolemia, and premature coronary atherosclerosis. We identified seven different mutations in two adjacent, oppositely oriented genes that encode new members of the adenosine triphosphate (ATP)-binding cassette (ABC) transporter family (six mutations in ABCG8 and one in ABCG5) in nine patients with sitosterolemia. The two genes are expressed at highest levels in liver and intestine and, in mice, cholesterol feeding up-regulates expressions of both genes. These data suggest that ABCG5 and ABCG8 normally cooperate to limit intestinal absorption and to promote biliary excretion of sterols, and that mutated forms of these transporters predispose to sterol accumulation and atherosclerosis.

ATP Binding Cassette Transporter, Subfamily G, Mem↗

The synthetase domains of cobalamin biosynthesis amidotransferases cobB and cobQ belong to a new family of ATP-dependent amidoligases, related to dethiobiotin synthetase.

Phosphotransacetylases of Escherichia coli and several other bacteria contain an additional 350-aa N-terminal fragment that is not required for phosphotransacetylase activity. Sequence analysis of this fragment revealed that it is closely related to a family of ATP-dependent enzymes that also includes dethiobiotin synthetase and the synthetase domains of two amidotransferases involved in cobalamin biosynthesis, cobyrinic acid a,c-diamide synthase (CobB) and cobyric acid synthase (CobQ). Further database searches showed that this enzyme family is also related to the MinD family of ATPases involved in regulation of cell division in bacteria and archaea. Analysis of sequence conservation in the members of this enzyme family using the structure of dethiobiotin synthetase active site as a guide allowed us to suggest a model for the interaction of CobB and CobQ with their respective substrates. CobB and CobQ were also found to contain unusual Triad family (class I) glutamine amidotransferase domains with conserved Cys and His residues, but lacking the Glu residue of the catalytic triad. These results should help in understanding the enzymology of cobalamin biosynthesis and in resolving the role of phosphotransacetylase in regulation of the carbon flow to and from acetate.

Adenosine Triphosphatases↗

Common fold in helix-hairpin-helix proteins.

Helix-hairpin-helix (HhH) is a widespread motif involved in non-sequence-specific DNA binding. The majority of HhH motifs function as DNA-binding modules, however, some of them are used to mediate protein-protein interactions or have acquired enzymatic activity by incorporating catalytic residues (DNA glycosylases). From sequence and structural analysis of HhH-containing proteins we conclude that most HhH motifs are integrated as a part of a five-helical domain, termed (HhH)(2) domain here. It typically consists of two consecutive HhH motifs that are linked by a connector helix and displays pseudo-2-fold symmetry. (HhH)(2) domains show clear structural integrity and a conserved hydrophobic core composed of seven residues, one residue from each alpha-helix and each hairpin, and deserves recognition as a distinct protein fold. In addition to known HhH in the structures of RuvA, RadA, MutY and DNA-polymerases, we have detected new HhH motifs in sterile alpha motif and barrier-to-autointegration factor domains, the alpha-subunit of Escherichia coli RNA-polymerase, DNA-helicase PcrA and DNA glycosylases. Statistically significant sequence similarity of HhH motifs and pronounced structural conservation argue for homology between (HhH)(2) domains in different protein families. Our analysis helps to clarify how non-symmetric protein motifs bind to the double helix of DNA through the formation of a pseudo-2-fold symmetric (HhH)(2) functional unit.

Amino Acid Sequence↗

Crystal structure of YbaK protein from Haemophilus influenzae (HI1434) at 1.8 A resolution: functional implications.

Structural genomics of proteins of unknown function most straightforwardly assists with assignment of biochemical activity when the new structure resembles that of proteins whose functions are known. When a new fold is revealed, the universe of known folds is enriched, and once the function is determined by other means, novel structure-function relationships are established. The previously unannotated protein HI1434 from H. influenzae provides a hybrid example of these two paradigms. It is a member of a microbial protein family, labeled in SwissProt as YbaK and ebsC. The crystal structure at 1.8 A resolution reported here reveals a fold that is only remotely related to the C-lectin fold, in particular to endostatin, and thus is not sufficiently similar to imply that YbaK proteins are saccharide binding proteins. However, a crevice that may accommodate a small ligand is evident. The putative binding site contains only one invariant residue, Lys46, which carries a functional group that could play a role in catalysis, indicating that YbaK is probably not an enzyme. Detailed sequence analysis, including a number of newly sequenced microbial organisms, highlights sequence homology to an insertion domain in prolyl-tRNA synthetases (proRS) from prokaryote, a domain whose function is unknown. A HI1434-based model of the insertion domain shows that it should also contain the putative binding site. Being part of a tRNA synthetases, the insertion domain is likely to be involved in oligonucleotide binding, with possible roles in recognition/discrimination or editing of prolyl-tRNA. By analogy, YbaK may also play a role in nucleotide or oligonucleotide binding, the nature of which is yet to be determined.

Amino Acid Sequence↗

C-terminal domains of Escherichia coli topoisomerase I belong to the zinc-ribbon superfamily.

Detection of remote evolutionary connections is increasingly difficult with sequence and structural divergence. A combination of sequence and structural analysis, in which statistically supported sequence similarity had a crucial impact, revealed that Escherichia coli topoisomerase I C-terminal fragment is evolutionarily related to the three tetracysteine zinc-binding domains of the enzyme. Spatial structure analysis of this C-terminal fragment indicates that it consists of two structurally similar domains and suggests homology between them. Sequence similarity between the zinc-binding domains of type Ia topoisomerases and transcription regulators of known spatial structure helps to conclude that E. coli topo I contains five copies of a zinc ribbon domain at the C terminus. Two of these domains, corresponding to the C-terminal fragment, lost their cysteine residues and are probably not able to bind zinc. Present analyses lead to the classification of the C-terminal fragment of E. coli topoisomerase I as a member of zinc ribbon superfamily, despite the absence of zinc-binding sites.

Amino Acid Sequence↗

Estimating the number of protein folds and families from complete genome data.

Using the data on proteins encoded in complete genomes, combined with a rigorous theory of the sampling process, we estimate the total number of protein folds and families, as well as the number of folds and families in each genome. The total number of folds in globular, water- soluble proteins is estimated at about 1000, with structural information currently available for about one-third of the number. The sequenced genomes of unicellular organisms encode from approximately 25%, for the minimal genomes of the Mycoplasmas, to 70-80% for larger genomes, such as Escherichia coli and yeast, of the total number of folds. The number of protein families with significant sequence conservation was estimated to be between 4000 and 7000, with structures available for about 20% of these.

Conserved Sequence↗

Two tricks in one bundle: helix-turn-helix gains enzymatic activity.

Many examples of enzymes that have lost their catalytic activity and perform other biological functions are known. The opposite situation is rare. A previously unnoticed structural similarity between the lambda integrase family (Int) proteins and the AraC family of transcriptional activators implies that the Int family evolved by duplication of an ancient DNA-binding homeodomain-like module, which acquired enzymatic activity. The two helix-turn-helix (HTH) motifs in Int proteins incorporate catalytic residues and participate in DNA binding. The active site of Int proteins, which include the type IB topoisomerases, is formed at the domain interface and the catalytic tyrosine residue is located in the second helix of the C-terminal HTH motif. Structural analysis of other 'tyrosine' DNA-breaking/rejoining enzymes with similar enzyme mechanisms, namely prokaryotic topoisomerase I, topoisomerase II and archaeal topoisomerase VI, reveals that the catalytic tyrosine is placed in a HTH domain as well. Surprisingly, the location of this tyrosine residue in the structure is not conserved, suggesting independent, parallel evolution leading to the same catalytic function by homologous HTH domains. The 'tyrosine' recombinases give a rare example of enzymes that evolved from ancient DNA-binding modules and present a unique case for homologous enzymatic domains with similar catalytic mechanisms but different locations of catalytic residues, which are placed at non-homologous sites.

Amino Acid Sequence↗