PubMed HealthSearch

Biomedical subjects

E V Koonin

Publications and source records attributed to E V Koonin.

At least 19 recordsLinked to original sources

Pineal serotonin N-acetyltransferase: expression cloning and molecular analysis.

Pineal serotonin N-acetyltransferase (arylalkylamine N-acetyltransferase, or AA-NAT) generates the large circadian rhythm in melatonin, the hormone that coordinates daily and seasonal physiology in some mammals. Complementary DNA encoding ovine AA-NAT was cloned. The abundance of AA-NAT messenger RNA (mRNA) during the day was high in the ovine pineal gland and somewhat lower in retina. AA-NAT mRNA was found unexpectedly in the pituitary gland and in some brain regions. The night-to-day ratio of ovine pineal AA-NAT mRNA is less than 2. In contrast, the ratio exceeds 150 in rats. AA-NAT represents a family within a large superfamily of acetyltransferases.

Amino Acid Sequence

The chromo superfamily: new members, duplication of the chromo domain and possible role in delivering transcription regulators to chromatin.

Using computer methods for detecting conserved amino acid sequence motifs, we show that the chromatin organization modifier (chromo) domain that has been previously identified in several proteins involved in transcription down-regulation is present in a much larger group of (putative) chromatin-binding proteins, some of which are positive rather than negative regulators of transcription. The most interesting new members of the chromo superfamily are Drosophila male-specific lethal (MSL-3) protein involved in the X chromosome gene dosage compensation in the males and human retinoblastoma-binding protein RBP-1. We show that the chromo domain is duplicated in several chromatin-binding proteins and use this observation to interpret recent results on chromatin binding obtained with chimeric chromo domain-containing proteins. We hypothesize that the chromo domain may be a vehicle that delivers both positive and negative transcription regulators to the sites of their action on chromatin.

Amino Acid Sequence

Bacteriophage P2: genes involved in baseplate assembly.

The sequences of two previously defined tail genes, V and J, of the temperate bacteriophage P2, and those of two new essential tail genes, W and I, were determined. Their order is the late gene promoter, VWJI, followed by the tail fiber genes H and G, and a transcription terminator. The V gene product is the small spike at the tip of the tail, and the J gene product lies at the edge of the baseplate. The W gene product may be homologous to the product of gene 25 of T4 phage, which is part of the T4 baseplate. A temperature-sensitive mutation in gene V affects satellite phage P4 production more than it affects the production of P2 helper phage. P4 mutations that partially compensate for this defect of gene V lie in the P4 capsid size determination gene, sid.

Amino Acid Sequence

Detection of new genes in a bacterial genome using Markov models for three gene classes.

We further investigated the statistical features of the three classes of Escherichia coli genes that have been previously delineated by factorial correspondence analysis and dynamic clustering methods. A phased Markov model for a nucleotide sequence of each gene class was developed and employed for gene prediction using the GeneMark program. The protein-coding region prediction accuracy was determined for class-specific Markov models of different orders when the programs implementing these models were applied to gene sequences from the same or other classes. It is shown that at least two training sets and two program versions derived for different classes of E. coli genes are necessary in order to achieve a high accuracy of coding region prediction for uncharacterized sequences. Some annotated E. coli genes from Class I and Class III are shown to be spurious, whereas many open reading frames (ORFs) that have not been annotated in GenBank as genes are predicted to encode proteins. The amino acid sequences of the putative products of these ORFs initially did not show similarity to already known proteins. However, conserved regions have been identified in several of them by screening the latest entries in protein sequence databases and applying methods for motif search, while some other of these new genes have been identified in independent experiments.

Algorithms

Male-specific lethal 2, a dosage compensation gene of Drosophila, undergoes sex-specific regulation and encodes a protein with a RING finger and a metallothionein-like cysteine cluster.

In Drosophila the equalization of X-linked gene products between males and females, i.e. dosage compensation, is the result of a 2-fold hypertranscription of most of these genes in males. At least four regulatory genes are required for this process. Three of these genes, maleless (mle), male-specific lethal 1 (msl-1) and male-specific lethal 3 (msl-3), have been cloned and their products have been shown to interact and to bind to numerous sites on the X chromosome of males, but not of females. Although binding to the X chromosome is negatively correlated with the function of the master regulatory gene Sex lethal (Sxl), the mechanisms that restrict this binding to males and to the X chromosome are not yet understood. We have cloned the last of the known autosomal genes involved in dosage compensation, male-specific lethal 2 (msl-2), and characterized its product. The encoded protein (MSL-2) consists of 769 amino acid residues and has a RING finger (C3HC4 zinc finger) and a metallothionein-like domain with eight conserved and two non-conserved cysteines. In addition, it contains a positively and a negatively charged amino acid residue cluster and a coiled coil domain that may be involved in protein-protein interactions. Males produce a msl-2 transcript that is shorter than in females, due to differential splicing of an intron of 132 bases in the untranslated leader. Using an antiserum against MSL-2 we have shown that the protein is expressed at a detectable level only in males, where it is physically associated with the X chromosome. Our observations suggest that MSL-2 may be the target of the master regulatory gene Sxl and provide the basic elements of a working hypothesis on the function of MSL-2 in mediating the 2-fold increase in transcription that is characteristic of dosage compensation.

Amino Acid Sequence

Identification of the primase active site of the herpes simplex virus type 1 helicase-primase.

Herpes simplex virus type 1 (HSV-1) encodes a heterotrimeric helicase-primase composed of the products of the three DNA replication-specific genes UL5, UL8, and UL52 (Crute, J. J., and Lehman, I. R. (1991) J. Biol. Chem. 266, 4484-4488). The UL5 and UL52 products constitute a heterodimeric subassembly of the holoenzyme that contains both helicase and primase activities (Calder, J. M., and Stow, N. D. (1990) Nucleic Acids Res. 18, 3573-3578; Dodson, M. S., and Lehman, I. R. (1991) Proc. Natl. Acad. Sci. U. S. A. 88, 1105-1109). The role of the UL52 product in the active HSV-1 helicase-primase was examined. A sequence located between residues 610 and 636 on the UL52 protein was found to be conserved among the UL52 homologues of eight herpesviruses. The carboxyl-terminal portion of this conserved sequence consisted of two Asp residues separated by a variable hydrophobic amino acid residue and is analogous to the divalent metal-binding site of DNA polymerases and several DNA primases. This motif has been designated the herpesvirus primase DXD motif. To study the role of the HSV-1 primase DXD motif in primase action, three site-directed changes were introduced into the UL52 gene. The helicase activity of the recombinant holoenzymes was unaffected by any of the introduced changes. Changing either of the two Asp residues that constitute the divalent metal-binding site (Asp628 or Asp630) to Ala dramatically reduced the primase activity of the HSV-1 helicase-primase holoenzyme in vitro, whereas alteration of the nearby conserved residue Asn624 to Gly had minimal effect. Therefore, in the three-subunit HSV-1 helicase-primase, the UL52 product provides at least a part of the primase catalytic site.

Amino Acid Sequence

Complete sequence of the citrus tristeza virus RNA genome.

The sequence of the entire genome of citrus tristeza virus (CTV), Florida isolate T36, was completed. The 19,296-nt CTV genome encodes 12 open reading frames (ORFs) potentially coding for at least 17 protein products. The 5'-proximal ORF 1a starts at nucleotide 108 and encodes a large polyprotein with calculated MW of 349 kDa containing domains characteristic of (from 5' to 3') two papain-like proteases (P-PRO), a methyltransferase (MT), and a helicase (HEL). Alignment of the putative P-PRO sequences of CTV with the related proteases of beet yellows closterovirus (BYV) and potyviruses allowed the prediction of catalytic cysteine and histidine residues as well as two cleavage sites, namely Val-Gly/Gly for the 5' proximal P-PRO domain and Met-Gly/Gly for the 5' distal P-PRO domain. The autoproteolytic cleavage of the polyprotein at these sites would release two N-terminal leader proteins of 54 and 55 kDa, respectively, and a 240-kDa C-terminal fragment containing MT and HEL domains. The apparent duplication of the leader domain distinguishes CTV from BYV and accounts for most of the size increase in the ORF 1a product of CTV. The downstream ORF 1b encodes a 57-kDa putative RNA-dependent RNA polymerase (RdRp), which is probably expressed via a +1 ribosomal frameshift. Sequence analysis of the frameshift region suggests that this +1 frameshift probably occurs at a rare arginine codon CGG and that elements of the RNA secondary structure are unlikely to be involved in this process. The complete polyprotein resulting from this frameshift event has a calculated MW of 401 kDa and after cleavage of the two N-terminal leaders would yield a 292-kDa protein containing the MT, HEL, and RdRp domains. Phylogenetic analysis of the three replication-associated domains, MT, HEL, and RdRp, indicates that CTV and BYV form a separate closterovirus lineage within the alpha-like supergroup of positive-strand RNA viruses. Two gene blocks or modules can be easily identified in the CTV genome. The first includes the replicative MT, HEL, and RdRp genes and is conserved throughout the entire alpha-like superfamily. The second block consists of five ORFs, 3 to 7, conserved among closteroviruses, including genes for the CTV homolog of HSP70 proteins and a duplicate of the coat protein gene. The 3'-terminal ORFs 8 to 11 encode a putative RNA-binding protein (ORF 11), and three proteins with unknown functions; this gene array is poorly conserved among closteroviruses.(ABSTRACT TRUNCATED AT 400 WORDS)

Amino Acid Sequence

A putative FAD-binding domain in a distinct group of oxidases including a protein involved in plant development.

Using methods for database screening with individual protein sequences and alignment blocks, a conserved domain is delineated in a group of proteins including several FAD-dependent oxidases. Two motifs within this domain resemble phosphate-binding loops and may be directly involved in FAD binding. These motifs can be readily distinguished from previously described nucleotide-binding sites using a method for database screening with position-dependent weight matrices derived from alignment blocks. Unexpectedly, this group of known and predicted FAD-dependent oxidases includes the product of the DIMINUTO gene, which is involved in Arabidopsis development, and its homologues from man and Mycobacterium leprae.

Amino Acid Sequence

The cytidylyltransferase superfamily: identification of the nucleotide-binding site and fold prediction.

The crystal structure of glycerol-3-phosphate cytidylyltransferase from B. subtilis (TagD) is about to be solved. Here, we report a testable structure prediction based on the identification by sequence analysis of a superfamily of functionally diverse but structurally similar nucleotide-binding enzymes. We predict that TagD is a member of this family. The most conserved region in this superfamily resembles the ATP-binding HiGH motif of class I aminoacyl-tRNA synthetases. The predicted secondary structure of cytidylyltransferase and its homologues is compatible with the alpha/beta topography of the class I aminoacyl-tRNA synthetases. The hypothesis of similarity of fold is strengthened by sequence-structure alignment and 3D model building using the known structure of tyrosyl tRNA synthetase as template. The proposed 3D model of TagD is plausible both structurally, with a well packed hydrophobic core, and functionally, as the most conserved residues cluster around the putative nucleotide binding site. If correct, the model would imply a very ancient evolutionary link between class I tRNA synthetases and the novel cytidylyltransferase superfamily.

Amino Acid Sequence

Identification and properties of the largest subunit of the DNA-dependent RNA polymerase of fish lymphocystis disease virus: dramatic difference in the domain organization in the family Iridoviridae.

Cytoplasmic DNA viruses encode a DNA-dependent RNA polymerase (DdRP) that is essential for transcription of viral genes. The amino acid sequences of the known largest subunits of DdRPs from different species contain highly conserved regions. Oligonucleotide primers, deduced from two conserved domains (RQP[T/S]LH and NADFDGDE) were used for detecting the corresponding gene of fish lymphocystis disease virus (FLCDV), a member of the family Iridoviridae, which replicates in the cytoplasm of infected cells of flatfish. The gene coding for the largest subunit of the DdRP was identified using a PCR-derived probe. The screening of the complete EcoRI gene library of the viral genome led to the identification of the gene locus of the largest subunit of the DdRP within the EcoRI DNA fragment B (12.4 kbp, 0.034 to 0.165 map units). The nucleotide sequence of a part (8334 bp) of the EcoRI DNA fragment B was determined and a large ORF on the lower strand (ATG = 5787; TAA = 2190) was detected which encodes a protein of 1199 amino acids. Comparison of the amino acid sequences of the largest subunits of the DdRP (RPO1) of FLCDV and Chilo iridescent virus (CIV) revealed a dramatic difference in their domain organization. Unlike the 1051 aa RPO1 of CIV, which lacks the C-terminal domain conserved in eukaryotic, eubacterial and other viral RNA polymerases, the 1199 aa RPO1 of FLCDV is fully collinear with its cellular and viral homologues. Despite this difference, comparative analysis of the amino acid sequences of viral and cellular RNA polymerases suggests a common origin for the largest RNA polymerase subunits of FLCDV and CIV.

Amino Acid Sequence

Plasmodium falciparum protein associated with the invasion junction contains a conserved oxidoreductase domain.

The merozoite cap protein-1 (MCP-1) of Plasmodium falciparum follows the distribution of the moving junction during invasion of erythrocytes. We have cloned the gene encoding this protein from a cDNA library using a monoclonal antibody. The protein lacks a signal sequence and has no predicted transmembrane domains; none of the antisera reacts with the surfaces of intact merozoites, indicating that the cap distribution is submembranous. MCP-1 is divided into three domains. The N-terminal domain includes a 52-amino-acid region that is highly conserved in a large family of bacterial and eukaryotic proteins. Based on the known functions of two proteins of this family and the pattern of amino acid conservation, it is predicted that this domain may possess oxido-reductase activity, since the active cysteine residue of this domain is invariant in all proteins of the family. The other two domains of MCP-1 are not found in any other members of this protein family and may reflect the specific function of MCP-1 in invasion. The middle domain is negatively charged and enriched in glutamate; the C-terminal domain is positively charged and enriched in lysine. By virtue of its positive charge, the C-terminal domain resembles domains in some cytoskeleton-associated proteins and may mediate the interaction of MCP-1 with cytoskeleton in Plasmodium.

Amino Acid Sequence

The gene for the longest known Escherichia coli protein is a member of helicase superfamily II.

The Escherichia coli rnt gene, which encodes the RNA-processing enzyme RNase T, is cotranscribed with a downstream gene. Complete sequencing of this gene indicates that its coding region encompasses 1,538 amino acids, making it the longest known protein in E. coli. The gene (tentatively termed lhr for long helicase related) contains the seven conserved motifs of the DNA and RNA helicase superfamily II. An approximately 170-kDa protein is observed by sodium dodecyl sulfate-polyacrylamide gel electrophoresis of 35S-labeled extracts prepared from cells in which lhr is under the control of an induced T7 promoter. This protein is absent when lhr is interrupted or when no plasmid is present. Downstream of lhr is the C-terminal region of a convergent gene with homology to glutaredoxin. Interruptions of chromosomal lhr at two different positions within the gene do not affect the growth of E. coli at various temperatures in rich or minimal medium, indicating that lhr is not essential for usual laboratory growth. lhr interruption also has no effect on anaerobic growth. In addition, cells lacking Lhr recover normally from starvation, plate phage normally, and display normal sensitivities to UV irradiation and H2O2. Southern analysis showed that no other gene closely related to lhr is present on the E. coli chromosome. These data expand the known size range of E. coli proteins and suggest that very large helicases are present in this organism.

Amino Acid Sequence

A novel RNA-binding motif in omnipotent suppressors of translation termination, ribosomal proteins and a ribosome modification enzyme?

Using computer methods for database search, multiple alignment, protein sequence motif analysis and secondary structure prediction, a putative new RNA-binding motif was identified. The novel motif is conserved in yeast omnipotent translation termination suppressor SUP1, the related DOM34 protein and its pseudogene homologue; three groups of eukaryotic and archaeal ribosomal proteins, namely L30e, L7Ae/S6e and S12e; an uncharacterized Bacillus subtilis protein related to the L7A/S6e group; and Escherichia coli ribosomal protein modification enzyme RimK. We hypothesize that a new type of RNA-binding domain may be utilized to deliver additional activities to the ribosome.

Amino Acid Sequence

Eukaryotic translation elongation factor 1 gamma contains a glutathione transferase domain--study of a diverse, ancient protein superfamily using motif search and structural modeling.

Using computer methods for multiple alignment, sequence motif search, and tertiary structure modeling, we show that eukaryotic translation elongation factor 1 gamma (EF1 gamma) contains an N-terminal domain related to class theta glutathione S-transferases (GST). GST-like proteins related to class theta comprise a large group including, in addition to typical GSTs and EF1 gamma, stress-induced proteins from bacteria and plants, bacterial reductive dehalogenases and beta-etherases, and several uncharacterized proteins. These proteins share 2 conserved sequence motifs with GSTs of other classes (alpha, mu, and pi). Tertiary structure modeling showed that in spite of the relatively low sequence similarity, the GST-related domain of EF1 gamma is likely to form a fold very similar to that in the known structures of class alpha, mu, and pi GSTs. One of the conserved motifs is implicated in glutathione binding, whereas the other motif probably is involved in maintaining the proper conformation of the GST domain. We predict that the GST-like domain in EF1 gamma is enzymatically active and that to exhibit GST activity, EF1 gamma has to form homodimers. The GST activity may be involved in the regulation of the assembly of multisubunit complexes containing EF1 and aminoacyl-tRNA synthetases by shifting the balance between glutathione, disulfide glutathione, thiol groups of cysteines, and protein disulfide bonds. The GST domain is a widespread, conserved enzymatic module that may be covalently or noncovalently complexed with other proteins. Regulation of protein assembly and folding may be 1 of the functions of GST.

Amino Acid Sequence

A P-loop-like motif in a widespread ATP pyrophosphatase domain: implications for the evolution of sequence motifs and enzyme activity.

A conserved amino acid sequence motif was identified in four distinct groups of enzymes that catalyze the hydrolysis of the alpha-beta phosphate bond of ATP, namely GMP synthetases, argininosuccinate synthetases, asparagine synthetases, and ATP sulfurylases. The motif is also present in Rhodobacter capsulata AdgA, Escherichia coli NtrL, and Bacillus subtilis OutB, for which no enzymatic activities are currently known. The observed pattern of amino acid residue conservation and predicted secondary structures suggest that this motif may be a modified version of the P-loop of nucleotide binding domains, and that it is likely to be involved in phosphate binding. We call it PP-motif, since it appears to be a part of a previously uncharacterized ATP pyrophophatase domain. ATP sulfurylases, NtrL, and OutB consist of this domain alone. In other proteins, the pyrophosphatase domain is associated with amidotransferase domains (type I or type II), a putative citrulline-aspartate ligase domain or a nitrilase/amidase domain. Unexpectedly, statistically significant overall sequence similarity was found between ATP sulfurylase and 3'-phosphoadenosine 5'-phosphosulfate (PAPS) reductase, another protein of the sulfate activation pathway. The PP-motif is strongly modified in PAPS reductases, but they share with ATP sulfurylases another conserved motif which might be involved in sulfate binding. We propose that PAPS reductases may have evolved from ATP sulfurylases; the evolution of the new enzymatic function appears to be accompanied by a switch of the strongest functional constraint from the PP-motif to the putative sulfate-binding motif.

Amino Acid Sequence