Gene order is not conserved in bacterial evolution.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to A R Mushegian.
Explore the source record for details and available documents.
The availability of complete genome sequences of cellular life forms creates the opportunity to explore the functional content of the genomes and evolutionary relationships between them at a new qualitative level. With the advent of these sequences, the construction of a minimal gene set sufficient for sustaining cellular life and reconstruction of the genome of the last common ancestor of bacteria, eukaryotes, and archaea become realistic, albeit challenging, research projects. A version of the minimal gene set for modern-type cellular life derived by comparative analysis of two bacterial genomes, those of Haemophilus influenzae and Mycoplasma genitalium, consists of approximately 250 genes. A comparison of the protein sequences encoded in these genes with those of the proteins encoded in the complete yeast genome suggests that the last common ancestor of all extant life might have had an RNA genome.
Most of the genes involved in the development of multicellular eukaryotes encode large, multidomain proteins. To decipher the major trends in the evolution of these proteins and make functional predictions for uncharacterized domains, we applied a strategy of sequence database search that includes construction of specialized data sets and iterative subsequence masking. This computational approach allowed us to detect previously unnoticed but potentially important sequence similarities. Developmental gene products are enriched in predicted nonglobular regions as compared to unbiased sets of eukaryotic and bacterial proteins. Developmental genes that act intracellularly, primarily at the level of transcription regulation, typically code for proteins containing highly conserved DNA-binding domains, most of which appear to have evolved before the radiation of bacteria and eukaryotes. We identified bacterial homologues, namely a protein family that includes the Escherichia coli universal stress protein UspA, for the MADS-box transcription regulators previously described only in eukaryotes. We also show that the FUS6 family of eukaryotic proteins contains a putative DNA-binding domain related to bacterial helix-turn-helix transcription regulators. Developmental proteins that act extracellularly are less conserved and often do not have bacterial homologues. Nevertheless, several provocative similarities between different groups of such proteins were detected.
Explore the source record for details and available documents.
The DNA genome of caulimoviruses contains a set of essential genes: I (movement gene), IV (major capsid protein gene), V (reverse transcriptase gene), and VI (gene coding for a post-transcriptional activator of the expression of other virus genes). In peanut chlorotic streak caulimovirus (PCISV), three ORFs, A, B, and C, are located between genes I and IV. They are dissimilar to other caulimovirus ORFs. ORF VII of PCISV is a homolog of ORF VII of soybean chlorotic mottle caulimovirus (SoCMV), but is not similar to the nonconserved ORF VII in other caulimoviruses. The sequence complementary to a portion of tRNA(Met), thought to be essential for the priming of minus-strand DNA synthesis in caulimoviruses, is located within the coding sequence of ORF A. To explore the functional significance of ORFs VII, A, B, and C, various mutations were engineered into an infectious DNA clone of PCISV. ORFs VII and B are shown to be dispensable, while ORFs A and C are essential. ORF C is a possible functional equivalent of gene III in other caulimoviruses. Sequences within ORF A that are required for efficient priming of minus-strand synthesis are likely to extend beyond the 12-bp tRNA-binding site. Complete deletion of ORF VII was correlated with severe symptoms, notably with the necrosis of apical meristems. Significance of these observations for the understanding of replication and pathogenesis of plant pararetroviruses and for the improvement of caulimovirus-based expression vectors is discussed.
Using methods for database screening with individual protein sequences and alignment blocks, a conserved domain is delineated in a group of proteins including several FAD-dependent oxidases. Two motifs within this domain resemble phosphate-binding loops and may be directly involved in FAD binding. These motifs can be readily distinguished from previously described nucleotide-binding sites using a method for database screening with position-dependent weight matrices derived from alignment blocks. Unexpectedly, this group of known and predicted FAD-dependent oxidases includes the product of the DIMINUTO gene, which is involved in Arabidopsis development, and its homologues from man and Mycobacterium leprae.
Gene I of peanut chlorotic streak virus (PCISV), a caulimovirus, is homologous to gene I of other caulimoviruses and may encode a protein for virus movement. To evaluate the function of gene I, several mutations were created in this gene of an infectious, partially redundant clone of PCISV. Constructs with an in-frame deletion and a single amino acid substitution in gene I were not infectious. To test for replication of these mutants in primarily infected cells, an immunosorbent PCR technique was devised. Virus particles formed by mutants in plants were recovered by binding to antivirus antibodies on a solid matrix and DNase treated to discriminate against residual inoculum, and DNA of trapped virions was subjected to PCR amplification. Gene I mutants were shown to direct formation of encapsidated DNA as revealed by a PCR product. Control gene V mutants (reverse transcriptase essential for replication) did not yield a PCR product. Quantitative PCR allowed estimation of the proportion of cells initially infected by gene I mutants and the amount of extractable virus per cell. It is concluded that PCISV gene I encodes a movement protein and that the immunoselection-PCR technique is useful in studying subliminal virus infection in plants.
Viruses have developed successful strategies for propagation at the expense of their host cells. Efficient gene expression, genome multiplication, and invasion of the host are enabled by virus-encoded genetic elements, many of which are well characterized. Sequences derived from plant DNA and RNA viruses can be used to control expression of other genes in vivo. The main groups of plant virus genetic elements useful in genetic engineering are reviewed, including the signals for DNA-dependent and RNA-dependent RNA synthesis, sequences on the virus mRNAs that enable translational control, and sequences that control processing and intracellular sorting of virus proteins. Use of plant viruses as extrachromosomal expression vectors is also discussed, along with the issue of their stability.
RNAse H (RNH1 protein) from the trypanosomatid Crithidia fasciculata has a functionally uncharacterized N-terminal domain dispensable for the RNAse H activity. Using computer methods for database search and multiple alignment, we show that the N-terminal domains of RNH1 and its homologue encoded by a cDNA from chicken lens are related to the conserved domain in caulimovirus ORF VI product that facilitates translation of polycistronic virus RNA in plant cells. We hypothesize that the N-terminal domain of eukaryotic RNAse H performs an as yet uncharacterized regulatory function, possibly in mRNA translation or turnover.
Amino acid sequences of enzymes that catalyze hydrolysis or phosphorolysis of the N-glycosidic bond in nucleosides and nucleotides (nucleosidases and phosphoribosyltransferases) were explored using computer methods for database similarity search and multiple alignment. Two new families, each including bacterial and eukaryotic enzymes, were identified. Family I consists of Escherichia coli AMP hydrolase (Amn), uridine phosphorylase (Udp), purine phosphorylase (DeoD), uncharacterized proteins from E. coli and Bacteroides uniformis, and, unexpectedly, a group of plant stress-inducible proteins. It is hypothesized that these plant proteins have evolved from nucleosidases and may possess nucleosidase activity. The proteins in this new family contain 3 conserved motifs, one of which was found also in eukaryotic purine nucleosidases, where it corresponds to the nucleoside-binding site. Family II is comprised of bacterial and eukaryotic thymidine phosphorylases and anthranilate phosphoribosyltransferases, the relationship between which has not been suspected previously. Based on the known tertiary structure of E. coli thymidine phosphorylase, structural interpretation was given to the sequence conservation in this family. The highest conservation is observed in the N-terminal alpha-helical domain, whose exact function is not known. Parts of the conserved active site of thymidine phosphorylases and anthranilate phosphoribosyltransferases were delineated. A motif in the putative phosphate-binding site is conserved in family II and in other phosphoribosyltransferases. Our analysis suggests that certain enzymes of very similar specificity, e.g., uridine and thymidine phosphorylases, could have evolved independently. In contrast, enzymes catalyzing such different reactions as AMP hydrolysis and uridine phosphorolysis or thymidine phosphorolysis and phosphoribosyl anthranilate synthesis are likely to have evolved from common ancestors.
Using computer methods for multiple alignment, sequence motif search, and tertiary structure modeling, we show that eukaryotic translation elongation factor 1 gamma (EF1 gamma) contains an N-terminal domain related to class theta glutathione S-transferases (GST). GST-like proteins related to class theta comprise a large group including, in addition to typical GSTs and EF1 gamma, stress-induced proteins from bacteria and plants, bacterial reductive dehalogenases and beta-etherases, and several uncharacterized proteins. These proteins share 2 conserved sequence motifs with GSTs of other classes (alpha, mu, and pi). Tertiary structure modeling showed that in spite of the relatively low sequence similarity, the GST-related domain of EF1 gamma is likely to form a fold very similar to that in the known structures of class alpha, mu, and pi GSTs. One of the conserved motifs is implicated in glutathione binding, whereas the other motif probably is involved in maintaining the proper conformation of the GST domain. We predict that the GST-like domain in EF1 gamma is enzymatically active and that to exhibit GST activity, EF1 gamma has to form homodimers. The GST activity may be involved in the regulation of the assembly of multisubunit complexes containing EF1 and aminoacyl-tRNA synthetases by shifting the balance between glutathione, disulfide glutathione, thiol groups of cysteines, and protein disulfide bonds. The GST domain is a widespread, conserved enzymatic module that may be covalently or noncovalently complexed with other proteins. Regulation of protein assembly and folding may be 1 of the functions of GST.
Explore the source record for details and available documents.
Cell-to-cell movement is a crucial step in plant virus infection. In many viruses, the movement function is secured by specific virus-encoded proteins. Amino acid sequence comparisons of these proteins revealed a vast superfamily containing a conserved sequence motif that may comprise a hydrophobic interaction domain. This superfamily combines proteins of viruses belonging to all principal groups of positive-strand RNA viruses, as well as single-stranded DNA containing geminiviruses, double-stranded DNA-containing pararetroviruses (caulimoviruses and badnaviruses), and tospoviruses that have negative-strand RNA genomes with two ambisense segments. In several groups of positive-strand RNA viruses, the movement function is provided by the proteins encoded by the so-called triple gene block including two putative small membrane-associated proteins and a putative RNA helicase. A distinct type of movement proteins with very high content of proline is found in tymoviruses. It is concluded that classification of movement proteins based on comparison of their amino acid sequences does not correlate with the type of genome nucleic acid or with grouping of viruses based on phylogenetic analysis of replicative proteins or with the virus host range. Recombination between unrelated or distantly related viruses could have played a major role in the evolution of the movement function. Limited sequence similarities were observed between i) movement proteins of dianthoviruses and the MIP family of cellular integral membrane proteins, and ii) between movement proteins of bromoviruses and cucumoviruses and M1 protein of influenza viruses which is involved in nuclear export of viral ribonucleoproteins. It is hypothesized that all movement proteins of plant viruses may mediate hydrophobic interactions between viral and cellular macromolecules.
Explore the source record for details and available documents.
Amino acid sequences of plant virus proteins mediating cell-to-cell movement were compared to each other and to protein sequences in databases. Two families of movement proteins have been identified, the members of which show statistically significant sequence similarity. The first, larger family (I) encompasses the movement proteins of tobamo-, tobra-, caulimo- and comoviruses, apple chlorotic leaf spot virus (ACLSV) and geminiviruses with bipartite genomes. Thus this family includes viruses which move by two methods, those requiring the coat protein for the cell-to-cell spread (comoviruses) and those not having this requirement (tobamoviruses). The previously unsuspected relationship between the movement proteins of RNA and DNA viruses having no RNA stage in their life cycle (geminiviruses) suggested that their movement mechanisms might be similar. The second, smaller family (II) consists of the movement proteins of tricornaviruses (bromoviruses, cucumoviruses, alfalfa mosaic virus and tobacco streak virus) and dianthoviruses. Alignment of the sequences of family I movement proteins highlighted two motifs, centred at conserved Gly and Asp residues, respectively, which are assumed to be crucial for the movement protein function(s). Screening the amino acid sequence database revealed another conserved motif that is shared by a large subset of family I movement proteins (those of caulimo- and comoviruses, and ACLSV) and the family of cellular 90K heat shock proteins (HSP90). Based on the analogy to HSP90, it is speculated that many plant virus movement proteins may mediate virus transport in a chaperone-like manner.