PubMed Health⌕ Search

Biomedical subjects

Arcady Mushegian

Publications and source records attributed to Arcady Mushegian.

At least 19 recordsLinked to original sources

Identification and Characterization of a Schizosaccharomyces pombe RNA Polymerase II Elongation Factor with Similarity to the Metazoan Transcription Factor ELL.

ELL family transcription factors activate the rate of transcript elongation by suppressing transient pausing by RNA polymerase II at many sites along the DNA. ELL-associated factors 1 and 2 (EAF1 and EAF2) bind stably to ELL family members and act as strong positive regulators of their transcription activities. Orthologs of ELL and EAF have been identified in metazoa, but it has been unclear whether such RNA polymerase II elongation factors are utilized in lower eukaryotes. Using bioinformatic and biochemical approaches, we have identified a new Schizosaccharomyces pombe RNA polymerase II elongation factor that is composed of two subunits designated SpELL and SpEAF, which share weak sequence similarity with members of the metazoan ELL and EAF families. Like mammalian ELL-EAF, SpELL-SpEAF stimulates RNA polymerase II transcription elongation and pyrophosphorolysis. In addition, like many yeast RNA polymerase II elongation factors, deletion of the SpELL gene renders S. pombe sensitive to the drug 6-azauracil. Finally, phylogenetic analyses suggest that the SpELL and SpEAF proteins are evolutionarily conserved in many fungi but not in Saccharomyces cerevisiae.

Amino Acid Sequence↗

The genome of the sea urchin Strongylocentrotus purpuratus.

We report the sequence and analysis of the 814-megabase genome of the sea urchin Strongylocentrotus purpuratus, a model for developmental and systems biology. The sequencing strategy combined whole-genome shotgun and bacterial artificial chromosome (BAC) sequences. This use of BAC clones, aided by a pooling strategy, overcame difficulties associated with high heterozygosity of the genome. The genome encodes about 23,300 genes, including many previously thought to be vertebrate innovations or known only outside the deuterostomes. This echinoderm genome provides an evolutionary outgroup for the chordates and yields insights into the evolution of deuterostomes.

Animals↗

A complex oscillating network of signaling genes underlies the mouse segmentation clock.

The segmental pattern of the spine is established early in development, when the vertebral precursors, the somites, are rhythmically produced from the presomitic mesoderm. Microarray studies of the mouse presomitic mesoderm transcriptome reveal that the oscillator associated with this process, the segmentation clock, drives the periodic expression of a large network of cyclic genes involved in cell signaling. Mutually exclusive activation of the notch-fibroblast growth factor and Wnt pathways during each cycle suggests that coordinated regulation of these three pathways underlies the clock oscillator.

Algorithms↗

The sea urchin kinome: a first look.

This paper reports a preliminary in silico analysis of the sea urchin kinome. The predicted protein kinases in the sea urchin genome were identified, annotated and classified, according to both function and kinase domain taxonomy. The results show that the sea urchin kinome, consisting of 353 protein kinases, is closer to the Drosophila kinome (239) than the human kinome (518) with respect to total kinase number. However, the diversity of sea urchin kinases is surprisingly similar to humans, since the urchin kinome is missing only 4 of 186 human subfamilies, while Drosophila lacks 24. Thus, the sea urchin kinome combines the simplicity of a non-duplicated genome with the diversity of function and signaling previously considered to be vertebrate-specific. More than half of the sea urchin kinases are involved with signal transduction, and approximately 88% of the signaling kinases are expressed in the developing embryo. These results support the strength of this nonchordate deuterostome as a pivotal developmental and evolutionary model organism.

Animals↗

Thermus thermophilus bacteriophage phiYS40 genome and proteomic characterization of virions.

We determined the sequence of the 152,372 bp genome of phiYS40, a lytic tailed bacteriophage of Thermus thermophilus. The genome contains 170 putative open reading frames and three tRNA genes. Functions for 25% of phiYS40 gene products were predicted on the basis of similarity to proteins of known function from diverse phages and bacteria. phiYS40 encodes a cluster of proteins involved in nucleotide salvage, such as flavin-dependent thymidylate synthase, thymidylate kinase, ribonucleotide reductase, and deoxycytidylate deaminase, and in DNA replication, such as DNA primase, helicase, type A DNA polymerase, and predicted terminal protein involved in initiation of DNA synthesis. The structural genes of phiYS40, most of which have no similarity to sequences in public databases, were identified by mass spectrometric analysis of purified virions. Various phiYS40 proteins have different phylogenetic neighbors, including myovirus, podovirus, and siphovirus gene products, bacterial genes and, in one case, a dUTPase from a eukaryotic virus. phiYS40 has apparently arisen through multiple acts of recombination between different phage genomes as well as through acquisition of bacterial genes.

Amino Acid Sequence↗

Intermediary metabolism in sea urchin: the first inferences from the genome sequence.

The genome sequence of the purple sea urchin Strongylocentrotus purpuratus recently became available. We report the results of functional annotation and initial analysis of more than 2300 proteins predicted to be involved in metabolite transport and enzymatic conversion in sea urchin. The comparison of various reconstructed biosynthetic and catabolic pathways in sea urchin to those known in other genomes suggests the overall similarity of the sea urchin metabolism to that of the vertebrates, with relatively small but non-trivial differences from both vertebrates and protostomes. There are several examples of two parallel, non-orthologous solutions for the same molecular function in sea urchin, in contrast with the other completely sequenced metazoans that tend to contain just one version of the same function. There are also genes that appear to be close phylogenetic neighbors of plant or bacterial homologs, as opposed to homologs in other Metazoa. The evolutionary and functional significance of these variations is discussed.

Amino Acids↗

The Sad1-UNC-84 homology domain in Mps3 interacts with Mps2 to connect the spindle pole body with the nuclear envelope.

The spindle pole body (SPB) is the sole site of microtubule nucleation in Saccharomyces cerevisiae; yet, details of its assembly are poorly understood. Integral membrane proteins including Mps2 anchor the soluble core SPB in the nuclear envelope. Adjacent to the core SPB is a membrane-associated SPB substructure known as the half-bridge, where SPB duplication and microtubule nucleation during G1 occurs. We found that the half-bridge component Mps3 is the budding yeast member of the SUN protein family (Sad1-UNC-84 homology) and provide evidence that it interacts with the Mps2 C terminus to tether the half-bridge to the core SPB. Mutants in the Mps3 SUN domain or Mps2 C terminus have SPB duplication and karyogamy defects that are consistent with the aberrant half-bridge structures we observe cytologically. The interaction between the Mps3 SUN domain and Mps2 C terminus is the first biochemical link known to connect the half-bridge with the core SPB. Association with Mps3 also defines a novel function for Mps2 during SPB duplication.

Amino Acid Sequence↗

Molecular dissection of arginyltransferases guided by similarity to bacterial peptidoglycan synthases.

Post-translational protein arginylation is essential for cardiovascular development and angiogenesis in mice and is mediated by arginyl-transfer RNA-protein transferases Ate1-a functionally conserved but poorly understood class of enzymes. Here, we used sequence analysis to detect the evolutionary relationship between the Ate1 family and bacterial FemABX family of aminoacyl-tRNA-peptide transferases, and to predict the functionally important residues in arginyltransferases, which were then used to construct a panel of mutants for further molecular dissection of mouse Ate1. Point mutations of the residues in the predicted regions of functional importance resulted in changes in enzymatic activity, including complete inactivation of mouse Ate1; other mutations altered its substrate specificity. Our results provide the first insights into the mechanisms of Ate1-mediated arginyl transfer reaction and substrate recognition, and define a new protein superfamily called Dupli-GNAT to reflect its origin by the duplication of the GNAT acetyltransferase domain.

Amino Acid Sequence↗

Similarity searches in genome-wide numerical data sets.

We present psi-square, a program for searching the space of gene vectors. The program starts with a gene vector, i.e., the set of measurements associated with a gene, and finds similar vectors, derives a probabilistic model of these vectors, then repeats search using this model as a query, and continues to update the model and search again, until convergence. When applied to three different pathway-discovery problems, psi-square was generally more sensitive and sometimes more specific than the ad hoc methods developed for solving each of these problems before.

Journal Article↗

Protein repertoire of double-stranded DNA bacteriophages.

The complexity and diversity of phage gene sets, which are produced by rapid evolution of phage genomes and rampant gene exchanges among phages, hamper the efforts to decipher the evolutionary relationships between individual phage proteins and reconstruct the complete set of evolutionary events leading to the known phages. To start unraveling the natural history of phages, we built the phage orthologous groups (POGs), a natural system of phage protein families that includes 6378 genes from 164 complete genome sequences of double-stranded DNA bacteriophages. Phage proteomes have high POG coverage: on average, 39 genes per phage genome belong to POGs, which is close to half of all genes in most phages. In an agreement with the notion of phage role in horizontal gene transfer, we see many cases of likely gene exchange between phages and their microbial hosts. At the same time, about 80% of all POGs are highly specific to phage genomes and are not commonly found in microbial genomes, indicating coherence and large degree of evolutionary independence of phage gene sets. The information on orthologous genes is essential for evolutionary classification of known bacteriophages and for reconstruction of ancestral phage genomes.

Bacteriophages↗

The choice of optimal distance measure in genome-wide datasets.

MOTIVATION: Many types of genomic data are naturally represented as binary vectors. Numerous tasks in computational biology can be cast as analysis of relationships between these vectors, and the first step is, frequently, to compute their pairwise distance matrix. Many distance measures have been proposed in the literature, but there is no theory justifying the choice of distance measure. RESULTS: We examine the approaches to measuring distances between binary vectors and study the characteristic properties of various distance measures and their performance in several tasks of genome analysis. Most distance measures between binary vectors turn out to belong to a single parametric family, namely generalized average-based distance with different exponents. We show that descriptive statistics of distance distribution, such as skewness and kurtosis, can guide the appropriate choice of the exponent. On the contrary, the more familiar distance properties, such as metric and additivity, appear to have much less effect on the performance of distances. AVAILABILITY: R code GADIST and Supplementary material are available at http://research.stowers-institute.org/bioinfo/

Algorithms↗

A mammalian chromatin remodeling complex with similarities to the yeast INO80 complex.

The mammalian Tip49a and Tip49b proteins belong to an evolutionarily conserved family of AAA+ ATPases. In Saccharomyces cerevisiae, orthologs of Tip49a and Tip49b, called Rvb1 and Rvb2, respectively, are subunits of two distinct ATP-dependent chromatin remodeling complexes, SWR1 and INO80. We recently demonstrated that the mammalian Tip49a and Tip49b proteins are integral subunits of a chromatin remodeling complex bearing striking similarities to the S. cerevisiae SWR1 complex (Cai, Y., Jin, J., Florens, L., Swanson, S. K., Kusch, T., Li, B., Workman, J. L., Washburn, M. P., Conaway, R. C., and Conaway, J. W. (2005) J. Biol. Chem. 280, 13665-13670). In this report, we identify a new mammalian Tip49a- and Tip49b-containing ATP-dependent chromatin remodeling complex, which includes orthologs of 8 of the 15 subunits of the S. cerevisiae INO80 chromatin remodeling complex as well as at least five additional subunits unique to the human INO80 (hINO80) complex. Finally, we demonstrate that, similar to the yeast INO80 complex, the hINO80 complex exhibits DNA- and nucleosome-activated ATPase activity and catalyzes ATP-dependent nucleosome sliding.

ATPases Associated with Diverse Cellular Activitie↗

Genome sequence and gene expression of Bacillus anthracis bacteriophage Fah.

Fah, a lytic bacteriophage of Bacillus anthracis, is used widely in the former Soviet Union to identify anthrax bacteria. Here, we present the analysis of a 37,974 bp sequence of the Fah genome and examine gene expression of the phage in a model host, Bacillus cereus. Half of the Fah genome contains genes coding for structural proteins and host lysis functions in an arrangement typical of Syphoviridae. The other half of the genome contains genes coding for enzymes of viral genome replication and for numerous predicted transcription factors that are likely to regulate viral gene expression. Primer extension, in vitro transcription assays, and gene array analysis identified temporal classes of Fah genes and allowed location of viral promoters. Fah does not execute host transcription shut-off and relies on host RNA polymerase (RNAP) sigma(A) holoenzyme for transcription of its early and late genes. In addition, Fah encodes a sigma factor, sigma(Fah), a close relative of Bacillus sporulation factor sigma(F) that directs bacterial RNAP to at least one late viral promoter. sigma(Fah) is negatively regulated by host SpoIIAB, an anti-sigma factor that controls sporulation. Thus, sigma(Fah) may link phage gene expression to sporulation of the host.

Bacillus Phages↗

Protein content of minimal and ancestral ribosome.

Minimal genome approaches seek to define the smallest gene complement compatible with modern-type cellular life on Earth. A consensus of computational and experimental approaches indicates that a minimal genome is close to 300 protein-coding genes, if a rich medium is provided for cell growth. I relate ribosomal gene content in completely sequenced genomes to ribosomal subunit structure and approximate the protein components of the putative minimal ribosome and the ribosome of the Last Universal Common Ancestor of Life. Both sets contain between 35 and 40 proteins. There is evidence of protein-protein and protein-RNA displacement in the evolution of both ribosomal subunits.

Amino Acid Sequence↗

Identification of Elongin C and Skp1 sequences that determine Cullin selection.

The multiprotein von Hippel-Lindau (VHL) tumor suppressor and Skp1-Cul1-F-box protein (SCF) complexes belong to families of structurally related E3 ubiquitin ligases. In the VHL ubiquitin ligase, the VHL protein serves as the substrate recognition subunit, which is linked by the adaptor protein Elongin C to a heterodimeric Cul2/Rbx1 module that activates ubiquitylation of target proteins by the E2 ubiquitin-conjugating enzyme Ubc5. In SCF ubiquitin ligases, F-box proteins serve as substrate recognition subunits, which are linked by the Elongin C-like adaptor protein Skp1 to a Cul1/Rbx1 module that activates ubiquitylation of target proteins, in most cases by the E2 Cdc34. In this report, we investigate the functions of the Elongin C and Skp1 proteins in reconstitution of VHL and SCF ubiquitin ligases. We identify Elongin C and Skp1 structural elements responsible for selective interaction with their cognate Cullin/Rbx1 modules. In addition, using altered specificity Elongin C and F-box protein mutants, we investigate models for the mechanism underlying E2 selection by VHL and SCF ubiquitin ligases. Our findings provide evidence that E2 selection by VHL and SCF ubiquitin ligases is determined not solely by the Cullin/Rbx1 module, the target protein, or the integrity of the substrate recognition subunit but by yet to be elucidated features of these macromolecular complexes.

Amino Acid Sequence↗

Chalcone isomerase family and fold: no longer unique to plants.

Chalcone isomerase, an enzyme in the isoflavonoid pathway in plants, catalyzes the cyclization of chalcone into (2S)-naringenin. Chalcone isomerase sequence family and three-dimensional fold appeared to be unique to plants and has been proposed as a plant-specific gene marker. Using sensitive methods of sequence comparison and fold recognition, we have identified genes homologous to chalcone isomerase in all completely sequenced fungi, in slime molds, and in many gammaproteobacteria. The residues directly involved in the enzyme's catalytic function are among the best conserved across species, indicating that the newly discovered homologs are enzymatically active. At the same time, fungal and bacterial species that have chalcone isomerase-like genes tend to lack the orthologs of the upstream enzyme chalcone synthase, suggesting a novel variation of the pathway in these species.

Amino Acid Sequence↗

Displacements of prohead protease genes in the late operons of double-stranded-DNA bacteriophages.

Most of the known prohead maturation proteases in double-stranded-DNA bacteriophages are shown, by computational methods, to fall into two evolutionarily independent clans of serine proteases, herpesvirus assemblin-like and ClpP-like. Phylogenetic analysis suggests that these two types of phage prohead protease genes displaced each other multiple times while preserving their exact location within the late operons of the phage genomes.

Amino Acid Sequence↗

Genome of Xanthomonas oryzae bacteriophage Xp10: an odd T-odd phage.

Xp10 is a lytic bacteriophage of the phytopathogenic bacterium Xanthomonas oryzae. Though morphologically Xp10 belongs to the Syphoviridae family, it encodes its own single-subunit RNA polymerase characteristic of T7-like phages of the Podoviridae family. Here, we report the determination and analysis of the 44,373 bp sequence of the Xp10 genome. The genome is a linear, double-stranded DNA molecule with 3' cohesive overhangs and no terminal repeats or redundancies. Half of the Xp10 genome contains genes coding for structural proteins and host lysis functions in an arrangement typical for temperate dairy phages that are related to the Escherichia coli lambda phage. The other half of the Xp10 genome contains genes coding for factors of host gene expression shut-off, enzymes of viral genome replication and expression. The two groups of genes are transcribed divergently and separated by a regulatory region, which contains divergent promoters recognized by the host RNA polymerase. Xp10 has apparently arisen through a recombination between genomes of widely different phages. Further evidence of extensive gene flux in the evolution of Xp10 includes a high fraction (10%) of genes derived from an HNH-family endonuclease, and a DNA-dependent DNA polymerase that is closer to a homolog from Leishmania than to DNA polymerases from other phages or bacteria.

Amino Acid Sequence↗