PubMed Health⌕ Search

Biomedical subjects

Michael W Gray

Publications and source records attributed to Michael W Gray.

At least 19 recordsLinked to original sources

TBestDB: a taxonomically broad database of expressed sequence tags (ESTs).

The TBestDB database contains approximately 370,000 clustered expressed sequence tag (EST) sequences from 49 organisms, covering a taxonomically broad range of poorly studied, mainly unicellular eukaryotes, and includes experimental information, consensus sequences, gene annotations and metabolic pathway predictions. Most of these ESTs have been generated by the Protist EST Program, a collaboration among six Canadian research groups. EST sequences are read from trace files up to a minimum quality cut-off, vector and linker sequence is masked, and the ESTs are clustered using phrap. The resulting consensus sequences are automatically annotated by using the AutoFACT program. The datasets are automatically checked for clustering errors due to chimerism and potential cross-contamination between organisms, and suspect data are flagged in or removed from the database. Access to data deposited in TBestDB by individual users can be restricted to those users for a limited period. With this first report on TBestDB, we open the database to the research community for free processing, annotation, interspecies comparisons and GenBank submission of EST data generated in individual laboratories. For instructions on submission to TBestDB, contact tbestdb@bch.umontreal.ca. The database can be queried at http://tbestdb.bcm.umontreal.ca/.

Animals↗

The frequency of eubacterium-to-eukaryote lateral gene transfers shows significant cross-taxa variation within amoebozoa.

Single-celled bacterivorous eukaryotes offer excellent test cases for evaluation of the frequency of prey-to-predator lateral gene transfer (LGT). Here we use analysis of expressed sequence tag (EST) data sets to quantify the extent of LGT from eubacteria to two amoebae, Acanthamoeba castellanii and Hartmannella vermiformis. Stringent screening for LGT proceeded in several steps intended to enrich for authentic events while at the same time minimizing the incidence of false positives due to factors such as limitations in database coverage and ancient paralogy. The results were compared with data obtained when the same methodology was applied to EST libraries from a number of other eukaryotic taxa. Significant differences in the extent of apparent eubacterium-to-eukaryote LGT were found between taxa. Our results indicate that there may be substantial inter-taxon variation in the number of LGT events that become fixed even between amoebozoan species that have similar feeding modalities.

Acanthamoeba castellanii↗

An early evolutionary origin for the minor spliceosome.

The minor spliceosome is a ribonucleoprotein complex that catalyses the removal of an atypical class of spliceosomal introns (U12-type) from eukaryotic messenger RNAs. It was first identified and characterized in animals, where it was found to contain several unique RNA constituents that share structural similarity with and seem to be functionally analogous to the small nuclear RNAs (snRNAs) contained in the major spliceosome. Subsequently, minor spliceosomal components and U12-type introns have been found in plants but not in fungi. Unlike that of the major spliceosome, which arose early in the eukaryotic lineage, the evolutionary history of the minor spliceosome is unclear because there is evidence of it in so few organisms. Here we report the identification of homologues of minor-spliceosome-specific proteins and snRNAs, and U12-type introns, in distantly related eukaryotic microbes (protists) and in a fungus (Rhizopus oryzae). Cumulatively, our results indicate that the minor spliceosome had an early origin: several of its characteristic constituents are present in representative organisms from all eukaryotic supergroups for which there is any substantial genome sequence information. In addition, our results reveal marked evolutionary conservation of functionally important sequence elements contained within U12-type introns and snRNAs.

Acanthamoeba castellanii↗

Analysis of Euglena gracilis plastid-targeted proteins reveals different classes of transit sequences.

The plastid of Euglena gracilis was acquired secondarily through an endosymbiotic event with a eukaryotic green alga, and as a result, it is surrounded by a third membrane. This membrane complexity raises the question of how the plastid proteins are targeted to and imported into the organelle. To further explore plastid protein targeting in Euglena, we screened a total of 9,461 expressed sequence tag (EST) clusters (derived from 19,013 individual ESTs) for full-length proteins that are plastid localized to characterize their targeting sequences and to infer potential modes of translocation. Of the 117 proteins identified as being potentially plastid localized whose N-terminal targeting sequences could be inferred, 83 were unique and could be classified into two major groups. Class I proteins have tripartite targeting sequences, comprising (in order) an N-terminal signal sequence, a plastid transit peptide domain, and a predicted stop-transfer sequence. Within this class of proteins are the lumen-targeted proteins (class IB), which have an additional hydrophobic domain similar to a signal sequence and required for further targeting across the thylakoid membrane. Class II proteins lack the putative stop-transfer sequence and possess only a signal sequence at the N terminus, followed by what, in amino acid composition, resembles a plastid transit peptide. Unexpectedly, a few unrelated plastid-targeted proteins exhibit highly similar transit sequences, implying either a recent swapping of these domains or a conserved function. This work represents the most comprehensive description to date of transit peptides in Euglena and hints at the complex routes of plastid targeting that must exist in this organism.

Algal Proteins↗

Twinkle, the mitochondrial replicative DNA helicase, is widespread in the eukaryotic radiation and may also be the mitochondrial DNA primase in most eukaryotes.

Recently, the human protein responsible for replicative mtDNA helicase activity was identified and designated Twinkle. Twinkle has been implicated in autosomal dominant progressive external ophthalmoplegia (adPEO), a mitochondrial disorder characterized by mtDNA deletions. The Twinkle protein appears to have evolved from an ancestor shared with the bifunctional primase-helicase found in the T-odd bacteriophages. However, the question has been raised as to whether human Twinkle possesses primase activity, due to amino acid sequence divergence and absence of a zinc-finger motif thought to play an integral role in DNA binding. To date, a primase protein participating in mtDNA replication has not been identified in any eukaryote. Here we investigate the wider phylogenetic distribution of Twinkle by surveying and analyzing data from ongoing EST and genome sequencing projects. We identify Twinkle homologues in representatives from five of six major eukaryotic assemblages ("supergroups") and present the sequence of the complete Twinkle gene from two members of Amoebozoa, a supergroup of amoeboid protists at the base of the opisthokont (fungal/metazoan) radiation. Notably, we identify conserved primase motifs including the zinc finger in all Twinkle sequences outside of Metazoa. Accordingly, we propose that Twinkle likely serves as the primase as well as the helicase for mtDNA replication in most eukaryotes whose genome encodes it, with the exception of Metazoa.

Amino Acid Sequence↗

Homologs of mitochondrial transcription factor B, sparsely distributed within the eukaryotic radiation, are likely derived from the dimethyladenosine methyltransferase of the mitochondrial endosymbiont.

Mitochondrial transcription factor B (mtTFB), an essential component in regulating the expression of mitochondrial DNA-encoded genes in both yeast and humans, is a dimethyladenosine methyltransferase (DMT) that has acquired a secondary role in mitochondrial transcription. So far, mtTFB has only been well studied in Opisthokonta (metazoan animals and fungi). Here we investigate the phylogenetic distribution of mtTFB homologs throughout the domain Eucarya, documenting the first examples of this protein outside of the opisthokonts. Surprisingly, we identified putative mtTFB homologs only in amoebozoan protists and trypanosomatids. Phylogenetic analysis together with conservation of intron positions in amoebozoan and human genes supports the grouping of the putative mtTFB homologs as a distinct clade. Phylogenetic analysis further demonstrates that the mtTFB is most likely derived from the DMT of the mitochondrial endosymbiont.

Acanthamoeba castellanii↗

A large collection of compact box C/D snoRNAs and their isoforms in Euglena gracilis: structural, functional and evolutionary insights.

In the domains Eucarya and Archaea, box C/D RNAs guide methylation at the 2'-position of selected ribose residues in ribosomal RNA (rRNA). Those eukaryotic box C/D RNAs that have been identified to date are larger and more variable in size than their archaeal counterparts. Here, we report the first extensive identification and characterization of box C/D small nucleolar (sno) RNAs from the protist Euglena gracilis. Among several unexpected findings, this organism contains a large assortment of methylation-guide RNAs that are smaller and more uniformly sized than those of other eukaryotes, and that consist of surprisingly few double-guide RNAs targeting sites of rRNA modification. Our comprehensive examination of the modification status of E.gracilis rRNA indicates that many of these box C/D snoRNAs target clustered methylation sites requiring extensive, overlapping guide RNA/rRNA pairings. An examination of the structure of the RNAs, in particular the location of the functional guide elements, suggests that the distances between adjacent box elements are an important factor in determining which of the potential guide elements is used to target a site of O(2')-methylation.

Animals↗

Bacteriophage origins of mitochondrial replication and transcription proteins.

Mounting evidence suggests that key components of the mitochondrial transcription and replication apparatus are derived from the T-odd lineage of bacteriophage rather than from an alpha-Proteobacterium, as the endosymbiont hypothesis would predict. We propose that several mitochondrial replication genes were acquired together from an ancestor of T-odd phage early in the evolution of the eukaryotic cell, at the time of the mitochondrial endosymbiosis. We further propose that at a later stage the single-subunit RNA polymerase, originally acquired for mitochondrial DNA replication, was co-opted to serve in mitochondrial transcription.

Bacteriophage T7↗

The tree of eukaryotes.

Recent advances in resolving the tree of eukaryotes are converging on a model composed of a few large hypothetical 'supergroups', each comprising a diversity of primarily microbial eukaryotes (protists, or protozoa and algae). The process of resolving the tree involves the synthesis of many kinds of data, including single-gene trees, multigene analyses, and other kinds of molecular and structural characters. Here, we review the recent progress in assembling the tree of eukaryotes, describing the major evidence for each supergroup, and where gaps in our knowledge remain. We also consider other factors emerging from phylogenetic analyses and comparative genomics, in particular lateral gene transfer, and whether such factors confound our understanding of the eukaryotic tree.

Journal Article↗

An ancient spliceosomal intron in the ribosomal protein L7a gene (Rpl7a) of Giardia lamblia.

BACKGROUND: Only one spliceosomal-type intron has previously been identified in the unicellular eukaryotic parasite, Giardia lamblia (a diplomonad). This intron is only 35 nucleotides in length and is unusual in possessing a non-canonical 5' intron boundary sequence, CT, instead of GT. RESULTS: We have identified a second spliceosomal-type intron in G. lamblia, in the ribosomal protein L7a gene (Rpl7a), that possesses a canonical GT 5' intron boundary sequence. A comparison of the two known Giardia intron sequences revealed extensive nucleotide identity at both the 5' and 3' intron boundaries, similar to the conserved sequence motifs recently identified at the boundaries of spliceosomal-type introns in Trichomonas vaginalis (a parabasalid). Based on these observations, we searched the partial G. lamblia genome sequence for these conserved features and identified a third spliceosomal intron, in an unassigned open reading frame. Our comprehensive analysis of the Rpl7a intron in other eukaryotic taxa demonstrates that it is evolutionarily conserved and is an ancient eukaryotic intron. CONCLUSION: An analysis of the phylogenetic distribution and properties of the Rpl7a intron suggests its utility as a phylogenetic marker to evaluate particular eukaryotic groupings. Additionally, analysis of the G. lamblia introns has provided further insight into some of the conserved and unique features possessed by the recently identified spliceosomal introns in related organisms such as T. vaginalis and Carpediemonas membranifera.

Animals↗

AutoFACT: an automatic functional annotation and classification tool.

BACKGROUND: Assignment of function to new molecular sequence data is an essential step in genomics projects. The usual process involves similarity searches of a given sequence against one or more databases, an arduous process for large datasets. RESULTS: We present AutoFACT, a fully automated and customizable annotation tool that assigns biologically informative functions to a sequence. Key features of this tool are that it (1) analyzes nucleotide and protein sequence data; (2) determines the most informative functional description by combining multiple BLAST reports from several user-selected databases; (3) assigns putative metabolic pathways, functional classes, enzyme classes, GeneOntology terms and locus names; and (4) generates output in HTML, text and GFF formats for the user's convenience. We have compared AutoFACT to four well-established annotation pipelines. The error rate of functional annotation is estimated to be only between 1-2%. Comparison of AutoFACT to the traditional top-BLAST-hit annotation method shows that our procedure increases the number of functionally informative annotations by approximately 50%. CONCLUSION: AutoFACT will serve as a useful annotation tool for smaller sequencing groups lacking dedicated bioinformatics staff. It is implemented in PERL and runs on LINUX/UNIX platforms. AutoFACT is available at http://megasun.bch.umontreal.ca/Software/AutoFACT.htm.

Acanthamoeba castellanii↗

Unusual features of fibrillarin cDNA and gene structure in Euglena gracilis: evolutionary conservation of core proteins and structural predictions for methylation-guide box C/D snoRNPs throughout the domain Eucarya.

Box C/D ribonucleoprotein (RNP) particles mediate O2'-methylation of rRNA and other cellular RNA species. In higher eukaryotic taxa, these RNPs are more complex than their archaeal counterparts, containing four core protein components (Snu13p, Nop56p, Nop58p and fibrillarin) compared with three in Archaea. This increase in complexity raises questions about the evolutionary emergence of the eukaryote-specific proteins and structural conservation in these RNPs throughout the eukaryotic domain. In protists, the primarily unicellular organisms comprising the bulk of eukaryotic diversity, the protein composition of box C/D RNPs has not yet been extensively explored. This study describes the complete gene, cDNA and protein sequences of the fibrillarin homolog from the protozoon Euglena gracilis, the first such information to be obtained for a nucleolus-localized protein in this organism. The E.gracilis fibrillarin gene contains a mixture of intron types exhibiting markedly different sizes. In contrast to most other E.gracilis mRNAs characterized to date, the fibrillarin mRNA lacks a spliced leader (SL) sequence. The predicted fibrillarin protein sequence itself is unusual in that it contains a glycine-lysine (GK)-rich domain at its N-terminus rather than the glycine-arginine-rich (GAR) domain found in most other eukaryotic fibrillarins. In an evolutionarily diverse collection of protists that includes E.gracilis, we have also identified putative homologs of the other core protein components of box C/D RNPs, thereby providing evidence that the protein composition seen in the higher eukaryotic complexes was established very early in eukaryotic cell evolution.

Animals↗

Gene discovery in the Acanthamoeba castellanii genome.

Acanthamoeba castellanii is a free-living amoeba found in soil, freshwater, and marine environments and an important predator of bacteria. Acanthamoeba castellanii is also an opportunistic pathogen of clinical interest, responsible for several distinct diseases in humans. In order to provide a genomic platform for the study of this ubiquitous and important protist, we generated a sequence survey of approximately 0.5 x coverage of the genome. The data predict that A. castellanii exhibits a greater biosynthetic capacity than the free-living Dictyostelium discoideum and the parasite Entamoeba histolytica, providing an explanation for the ability of A. castellanii to inhabit a diversity of environments. Alginate lyase may provide access to bacteria within biofilms by breaking down the biofilm matrix, and polyhydroxybutyrate depolymerase may facilitate utilization of the bacterial storage compound polyhydroxybutyrate as a food source. Enzymes for the synthesis and breakdown of cellulose were identified, and they likely participate in encystation and excystation as in D. discoideum. Trehalose-6-phosphate synthase is present, suggesting that trehalose plays a role in stress adaptation. Detection and response to a number of stress conditions is likely accomplished with a large set of signal transduction histidine kinases and a set of putative receptor serine/threonine kinases similar to those found in E. histolytica. Serine, cysteine and metalloproteases were identified, some of which are likely involved in pathogenicity.

Acanthamoeba castellanii↗

In vitro characterization of a tRNA editing activity in the mitochondria of Spizellomyces punctatus, a Chytridiomycete fungus.

In the chytridiomycete fungus, Spizellomyces punctatus, all eight of the mitochondrially encoded tRNAs are predicted to have one or more base pair mismatches at the first three positions of their aminoacyl acceptor stems. These tRNAs are edited post-transcriptionally by replacement of the 5'-nucleotide in each mismatched pair with a nucleotide that can form a standard Watson-Crick base pair with its counterpart in the 3'-half of the stem. The type of mitochondrial tRNA editing found in S. punctatus also occurs in Acanthamoeba castellanii, a distantly related amoeboid protist. Using an S. punctatus mitochondrial extract, we have developed an in vitro assay of tRNA editing in which nucleotides are incorporated into various tRNA substrates. Experiments employing synthetic transcripts revealed that the S. punctatus tRNA editing activity incorporates nucleotides on the 5'-side of substrate tRNAs, uses the 3'-sequence as a template for incorporation, and adds nucleotides in a 3'-to-5' direction. This activity can add nucleotides to a triphosphorylated 5'-end in the absence of ATP but requires ATP to add nucleotides to a monophosphorylated 5'-end; moreover, it functions independently of the state of tRNA 3' processing. These data parallel results obtained in a previous in vitro study of A. castellanii tRNA editing, suggesting that remarkably similar activities function in the mitochondria of these two organisms. The evolutionary origins of these activities are discussed.

Adenosine Triphosphate↗

Evolution of the mitochondrial genome: protist connections to animals, fungi and plants.

The past decade has seen the determination of complete mitochondrial genome sequences from a taxonomically diverse set of organisms. These data have allowed an unprecedented understanding of the evolution of the mitochondrial genome in terms of gene content and order, as well as genome size and structure. In addition, phylogenetic reconstructions based on mitochondrial DNA (mtDNA)-encoded protein sequences have firmly established the identities of protistan relatives of the animal, fungal and plant lineages. Analysis of the mtDNAs of these protists has provided insight into the structure of the mitochondrial genome at the origin of these three, mainly multicellular, eukaryotic groups. Further research into mtDNAs of taxa ancestral and intermediate to currently characterized organisms will help to refine pathways and modes of mtDNA evolution, as well as provide valuable phylogenetic characters to assist in unraveling the deep branching order of all eukaryotes.

Animals↗

Mitochondria of protists.

Over the past several decades, our knowledge of the origin and evolution of mitochondria has been greatly advanced by determination of complete mitochondrial genome sequences. Among the most informative mitochondrial genomes have been those of protists (primarily unicellular eukaryotes), some of which harbor the most gene-rich and most eubacteria-like mitochondrial DNAs (mtDNAs) known. Comparison of mtDNA sequence data has provided insights into the radically diverse trends in mitochondrial genome evolution exhibited by different phylogenetically coherent groupings of eukaryotes, and has allowed us to pinpoint specific protist relatives of the multicellular eukaryotic lineages (animals, plants, and fungi). This comparative genomics approach has also revealed unique and fascinating aspects of mitochondrial gene expression, highlighting the mitochondrion as an evolutionary playground par excellence.

Animals↗