PubMed Health⌕ Search

Biomedical subjects

Yuji Inagaki

Publications and source records attributed to Yuji Inagaki.

14 recordsLinked to original sources

Phylogenetic estimation under codon models can be biased by codon usage heterogeneity.

In theory, codon models that account for the dependence of nucleotide substitutions between codon positions as well as differences between synonymous and non-synonymous changes best describe the sequence evolution in protein coding genes. However, in practice we know little about the degree to which violations of the assumptions of codon model-based estimates occur, and how significant these artifacts may be. In nucleotide-based phylogenies from first and second codon positions in a concatenated plastid gene data set, two distantly related taxa--dinoflagellate and haptophyte plastids--were robustly grouped together. This artifactual grouping is attributed to the parallel heterogeneity in leucine (Leu) and serine (Ser) codon usages in the data set. Here, by using this data set, we demonstrated that codon-based phylogenetic estimations are seriously biased, robustly uniting the dinoflagellate and haptophyte plastids into a monophyletic clade, when the model assumption of homogeneity of codon composition was violated. Our results suggest that similar phylogenetic artifacts may occur via codon usage heterogeneity in any amino acids in codon model-based estimations. We advise that homogeneity in codon usage across taxa in a data set be confirmed before codon model-based phylogenetic estimation is attempted.

Codon↗

Recombination between elongation factor 1alpha genes from distantly related archaeal lineages.

Homologous recombination (HR) and lateral gene transfer are major processes in genome evolution. The combination of the two processes, HR between genes in different species, has been documented but is thought to be restricted to very similar sequences in relatively closely related organisms. Here we report two cases of interspecific HR in the gene encoding the core translational protein translation elongation factor 1alpha (EF-1alpha) between distantly related archaeal groups. Maximum-likelihood sliding window analyses indicate that a fragment of the EF-1alpha gene from the archaeal lineage represented by Methanopyrus kandleri was recombined into the orthologous gene in a common ancestor of the Thermococcales. A second recombination event appears to have occurred between the EF-1alpha gene of the genus Methanothermobacter and its ortholog in a common ancestor of the Methanosarcinales, a distantly related euryarchaeal lineage. These findings suggest that HR occurs across a much larger evolutionary distance than generally accepted and affects highly conserved essential "informational" genes. Although difficult to detect by standard whole-gene phylogenetic analyses, interspecific HR in highly conserved genes may occur at an appreciable frequency, potentially confounding deep phylogenetic inference and hypothesis testing.

Archaea↗

Comprehensive multigene phylogenies of excavate protists reveal the evolutionary positions of "primitive" eukaryotes.

Many of the protists thought to represent the deepest branches on the eukaryotic tree are assigned to a loose assemblage called the "excavates." This includes the mitochondrion-lacking diplomonads and parabasalids (e.g., Giardia and Trichomonas) and the jakobids (e.g., Reclinomonas). We report the first multigene phylogenetic analyses to include a comprehensive sampling of excavate groups (six nuclear-encoded protein-coding genes, nine of the 10 recognized excavate groups). Excavates coalesce into three clades with relatively strong maximum likelihood bootstrap support. Only the phylogenetic position of Malawimonas is uncertain. Diplomonads, parabasalids, and the free-living amitochondriate protist Carpediemonas are closely related to each other. Two other amitochondriate excavates, oxymonads and Trimastix, form the second monophyletic group. The third group is comprised of Euglenozoa (e.g., trypanosomes), Heterolobosea, and jakobids. Unexpectedly, jakobids appear to be specifically related to Heterolobosea. This tree topology calls into question the concept of Discicristata as a supergroup of eukaryotes united by discoidal mitochondrial cristae and makes it implausible that jakobids represent an independent early-diverging eukaryotic lineage. The close jakobids-Heterolobosea-Euglenozoa connection demands complex evolutionary scenarios to explain the transition between the presumed ancestral bacterial-type mitochondrial RNA polymerase found in jakobids and the phage-type protein in other eukaryotic lineages, including Euglenozoa and Heterolobosea.

Animals↗

A close relationship between Cercozoa and Foraminifera supported by phylogenetic analyses based on combined amino acid sequences of three cytoskeletal proteins (actin, alpha-tubulin, and beta-tubulin).

Recently, there has been increasing molecular evidence of phylogenetic affinity between Cercozoa and Foraminifera in the eukaryotic lineage. We performed phylogenetic analyses based on the combined (concatenated) amino acid sequence data of actin, alpha-tubulin, and beta-tubulin from a wide variety of eukaryotes, including the foraminifers Planoglabratella opercularis and Reticulomyxa filosa, as well as cercomonad and chlorarachniophyte members of Cercozoa. A monophyletic lineage composed of two foraminiferan species branched with the centroheliozoan species Raphidiophrys contractilis was reconstructed in both Bayesian and maximum-likelihood (ML) analyses under 'linked' models, enforcing a single set of the parameters (the parameter for among-site rate variation and branch lengths) on the entire combined alignment. Considering the extremely divergent nature of Foraminifera and Raphidiophyrs tubulins, the union of these lineages recovered is most probably a long-branch attraction artifact due to ignoring gene-specific evolutionary processes. On the other hand, the foraminiferan lineage was within the radiation of Cercozoa in Bayesian analyses under 'unlinked' model conditions, accommodating differences in evolutionary processes across the three genes in the combined alignment. The Foraminifera+Cercozoa affinity recovered in the latter multi-gene analyses is most likely genuine, and thus our data presented here provide further support for the close relationship between these two protist lineages.

Actins↗

Copper-mediated oxidative DNA damage induced by eugenol: possible involvement of O-demethylation.

Eugenol used as a flavor has potential carcinogenicity. DNA adduct formation via 2,3-epoxidation pathway has been thought to be a major mechanism of DNA damage by carcinogenic allylbenzene analogs including eugenol. We examined whether eugenol can induce oxidative DNA damage in the presence of cytochrome P450 using [32P]-5'-end-labeled DNA fragments obtained from human genes relevant to cancer. Eugenol induced Cu(II)-mediated DNA damage in the presence of cytochrome P450 (CYP)1A1, 1A2, 2C9, 2D6, or 2E1. CYP2D6 mediated eugenol-dependent DNA damage most efficiently. Piperidine and formamidopyrimidine-DNA glycosylase treatment induced cleavage sites mainly at T and G residues of the 5'-TG-3' sequence, respectively. Interestingly, CYP2D6-treated eugenol strongly damaged C and G of the 5'-ACG-3' sequence complementary to codon 273 of the p53 gene. These results suggest that CYP2D6-treated eugenol can cause double base lesions. DNA damage was inhibited by both catalase and bathocuproine, suggesting that H2O2 and Cu(I) are involved. These results suggest that Cu(I)-hydroperoxo complex is primary reactive species causing DNA damage. Formation of 8-oxo-7,8-dihydro-2'-deoxyguanosine was significantly increased by CYP2D6-treated eugenol in the presence of Cu(II). Time-of-flight-mass spectrometry demonstrated that CYP2D6 catalyzed O-demethylation of eugenol to produce hydroxychavicol, capable of causing DNA damage. Therefore, it is concluded that eugenol may express carcinogenicity through oxidative DNA damage by its metabolite.

8-Hydroxy-2'-Deoxyguanosine↗

A class of eukaryotic GTPase with a punctate distribution suggesting multiple functional replacements of translation elongation factor 1alpha.

Translation elongation factor 1alpha (EF-1alpha, or EF-Tu in bacteria) is a highly conserved core component of the translation machinery that is shared by all cellular life. It is part of a large superfamily of GTPases that are involved in translation initiation, elongation, and termination, as well as several other cellular functions. Eukaryotic EF-1alpha (eEF-1alpha) is well studied and widely sampled and has been used extensively for phylogenetic analyses. It is generally thought that such highly conserved and functionally integrated proteins are unlikely to be involved in events such as lateral gene transfer or ancient duplication and gene sorting, which would undermine phylogenetic reconstruction. Here we describe a GTPase called EF-like (EFL), which is very similar to, but also distinct from, canonical eEF-1alpha. EFL is found in a wide variety of eukaryotes (dinoflagellates, haptophytes, cercozoa, green algae, choanoflagellates, and fungi), but its distribution is punctate: organisms that possess EFL are not closely related to one another, and EFL appears to be absent from the closest relatives of organisms that do possess it. Moreover, in most genomes where EFL is present, canonical eEF-1alpha appears to be absent. Analysis of functional divergence suggests that, whereas EFL is divergent in general, putative functional binding sites involved in translation are not significantly divergent as a whole. Altogether, it appears that EFL has replaced eEF-1alpha several times independently. This finding could be an indication of an ancient paralogy or, more likely, eukaryote-to-eukaryote lateral gene transfer.

Eukaryotic Cells↗

On inconsistency of the neighbor-joining, least squares, and minimum evolution estimation when substitution processes are incorrectly modeled.

Using analytical methods, we show that under a variety of model misspecifications, Neighbor-Joining, minimum evolution, and least squares estimation procedures are statistically inconsistent. Failure to correctly account for differing rates-across-sites processes, failure to correctly model rate matrix parameters, and failure to adjust for parallel rates-across-sites changes (a rates-across-subtrees process) are all shown to lead to a "long branch attraction" form of inconsistency. In addition, failure to account for rates-across-sites processes is also shown to result in underestimation of evolutionary distances for a wide variety of substitution models, generalizing an earlier analytical result for the Jukes-Cantor model reported in Golding and a similar bias result for the GTR or REV model in Kelly and Rice (1996). Although standard rates-across-sites models can be employed in many of these cases to restore consistency, current models cannot account for other kinds of misspecification. We examine an idealized but biologically relevant case, where parallel changes in rates at sites across subtrees is shown to give rise to inconsistency. This changing rates-across-subtrees type model misspecification cannot be adjusted for with conventional methods or without carefully considering the rate variation in the larger tree. The results are presented for four-taxon trees, but the expectation is that they have implications for larger trees as well. To illustrate this, a simulated 42-taxon example is given in which the microsporidia, an enigmatic group of eukaryotes, are incorrectly placed at the archaebacteria-eukaryotes split because of incorrectly specified pairwise distances. The analytical nature of the results lend insight into the reasons that long branch attraction tends to be a common form of inconsistency and reasons that other forms of inconsistency like "long branches repel" can arise in some settings. In many of the cases of inconsistency presented, a particular incorrect topology is estimated with probability converging to one, the implication being that measures of uncertainty like bootstrap support will be unable to detect that there is a problem with the estimation. The focus is on distance methods, but previous simulation results suggest that the zones of inconsistency for distance methods contain the zones of inconsistency for maximum likelihood methods as well.

Archaea↗

Covarion shifts cause a long-branch attraction artifact that unites microsporidia and archaebacteria in EF-1alpha phylogenies.

Microsporidia branch at the base of eukaryotic phylogenies inferred from translation elongation factor 1alpha (EF-1alpha) sequences. Because these parasitic eukaryotes are fungi (or close relatives of fungi), it is widely accepted that fast-evolving microsporidian sequences are artifactually "attracted" to the long branch leading to the archaebacterial (outgroup) sequences ("long-branch attraction," or "LBA"). However, no previous studies have explicitly determined the reason(s) why the artifactual allegiance of microsporidia and archaebacteria ("M + A") is recovered by all phylogenetic methods, including maximum likelihood, a method that is supposed to be resistant to classical LBA. Here we show that the M + A affinity can be attributed to those alignment sites associated with large differences in evolutionary site rates between the eukaryotic and archaebacterial subtrees. Therefore, failure to model the significant evolutionary rate distribution differences (covarion shifts) between the ingroup and outgroup sequences is apparently responsible for the artifactual basal position of microsporidia in phylogenetic analyses of EF-1alpha sequences. Currently, no evolutionary model that accounts for discrete changes in the site rate distribution on particular branches is available for either protein or nucleotide level phylogenetic analysis, so the same artifacts may affect many other "deep" phylogenies. Furthermore, given the relative similarity of the site rate patterns of microsporidian and archaebacterial EF-1alpha proteins ("parallel site rate variation"), we suggest that the microsporidian orthologs may have lost some eukaryotic EF-1alpha-specific nontranslational functions, exemplifying the extreme degree of reduction in this parasitic lineage.

Animals↗

Phylogenetic artifacts can be caused by leucine, serine, and arginine codon usage heterogeneity: dinoflagellate plastid origins as a case study.

Phylogenetic analyses of first and second codon positions (DNA1 + 2 analysis) and amino acid sequences (protein analysis) are often thought to provide similar estimates of deep-level phylogeny. However, here we report a novel artifact influencing DNA level phylogenetic inference of protein-coding genes introduced by codon usage heterogeneity that causes significant incongruities between DNA1 + 2 and protein analyses. DNA1 + 2 analyses of plastid-encoded psbA genes (encoding of photosystem II D1 proteins) strongly suggest a relationship between haptophyte plastids and typical (peridinin-containing) dinoflagellate plastids. The psbA genes from haptophytes and a subset of the peridinin-type plastids display similar codon usage patterns for Leu, Ser, and Arg, which are each encoded by two separated codon sets that differ at first or first plus second codon positions. Our detailed analyses clearly indicate that these unusual preferences shared by haptophyte and some peridinin-type plastid genes are largely responsible for their strong affinity in DNA analyses. In particular, almost all of the support from DNA level analyses for the monophyly of haptophyte and peridinin-type plastids is lost when the codons corresponding to constant Leu, Ser, and Arg amino acids are excluded, suggesting that this signal comes from rapidly evolving synonymous substitutions, rather than from substitutions that result in amino acid changes. Indeed, protein maximum-likelihood analyses of concatenated PsaA and PsbA amino acid sequences indicate that, although 19' hexanoyloxyfucoxanthin-type (19' HNOF-type) plastids in dinoflagellates group with haptophyte plastids, peridinin-type plastids group weakly with those of stramenopiles. Consequently our results cast doubt on the single origin of peridinin-type and 19' HNOF-type plastids in dinoflagellates previously suggested on the basis of psaA and psbA concatenated gene phylogenetic analyses. We suggest that codon usage heterogeneity could be a more general problem for DNA level analyses of protein-coding genes, even when third codon positions are excluded.

Animals↗

Assessing functional divergence in EF-1alpha and its paralogs in eukaryotes and archaebacteria.

A number of methods have recently been published that use phylogenetic information extracted from large multiple sequence alignments to detect sites that have changed properties in related protein families. In this study we use such methods to assess functional divergence between eukaryotic EF-1alpha (eEF-1alpha), archaebacterial EF-1alpha (aEF-1alpha) and two eukaryote-specific EF-1alpha paralogs-eukaryotic release factor 3 (eRF3) and Hsp70 subfamily B suppressor 1 (HBS1). Overall, the evolutionary modes of aEF-1alpha, HBS1 and eRF3 appear to significantly differ from that of eEF-1alpha. However, functionally divergent (FD) sites detected between aEF-1alpha and eEF-1alpha only weakly overlap with sites implicated as putative EF-1beta or aminoacyl-tRNA (aa-tRNA) binding residues in EF-1alpha, as expected based on the shared ancestral primary translational functions of these two orthologs. In contrast, FD sites detected between eEF-1alpha and its paralogs significantly overlap with the putative EF-1beta and/or aa-tRNA binding sites in EF-1alpha. In eRF3 and HBS1, these sites appear to be released from functional constraints, indicating that they bind neither eEF-1beta nor aa-tRNA. These results are consistent with experimental observations that eRF3 does not bind to aa-tRNA, but do not support the 'EF-1alpha-like' function recently proposed for HBS1. We re-assess the available genetic data for HBS1 in light of our analyses, and propose that this protein may function in stop codon-independent peptide release.

Amino Acid Sequence↗

Lateral transfer of an EF-1alpha gene: origin and evolution of the large subunit of ATP sulfurylase in eubacteria.

It is generally accepted that new genes arise via duplication and functional divergence of existing genes, in accordance with Ohno's model, now called "Mutation During Redundancy," or MDR. In this model, one of the two gene copies is free to acquire novel (although likely related) activities through mutation, since only one copy is required for its original function. However, duplication within a genome is not the only process that might give rise to this situation: acquisition of a functionally redundant gene by lateral gene transfer (LGT) could also initiate the MDR process. Here we describe a probable instance, involving LGT of an archaeal or eukaryotic elongation factor 1alpha (EF-1alpha) gene. The large subunit of ATP sulfurylase (CysN or the N-terminal portion of NodQ), found mainly in proteobacteria, is clearly related to translation elongation factors. However, our analyses show that cysN arose from an EF-1alpha gene initially acquired by LGT, not from a within-genome duplication of the resident EF-Tu gene. To our knowledge, this is the first unequivocal case of LGT followed by functional modification to be described; this mechanism could be a potentially important force in establishing genes with novel functions in genomes.

Amino Acid Sequence↗

Convergence and constraint in eukaryotic release factor 1 (eRF1) domain 1: the evolution of stop codon specificity.

Class 1 release factor in eukaryotes (eRF1) recognizes stop codons and promotes peptide release from the ribosome. The 'molecular mimicry' hypothesis suggests that domain 1 of eRF1 is analogous to the tRNA anticodon stem-loop. Recent studies strongly support this hypothesis and several models for specific interactions between stop codons and residues in domain 1 have been proposed. In this study we have sequenced and identified novel eRF1 sequences across a wide diversity of eukaryotes and re-evaluated the codon-binding site by bioinformatic analyses of a large eRF1 dataset. Analyses of the eRF1 structure combined with estimates of evolutionary rates at amino acid sites allow us to define the residues that are under structural (i.e. those involved in intramolecular interactions) versus non-structural selective constraints. Furthermore, we have re-assessed convergent substitutions in the ciliate variant code eRF1s using maximum likelihood-based phylogenetic approaches. Our results favor the model proposed by Bertram et al. that stop codons bind to three 'cavities' on the protein surface, although we suggest that the stop codon may bind in the opposite orientation to the original model. We assess the feasibility of this alternative binding orientation with a triplet stop codon and the eRF1 domain 1 structures using molecular modeling techniques.

Amino Acid Sequence↗

Testing for differences in rates-across-sites distributions in phylogenetic subtrees.

It has long been recognized that the rates of molecular evolution vary amongst sites in proteins. The usual model for rate heterogeneity assumes independent rate variation according to a rate distribution. In such models the rate at a site, although random, is assumed fixed throughout the evolutionary tree. Recent work by several groups has suggested that rates at sites often vary across subtrees of the larger tree as well as across sites. This phenomenon is not captured by most phylogenetic models but instead is more similar to the covarion model of Fitch and coworkers. In this article we present methods that can be useful in detecting whether different rates occur in two different subtrees of the larger tree and where these differences occur. Parametric bootstrapping and orthogonal regression methodologies are used to test for rate differences and to make statements about the general differences in the rates at sites. Confidence intervals based on the conditional distributions of rates at sites are then used to detect where the rate differences occur. Such methods will be helpful in studying the phylogenetic, structural, and functional bases of changes in evolutionary rates at sites, a phenomenon that has important consequences for deep phylogenetic inference.

Confidence Intervals↗