PubMed Health⌕ Search

Biomedical subjects

Edward Susko

Publications and source records attributed to Edward Susko.

14 recordsLinked to original sources

Testing for covarion-like evolution in protein sequences.

The covarion hypothesis of molecular evolution proposes that selective pressures on an amino acid or nucleotide site change through time, thus causing changes of evolutionary rate along the edges of a phylogenetic tree. Several kinds of Markov models for the covarion process have been proposed. One model, proposed by Huelsenbeck (2002), has 2 substitution rate classes: the substitution process at a site can switch between a single variable rate, drawn from a discrete gamma distribution, and a zero invariable rate. A second model, suggested by Galtier (2001), assumes rate switches among an arbitrary number of rate classes but switching to and from the invariable rate class is not allowed. The latter model allows for some sites that do not participate in the rate-switching process. Here we propose a general covarion model that combines features of both models, allowing evolutionary rates not only to switch between variable and invariable classes but also to switch among different rates when they are in a variable state. We have implemented all 3 covarion models in a maximum likelihood framework for amino acid sequences and tested them on 23 protein data sets. We found significant likelihood increases for all data sets for the 3 models, compared with a model that does not allow site-specific rate switches along the tree. Furthermore, we found that the general model fit the data better than the simpler covarion models in the majority of the cases, highlighting the complexity in modeling the covarion process. The general covarion model can be used for comparing tree topologies, molecular dating studies, and the investigation of protein adaptation.

Algorithms↗

A selective barrier to horizontal gene transfer in the T4-type bacteriophages that has preserved a core genome with the viral replication and structural genes.

Genomic analysis of bacteriophages frequently reveals a mosaic structure made up from modules that come from disparate sources. This fact has led to the general acceptance of the notion that rampant and promiscuous lateral gene transfer (LGT) plays a critical role in phage evolution. However, recent sequencing of a series of the T4-type phages has revealed that these large and complex genomes all share 2 substantial syntenous blocks of genes encoding the replication and virion structural genes. To analyze the pattern of inheritance of this core T4 genome, we compared the complete genome sequences of 16 T4-type phages. We identified a set of 24 genes present in all these T4-type genomes. Somewhat surprisingly, only one of these genes, that encodes for ribonucleotide reductase (NrdA), displayed evidence of LGT with the bacterial host. We test the congruence of the inheritance of the other 23 markers using heat map analyses and comparison of a reference topology with the 23 individual gene phylogenies. The vast majority of these core genes share a common evolutionary history. In contrast, analyses of all the noncore genes present in the same 16 genomes, located in the hyperplastic regions of the genome, show considerable evidence of frequent LGT. The similar evolution of the core replication and virion structural genes in the T4-type phage genomes suggests that, unlike the situation in many other phage groups, such portions of T4-type genome have been inherited as a block, without significant LGT, from a distant common ancestor. The preservation of the synteny of the core T4 genome could result from several factors acting in synergy, such as the constraints imposed by the sophisticated regulation of the transcription. Moreover, numerous and complex protein-protein interactions during virion morphogenesis could also impose a supplementary barrier against LGT. Finally, there may be some real evolutionary advantage to maintaining large regions of conserved sequence. Such segments could be a sort of genetic glue that maintains the genetic cohesion of the T4-type phages via recombination within the most conserved sequences. This could mediate the swapping of nonconserved sequences that they flank.

Bacteriophage T4↗

Recombination between elongation factor 1alpha genes from distantly related archaeal lineages.

Homologous recombination (HR) and lateral gene transfer are major processes in genome evolution. The combination of the two processes, HR between genes in different species, has been documented but is thought to be restricted to very similar sequences in relatively closely related organisms. Here we report two cases of interspecific HR in the gene encoding the core translational protein translation elongation factor 1alpha (EF-1alpha) between distantly related archaeal groups. Maximum-likelihood sliding window analyses indicate that a fragment of the EF-1alpha gene from the archaeal lineage represented by Methanopyrus kandleri was recombined into the orthologous gene in a common ancestor of the Thermococcales. A second recombination event appears to have occurred between the EF-1alpha gene of the genus Methanothermobacter and its ortholog in a common ancestor of the Methanosarcinales, a distantly related euryarchaeal lineage. These findings suggest that HR occurs across a much larger evolutionary distance than generally accepted and affects highly conserved essential "informational" genes. Although difficult to detect by standard whole-gene phylogenetic analyses, interspecific HR in highly conserved genes may occur at an appreciable frequency, potentially confounding deep phylogenetic inference and hypothesis testing.

Archaea↗

On the correlation between genomic G+C content and optimal growth temperature in prokaryotes: data quality and confounding factors.

The correlation between genomic G+C content and optimal growth temperature in prokaryotes has gained renewed interest after Musto et al. [H. Musto, H. Naya, A. Zavala, H. Romero, F. Alvarex-Valin, G. Bernardi, Correlations between genomic GC levels and optimal growth temperatures in prokaryotes, FEBS Lett. 573 (2004) 73-77], reported that positive correlations exist in 15 families studied. We have reanalyzed their data and found that when genome size and data quality were adjusted for, there was no significant evidence of relationship between optimal temperature and GC content for two of the families that had previously shown strongly significant correlations. Using updated temperature optima for Halobacteriaceae species we found the correlation is insignificant in this family. For the family Enterobacteriaceae when genome size and optimal temperature are included in a multiple linear regression, only genome size is significant as a predictor of GC content. We showed that more profound statistical methods than simple two factor correlation analysis should be used for analyzing complex intrinsic and extrinsic factors that affect genomic GC content. We further found that a positive correlation between temperature and genomic GC is only evident in free-living species of low optimal growth temperatures.

Base Composition↗

The comparison of the confidence regions in phylogeny.

In this paper, several different procedures for constructing confidence regions for the true evolutionary tree are evaluated both in terms of coverage and size without considering model misspecification. The regions are constructed on the basis of tests of hypothesis using six existing tests: Shimodaira Hasegawa (SH), SOWH, star form of SOWH (SSOWH), approximately unbiased (AU), likelihood weight (LW), generalized least squares, plus two new tests proposed in this paper: single distribution nonparametric bootstrap (SDNB) and single distribution parametric bootstrap (SDPB). The procedures are evaluated on simulated trees both with small and large number of taxa. Overall, the SH, SSOWH, AU, and LW tests led to regions with higher coverage than the nominal level at the price of including large numbers of trees. Under the specified model, the SOWH test gives accurate coverage and relatively small regions. The SDNB and SDPB tests led to the small regions with occasional undercoverage. These two procedures have a substantial computational advantage over the SOWH test. Finally, the cutoff levels for the SDNB test are shown to be more variable than those for the SDPB test.

Classification↗

Biases in phylogenetic estimation can be caused by random sequence segments.

We consider the effects of fully or partially random sequences on the estimation of four-taxon phylogenies. Fully or partially random sequences occur when whole subsets of sequences or some sites for subsets of sequences are independent of sequence data for the other taxa. Random sequences can be a consequence of misalignment or because sites evolve at very fast rates in some portions of a tree, a situation that occurs especially in analyses involving deep divergence times. One might reasonably speculate that random sites will only add noise to the estimation of a phylogeny. We show that in the case that a random sequence is added to a three-taxa alignment, it is more likely to be a neighbor of the sequence corresponding to the longest branch in the three-taxon tree. Surprisingly, when only about half of the sites show randomness, a long-branch-repels form of small sample bias occurs, and when a minority of sites show randomness this becomes a long-branch-attraction bias again. The most serious bias, one that does not vanish with increasing sequence length, occurs when more than one sequence is partially random. If there is a large amount of overlap in the random sites for two sequences, those two sequences will be attracted to each other; otherwise, they will repel each other. Random sequences or sites can, therefore, cause complicated biases in phylogenetic inference. We suggest performing analyses with and without potentially saturated sequences and/or misaligned sites, to check that these biases are not affecting the inferred branching pattern.

Animals↗

Likelihood, parsimony, and heterogeneous evolution.

Evolutionary rates vary among sites and across the phylogenetic tree (heterotachy). A recent analysis suggested that parsimony can be better than standard likelihood at recovering the true tree given heterotachy. The authors recommended that results from parsimony, which they consider to be nonparametric, be reported alongside likelihood results. They also proposed a mixture model, which was inconsistent but better than either parsimony or standard likelihood under heterotachy. We show that their main conclusion is limited to a special case for the type of model they study. Their mixture model was inconsistent because it was incorrectly implemented. A useful nonparametric model should perform well over a wide range of possible evolutionary models, but parsimony does not have this property. Likelihood-based methods are therefore the best way to deal with heterotachy.

Animals↗

On inconsistency of the neighbor-joining, least squares, and minimum evolution estimation when substitution processes are incorrectly modeled.

Using analytical methods, we show that under a variety of model misspecifications, Neighbor-Joining, minimum evolution, and least squares estimation procedures are statistically inconsistent. Failure to correctly account for differing rates-across-sites processes, failure to correctly model rate matrix parameters, and failure to adjust for parallel rates-across-sites changes (a rates-across-subtrees process) are all shown to lead to a "long branch attraction" form of inconsistency. In addition, failure to account for rates-across-sites processes is also shown to result in underestimation of evolutionary distances for a wide variety of substitution models, generalizing an earlier analytical result for the Jukes-Cantor model reported in Golding and a similar bias result for the GTR or REV model in Kelly and Rice (1996). Although standard rates-across-sites models can be employed in many of these cases to restore consistency, current models cannot account for other kinds of misspecification. We examine an idealized but biologically relevant case, where parallel changes in rates at sites across subtrees is shown to give rise to inconsistency. This changing rates-across-subtrees type model misspecification cannot be adjusted for with conventional methods or without carefully considering the rate variation in the larger tree. The results are presented for four-taxon trees, but the expectation is that they have implications for larger trees as well. To illustrate this, a simulated 42-taxon example is given in which the microsporidia, an enigmatic group of eukaryotes, are incorrectly placed at the archaebacteria-eukaryotes split because of incorrectly specified pairwise distances. The analytical nature of the results lend insight into the reasons that long branch attraction tends to be a common form of inconsistency and reasons that other forms of inconsistency like "long branches repel" can arise in some settings. In many of the cases of inconsistency presented, a particular incorrect topology is estimated with probability converging to one, the implication being that measures of uncertainty like bootstrap support will be unable to detect that there is a problem with the estimation. The focus is on distance methods, but previous simulation results suggest that the zones of inconsistency for distance methods contain the zones of inconsistency for maximum likelihood methods as well.

Archaea↗

Estimating and comparing the rates of gene discovery and expressed sequence tag (EST) frequencies in EST surveys.

MOTIVATION: Expressed sequence tag (EST) surveys are an efficient way to characterize large numbers of genes from an organism. The rate of gene discovery in an EST survey depends on the degree of redundancy of the cDNA libraries from which sequences are obtained. However, few statistical methods have been developed to assess and compare redundancies of various libraries from preliminary EST surveys. RESULTS: We consider statistics for the comparison of EST libraries based upon the frequencies with which genes occur in subsamples of reads. These measures are useful in determining which one of several libraries is more likely to yield new genes in future reads and what proportion of additional reads one might want to take from the libraries in order to be likely to obtain new genes. One approach is to compare single sample measures that have been successfully used in species estimation problems, such as coverage of a library, defined as the proportion of the library that is represented in the given sample of reads. Another single library measure is an estimate of the expected number of additional genes that will be found in a new sample of reads. We also propose statistics that jointly use data from all the libraries. Analogous formulas for coverage and the expected numbers of new genes are presented. These measures consider coverage in a single library based upon reads from all libraries and similarly, the expected numbers of new genes that will be discovered by taking reads from all libraries with fixed proportions. Together, the statistics presented provide useful comparative measures for the libraries that can be used to guide sampling from each of the libraries to maximize the rate of gene discovery. Finally, we present tests for whether genes are equally represented or expressed in a set of libraries. Binomial and chi2 tests are presented for gene-by-gene comparisons of expression. Overall tests of the equality of proportional representation are presented and multiple comparisons issues are addressed. These methods can be used to evaluate changes in gene expression reflected in the composition of EST libraries prepared from different tissue types or cells exposed to different environmental conditions. AVAILABILITY: Software will be made available at http://www.mathstat.dal.ca/~tsusko

Algorithms↗

Covarion shifts cause a long-branch attraction artifact that unites microsporidia and archaebacteria in EF-1alpha phylogenies.

Microsporidia branch at the base of eukaryotic phylogenies inferred from translation elongation factor 1alpha (EF-1alpha) sequences. Because these parasitic eukaryotes are fungi (or close relatives of fungi), it is widely accepted that fast-evolving microsporidian sequences are artifactually "attracted" to the long branch leading to the archaebacterial (outgroup) sequences ("long-branch attraction," or "LBA"). However, no previous studies have explicitly determined the reason(s) why the artifactual allegiance of microsporidia and archaebacteria ("M + A") is recovered by all phylogenetic methods, including maximum likelihood, a method that is supposed to be resistant to classical LBA. Here we show that the M + A affinity can be attributed to those alignment sites associated with large differences in evolutionary site rates between the eukaryotic and archaebacterial subtrees. Therefore, failure to model the significant evolutionary rate distribution differences (covarion shifts) between the ingroup and outgroup sequences is apparently responsible for the artifactual basal position of microsporidia in phylogenetic analyses of EF-1alpha sequences. Currently, no evolutionary model that accounts for discrete changes in the site rate distribution on particular branches is available for either protein or nucleotide level phylogenetic analysis, so the same artifacts may affect many other "deep" phylogenies. Furthermore, given the relative similarity of the site rate patterns of microsporidian and archaebacterial EF-1alpha proteins ("parallel site rate variation"), we suggest that the microsporidian orthologs may have lost some eukaryotic EF-1alpha-specific nontranslational functions, exemplifying the extreme degree of reduction in this parasitic lineage.

Animals↗

Assessing functional divergence in EF-1alpha and its paralogs in eukaryotes and archaebacteria.

A number of methods have recently been published that use phylogenetic information extracted from large multiple sequence alignments to detect sites that have changed properties in related protein families. In this study we use such methods to assess functional divergence between eukaryotic EF-1alpha (eEF-1alpha), archaebacterial EF-1alpha (aEF-1alpha) and two eukaryote-specific EF-1alpha paralogs-eukaryotic release factor 3 (eRF3) and Hsp70 subfamily B suppressor 1 (HBS1). Overall, the evolutionary modes of aEF-1alpha, HBS1 and eRF3 appear to significantly differ from that of eEF-1alpha. However, functionally divergent (FD) sites detected between aEF-1alpha and eEF-1alpha only weakly overlap with sites implicated as putative EF-1beta or aminoacyl-tRNA (aa-tRNA) binding residues in EF-1alpha, as expected based on the shared ancestral primary translational functions of these two orthologs. In contrast, FD sites detected between eEF-1alpha and its paralogs significantly overlap with the putative EF-1beta and/or aa-tRNA binding sites in EF-1alpha. In eRF3 and HBS1, these sites appear to be released from functional constraints, indicating that they bind neither eEF-1beta nor aa-tRNA. These results are consistent with experimental observations that eRF3 does not bind to aa-tRNA, but do not support the 'EF-1alpha-like' function recently proposed for HBS1. We re-assess the available genetic data for HBS1 in light of our analyses, and propose that this protein may function in stop codon-independent peptide release.

Amino Acid Sequence↗

Confidence regions and hypothesis tests for topologies using generalized least squares.

A confidence region for topologies is a data-dependent set of topologies that, with high probability, can be expected to contain the true topology. Because of the connection between confidence regions and hypothesis tests, implicitly or explicitly, the construction of confidence regions for topologies is a component of many phylogenetic studies. Existing methods for constructing confidence regions, however, often give conflicting results. The Shimodaira-Hasegawa test seems too conservative, including too many topologies, whereas the other commonly used method, the Swofford-Olsen-Waddell-Hillis test, tends to give confidence regions with too few topologies. Confidence regions are constructed here based on a generalized least squares test statistic. The methodology described is computationally inexpensive and broadly applicable to maximum likelihood distances. Assuming the model used to construct the distances is correct, the coverage probabilities are correct with large numbers of sites.

Animals↗

Estimation of rates-across-sites distributions in phylogenetic substitution models.

Previous work has shown that it is often essential to account for the variation in rates at different sites in phylogenetic models in order to avoid phylogenetic artifacts such as long branch attraction. In most current models, the gamma distribution is used for the rates-across-sites distributions and is implemented as an equal-probability discrete gamma. In this article, we introduce discrete distribution estimates with large numbers of equally spaced rate categories allowing us to investigate the appropriateness of the gamma model. With large numbers of rate categories, these discrete estimates are flexible enough to approximate the shape of almost any distribution. Likelihood ratio statistical tests and a nonparametric bootstrap confidence-bound estimation procedure based on the discrete estimates are presented that can be used to test the fit of a parametric family. We applied the methodology to several different protein data sets, and found that although the gamma model often provides a good parametric model for this type of data, rate estimates from an equal-probability discrete gamma model with a small number of categories will tend to underestimate the largest rates. In cases when the gamma model assumption is in doubt, rate estimates coming from the discrete rate distribution estimate with a large number of rate categories provide a robust alternative to gamma estimates. An alternative implementation of the gamma distribution is proposed that, for equal numbers of rate categories, is computationally more efficient during optimization than the standard gamma implementation and can provide more accurate estimates of site rates.

Evolution, Molecular↗

Testing for differences in rates-across-sites distributions in phylogenetic subtrees.

It has long been recognized that the rates of molecular evolution vary amongst sites in proteins. The usual model for rate heterogeneity assumes independent rate variation according to a rate distribution. In such models the rate at a site, although random, is assumed fixed throughout the evolutionary tree. Recent work by several groups has suggested that rates at sites often vary across subtrees of the larger tree as well as across sites. This phenomenon is not captured by most phylogenetic models but instead is more similar to the covarion model of Fitch and coworkers. In this article we present methods that can be useful in detecting whether different rates occur in two different subtrees of the larger tree and where these differences occur. Parametric bootstrapping and orthogonal regression methodologies are used to test for rate differences and to make statements about the general differences in the rates at sites. Confidence intervals based on the conditional distributions of rates at sites are then used to detect where the rate differences occur. Such methods will be helpful in studying the phylogenetic, structural, and functional bases of changes in evolutionary rates at sites, a phenomenon that has important consequences for deep phylogenetic inference.

Confidence Intervals↗