PubMed HealthSearch

PubMed · 8790461

Reconstructing trees when sequence sites evolve at variable rates.

Abstract

For a sequence of colors independently evolving on a tree under a simple Markov model, we consider conditions under which the tree can be uniquely recovered from the "sequence spectrum"-the expected frequencies of the various leaf colorations. This is relevant for phylogenetic analysis (where colors represent nucleotides or amino acids; leaves represent extant taxa) as the sequence spectrum is estimated directly from a collection of aligned sequences. Allowing the rate of the evolutionary process to vary across sites is an important extension over most previous studies-we show that, given suitable restrictions on the rate distribution, the true tree (up to the placement of its root) is uniquely identified by its sequence spectrum. However, if the rate distribution is unknown and arbitrary, then, for simple models, it is possible for every tree to produce the same sequence spectrum. Hence there is a logical barrier to accurate, consistent phylogenetic inference for these models when assumptions about the rate distribution are not made. This result exploits a novel theorem on the action of polynomials with non-negative coefficients on sequences.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

M A Steel, L A Székely, M D Hendy. 1994. Reconstructing trees when sequence sites evolve at variable rates.. https://doi.org/10.1089/cmb.1994.1.153

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A free energy analysis by unfolding applied to 125-mers on a cubic lattice.

BACKGROUND: A common approach to the protein folding problem involves computer simulation of folding using lattice models of amino acid sequences. Key factors for good performance in such models are the correct choice of the temperature and the average interaction energy between residues. In order to push the lattice approach to its limit it is important to have a method to adjust these parameters for optimal folding that is not limited by our ability to successfully simulate folding in a reasonable time. RESULTS: In this study, we adopt a simple cubic-lattice model and present a method for calculating the free energy of a chain as a function of the number of native contacts. This does not require that we are able to fold the sequence by simulation and it provides a method of estimating the folding transition temperature. For a given set of parameters, the free energy analysis also allows an estimate of foldability. By applying the method to sequences with 27 and 125 residues, we show that optimal folding occurs near the folding transition temperature and at either zero or small negative average interaction energy. We find ourselves able to fold only 125-mers that have significant short-range native contacts. CONCLUSIONS: A free energy analysis during unfolding is a useful tool for the study of foldability and should be applicable to a variety of folding models. In this way we are able to fold some 125-mer designed sequences and our results confirm the finding that short-range contacts contribute to foldability.

Markov Chains

A covariotide model explains apparent phylogenetic structure of oxygenic photosynthetic lineages.

The aims of the work were (1) to develop statistical tests to identify whether substitution takes place under a covariotide model in sequences used for phylogenetic inference and (2) to determine the influence of covariotide substitution on phylogenetic trees inferred for photosynthetic and other organisms. (Covariotide and covarion models are ones in which sites that are variable in some parts of the underlying tree are invariable in others and vice versa.) Two tests were developed. The first was a contingency test, and the second was an inequality test comparing the expected number of variable sites in two groups with the observed number. Application of these tests to 16S rDNA and tufA sequences from a range of nonphotosynthetic prokaryotes and oxygenic photosynthetic prokaryotes and eukaryotes suggests the occurrence of a covariotide mechanism. The degree of support for partitioning of taxa in reconstructed trees involving these organisms was determined in the presence or absence of sites showing particular substitution patterns. This analysis showed that the support for splits between (1) photosynthetic eukaryotes and prokaryotes and (2) photosynthetic and nonphotosynthetic organisms could be accounted for by patterns arising from covariotide substitution. We show that the additional problem of compositional bias in sequence data needs to be considered in the context of patterns of covariotide/covarion substitution. We argue that while covariotide or covarion substitution may give rise to phylogenetically informative patterns in sequence data, this may not always be so.

Markov Chains

Maximum likelihood estimation of aggregated Markov processes.

We present a maximum likelihood method for the modelling of aggregated Markov processes. The method utilizes the joint probability density of the observed dwell time sequence as likelihood. A forward-backward recursive procedure is developed for efficient computation of the likelihood function and its derivatives with respect to the model parameters. Based on the calculated forward and backward vectors, analytical formulae for the derivatives of the likelihood function are derived. The method exploits the variable metric optimizer for search of the likelihood space. It converges rapidly and is numerically stable. Numerical examples are given to show the effectiveness of the method.

Markov Chains