PubMed HealthSearch

Biomedical subjects

J Felsenstein

Publications and source records attributed to J Felsenstein.

At least 19 recordsLinked to original sources

Estimating effective population size and mutation rate from sequence data using Metropolis-Hastings sampling.

We present a new way to make a maximum likelihood estimate of the parameter 4N mu (effective population size times mutation rate per site, or theta) based on a population sample of molecular sequences. We use a Metropolis-Hastings Markov chain Monte Carlo method to sample genealogies in proportion to the product of their likelihood with respect to the data and their prior probability with respect to a coalescent distribution. A specific value of theta must be chosen to generate the coalescent distribution, but the resulting trees can be used to evaluate the likelihood at other values of theta, generating a likelihood curve. This procedure concentrates sampling on those genealogies that contribute most of the likelihood, allowing estimation of meaningful likelihood curves based on relatively small samples. The method can potentially be extended to cases involving varying population size, recombination, and migration.

Base Sequence

Inching toward reality: an improved likelihood model of sequence evolution.

Our previous evolutionary model is generalized to permit approximate treatment of multiple-base insertions and deletions as well as regional heterogeneity of substitution rates. Parameter estimation and alignment procedures that incorporate these generalizations are developed. Simulations are used to assess the accuracy of the parameter estimation procedure and an example of an inferred alignment is included.

Animals

Estimating effective population size from samples of sequences: inefficiency of pairwise and segregating sites as compared to phylogenetic estimates.

It is known that under neutral mutation at a known mutation rate a sample of nucleotide sequences, within which there is assumed to be no recombination, allows estimation of the effective size of an isolated population. This paper investigates the case of very long sequences, where each pair of sequences allows a precise estimate of the divergence time of those two gene copies. The average divergence time of all pairs of copies estimates twice the effective population number and an estimate can also be derived from the number of segregating sites. One can alternatively estimate the genealogy of the copies. This paper shows how a maximum likelihood estimate of the effective population number can be derived from such a genealogical tree. The pairwise and the segregating sites estimates are shown to be much less efficient than this maximum likelihood estimate, and this is verified by computer simulation. The result implies that there is much to gain by explicitly taking the tree structure of these genealogies into account.

Likelihood Functions

Estimating effective population size from samples of sequences: a bootstrap Monte Carlo integration method.

We would like to use maximum likelihood to estimate parameters such as the effective population size N(e) or, if we do not know mutation rates, the product 4N(e) mu of mutation rate per site and effective population size. To compute the likelihood for a sample of unrecombined nucleotide sequences taken from a random-mating population it is necessary to sum over all genealogies that could have led to the sequences, computing for each one the probability that it would have yielded the sequences, and weighting each one by its prior probability. The genealogies vary in tree topology and in branch lengths. Although the likelihood and the prior are straightforward to compute, the summation over all genealogies seems at first sight hopelessly difficult. This paper reports that it is possible to carry out a Monte Carlo integration to evaluate the likelihoods approximately. The method uses bootstrap sampling of sites to create data sets for each of which a maximum likelihood tree is estimated. The resulting trees are assumed to be sampled from a distribution whose height is proportional to the likelihood surface for the full data. That it will be so is dependent on a theorem which is not proven, but seems likely to be true if the sequences are not short. One can use the resulting estimated likelihood curve to make a maximum likelihood estimate of the parameter of interest, N(e) or of 4N(e) mu. The method requires at least 100 times the computational effort required for estimation of a phylogeny by maximum likelihood, but is practical on today's work stations. The method does not at present have any way of dealing with recombination.

Base Sequence

Counting phylogenetic invariants in some simple cases.

An informal degrees of freedom argument is used to count the number of phylogenetic invariants in cases where we have three or four species and can assume a Jukes-Cantor model of base substitution with or without a molecular clock. A number of simple cases are treated and in each the number of invariants can be found. Two new classes of invariants are found: non-phylogenetic cubic invariants testing independence of evolutionary events in different lineages, and linear phylogenetic invariants which occur when there is a molecular clock. Most of the linear invariants found by Cavender (1989, Molec. Biol. Evol. 6, 301-316) turn out in the Jukes-Cantor case to be simple tests of symmetry of the substitution model, and not phylogenetic invariants.

Animals

An evolutionary model for maximum likelihood alignment of DNA sequences.

Most algorithms for the alignment of biological sequences are not derived from an evolutionary model. Consequently, these alignment algorithms lack a strong statistical basis. A maximum likelihood method for the alignment of two DNA sequences is presented. This method is based upon a statistical model of DNA sequence evolution for which we have obtained explicit transition probabilities. The evolutionary model can also be used as the basis of procedures that estimate the evolutionary parameters relevant to a pair of unaligned DNA sequences. A parameter-estimation approach which takes into account all possible alignments between two sequences is introduced; the danger of estimating evolutionary parameters from a single alignment is discussed.

Algorithms

A maximum likelihood approach to the detection of selection from a phylogeny.

A large amount of information is contained within the phylogenetic relationships between species. In addition to their branching patterns it is also possible to examine other aspects of the biology of the species. The influence that deleterious selection might have is determined here. The likelihood of different phylogenies in the presence of selection is explored to determine the properties of such a likelihood surface. The calculation of likelihoods for a phylogeny in the presence and absence of selection, permits the application of a likelihood ratio test to search for selection. It is shown that even a single selected site can have a strong effect on the likelihood. The method is illustrated with an example from Drosophila melanogaster and suggests that deleterious selection may be acting on transposable elements.

Animals

Estimation of hominoid phylogeny from a DNA hybridization data set.

Analysis of the expanded data set of Sibley and Ahlquist (1987) on primate phylogeny using a maximum likelihood mixed model analysis of variance method shows that there is significant evidence for resolving the Homo-Pan-Gorilla trifurcation in favor of a Homo-Pan clade. The resulting tree is close to that estimated by Sibley and Ahlquist (1984). The mixed model can be used to test a number of hypotheses about the existence of components of variance and the linearity of the relationship between branch length and expected distance. No evidence is found that there is a variance component for extract, or for the individual from which the extract was taken. A variance component for experiment does seem to exist, presumably arising as a result of error of measurement of the common standard from which all values in the same experiment were subtracted. There is significant evidence that the relationship between total branch length between species and their expected distances is nonlinear, or else that the measurement error on larger distances is greater than on smaller ones. Allowing for the nonlinearity might cause one to infer the time of distant common ancestors as less remote than the measured hybridization values would imply if used directly.

Analysis of Variance

An efficient method for matching nucleic acid sequences.

A method of computing the fraction of matches between two nucleic acid sequences at all possible alignments is described. It makes use of the Fast Fourier Transform. It should be particularly efficient for very long sequences, achieving its result in a number of operations proportional to n ln n, where n is the length of the longer of the two sequences. Though the objective achieved is of limited interest, this method will complement algorithms for efficiently finding the longest matching parts of two sequences, and is faster than existing algorithms for finding matches allowing deletions and insertions. A variety of economies can be achieved by this Fast Fourier Transform technique in matching multiple sequences, looking for complementarity rather than identity, and matching the same sequences both in forward and reversed orientations.

Base Sequence

A continuous migration model with stable demography.

A probability model of a population undergoing migration, mutation, and mating in a geographic continuum R is constructed, and an integro-differential equation is derived for the probability of genetic identity. The equation is solved in one case, and asymptotic analysis done in others. Individuals at x, y epsilon R in the model mate with probability V(x, y) dt in any time interval (t, t + dt). In two dimensions, if V(x, y) = V(x - y) where V(x) approximately V(x/beta)/beta2 approaches a delta function, the equilibrium probability of identity vanishes as beta Leads to 0. The asymptotic rate at which this occurs is discussed for mutation rates u = u0 Greater than 0 and for beta approximately cua, alpha Greater than 0, and u Leads to 0.

Animals

Evolutionary trees from DNA sequences: a maximum likelihood approach.

The application of maximum likelihood techniques to the estimation of evolutionary trees from nucleic acid sequence data is discussed. A computationally feasible method for finding such maximum likelihood estimates is developed, and a computer program is available. This method has advantages over the traditional parsimony algorithms, which can give misleading results if rates of evolution differ in different lineages. It also allows the testing of hypotheses about the constancy of evolutionary rates by likelihood ratio tests, and gives rough indication of the error of ;the estimate of the tree.

Base Sequence

Isolation by distance: reply to Lalouel and Morton.

Lalouel's assertion that I misinterpreted Malécot's work on isolation by distance may or may not be correct. If so, my assertions of error in Malécot's derivation are wrong, although they do apply to others who have used models involving a spatial continuum. Lalouel's other claims of error in my derivations of the consequences of a spatially continuous model of population reproduction and migration are incorrect, with the exception of one isolated misprint.

Genetics, Population

A model of kin selection for an altruistic trait considered as a quantitative character.

Conditions for natural selection to favor increase of a quantitative character are derived for a model in which individuals associate in groups of size n. It is assumed that the logarithm of the fitness of an individual is the sum of two parts, one proportional to the individual's own phenotype, and the other to the mean phenotype in its group. The resulting conditions for the trait to increase under natural selection are analogous to the results found previously in single-locus kin selection models.

Altruism