PubMed HealthSearch

Biomedical subjects

S Tavaré

Publications and source records attributed to S Tavaré.

At least 19 recordsLinked to original sources

The mutation rate of the human mtDNA deletion mtDNA4977.

The human mitochondrial mutation mtDNA4977 is a 4,977-bp deletion that originates between two 13-bp direct repeats. We grew 220 colonies of cells, each from a single human cell. For each colony, we counted the number of cells and amplified the DNA by PCR to test for the presence of a deletion. To estimate the mutation fate, we used a model that describes the relationship between the mutation rate and the probability that a colony of a given size will contain no mutants, taking into account such factors as possible mitochondrial turnover and mistyping due to PCR error. We estimate that the mutation rate for mtDNA4977 in cultured human cells is 5.95 x 10(-8) per mitochondrial genome replication. This method can be applied to specific chromosomal, as well as mitochondrial, mutations.

Adult

The distribution of rare alleles.

Population geneticists have long been interested in the behavior of rare variants. The definition of a rare variant has been the subject of some debate, centered mainly on whether alleles with small relative frequency should be considered rare, or whether alleles with small numbers should be. We study the behavior of the counts of rare alleles in samples taken from a population genetics model that allows for selection and infinitely-many-alleles mutation structure. We show that in large samples the counts of rare alleles--those represented once, twice, ...--are approximately distributed as a Poisson process, with a parameter that depends on the total mutation rate, but not on the selection parameters. This result is applied to the problem of estimating the fraction of neutral mutations.

Alleles

Unrooted genealogical tree probabilities in the infinitely-many-sites model.

The infinitely-many-sites process is often used to model the sequence variability observed in samples of DNA sequences. Despite its popularity, the sampling theory of the process is rather poorly understood. We describe the tree structure underlying the model and show how this may be used to compute the probability of a sample of sequences. We show how to produce the unrooted genealogy from a set of sites in which the ancestral labeling is unknown and from this the corresponding rooted genealogies. We derive recursions for the probability of the configuration of sequences (equivalently, of trees) in both the rooted and unrooted cases. We give a computational method based on Monte Carlo recursion that provides approximates to sampling probabilities for samples of any size. Among several applications, this algorithm may be used to find maximum likelihood estimators of the substitution rate, both when the ancestral labeling of sites is known and when it is unknown.

Biological Evolution

Single sperm analysis of the trinucleotide repeats in the Huntington's disease gene: quantification of the mutation frequency spectrum.

The CAG triplet repeat region of the Huntington's disease gene was amplified in 923 single sperm from three affected and two normal individuals. Average-size alleles (15-18 repeats) showed only three contraction mutations among 475 sperm (0.6%). A 30 repeat normal allele showed an 11% mutation frequency. The mutation frequency of a 36 repeat intermediate allele was 53% with 8% of all gametes having expansions which brought the allele size into the HD disease range (> or = 38 repeats). Disease alleles (38-51 repeats) showed a very high mutation frequency (92-99%). As repeat number increased there was a marked elevation in the frequency of expansions, in the mean number of repeats added per expansion and the size of the largest observed expansion. Contraction frequencies also appeared to increase with allele size but decreased as repeat number exceeded 36. Our sperm typing data are of a discrete nature rather than consisting of smears of PCR product from pooled sperm. This allowed the observed mutation frequency spectra to be compared to the distribution calculated using discrete stochastic models based on current molecular ideas of the expansion process. An excellent fit was found when the model specified that a random number of repeats are added during the progression of the polymerase through the repeated region.

Alleles

Sampling theory for neutral alleles in a varying environment.

We develop a sampling theory for genes sampled from a population evolving with deterministically varying size. We use a coalescent approach to provide recursions for the probabilities of particular sample configurations, and describe a Monte Carlo method by which the solutions to such recursions can be approximated. We focus on infinite-alleles, infinite-sites and finite-sites models. This approach may be used to find maximum likelihood estimates of parameters of genetic interest, and to test hypotheses about the varying environment. The methods are illustrated with data from the mitochondrial control region sampled from a North American Indian tribe.

Alleles

Estimating substitution rates from molecular data using the coalescent.

A coalescent model is used to estimate the rate at which neutral substitutions occur in a DNA sequence, without the necessity for an independent estimate of divergence times. Given a random sample of molecular sequences from a finite population, the distribution of the time to a common ancestor can be obtained from the coalescent model. With this principle, summary statistics are developed that use the distribution of molecular diversity within the sample to estimate the relative magnitude of nucleotide substitution rates. If, in addition, the effective population size is known, absolute substitution rates can also be estimated. These techniques are illustrated by estimating the transition rates that underlie the evolution of the first 360 nucleotides of the mitochondrial control region in an Amerindian tribal population.

Biological Evolution

Modeling the evolution of the human mitochondrial genome.

Mitochondrial DNA data have been used extensively to study evolution and early human origins. These applications require estimates of the rate at which nucleotide substitutions occur in the DNA sequence. We consider the problem of estimating substitution rates in the presence of site-to-site rate variation. A coalescent model is presented that allows for different substitution rates for purines and pyrimidines, as well as more detailed models that allow fast and slow rates within each of the purine and pyrimidine classes. A method for estimating such rates is presented. Even for these simple models of site heterogeneity, there are, typically, insufficient data to obtain reliable estimates of site-specific substitution rates. However, estimates of the average rate across all sites appear to be relatively stable even in the presence of site heterogeneity. Simulations of models with site-to-site variation in mutation rate show that hypervariable sites can produce peaks in the pairwise difference curves that have previously been attributed to population dynamics.

Biological Evolution

Genomic mapping by anchoring random clones: a mathematical analysis.

A complete physical map of the DNA of an organism, consisting of overlapping clones spanning the genome, is an extremely useful tool for genomic analysis. Various methods for the construction of such physical maps are available. One approach is to assemble the physical map by "fingerprinting" a large number of random clones and inferring overlap between clones with sufficiently similar fingerprints. E.S. Lander and M.S. Waterman (1988, Genomics 2:231-239) have recently provided a mathematical analysis of such physical mapping schemes, useful for planning such a project. Another approach is to assemble the physical map by "anchoring" a large number of random clones--that is, by taking random short regions called anchors and identifying the clones containing each anchor. Here, we provide a mathematical analysis of such a physical mapping scheme.

Chromosome Mapping

Codon preference and primary sequence structure in protein-coding regions.

The stochastic complexity of a data base of 365 protein-coding regions is analysed. When the primary sequence is modeled as a spatially homogeneous Markov source, the fit to observed codon preference is very poor. The situation improves substantially when a non-homogeneous model is used. Some implications for the estimation of species phylogeny and substitution rates are discussed.

Amino Acids

Is knowing the age-order of alleles in a sample useful in testing for selective neutrality?

The most powerful, and most frequently used, test of selective neutrality, based on data consisting of observed allelic frequencies in a sample of genes at some locus, is the procedure of G. A. Watterson. This procedure uses the sample homozygosity F* as the test statistic, and in effect leads to rejection of the hypothesis of selective neutrality if the observed value of F* differs significantly from neutral theory expectations. The homozygosity statistic is invariant under relabeling of the alleles and thus cannot use any further information on the alleles which might be available. We present results which suggest that information concerning the age order of the alleles cannot be used to provide a more powerful testing procedure than that of Watterson.

Alleles

The birth process with immigration, and the genealogical structure of large populations.

This paper studies a version of the birth and immigration process in which families are followed in the order of their appearance. This age structure is related to a number of results from population genetics, in particular the genealogical structure of the infinitely-many neutral alleles model. The asymptotic behavior of this genealogy is an easy consequence of the structure of the age-ordered family size process.

Biometry

The population genealogy of the infinitely--many neutral alleles model.

A process analogous to Kingman's coalescent is introduced to describe the genealogy of populations evolving according to the infinitely- many neutral alleles model. The process records population frequencies in old and new classes, and labels the new classes in order of decreasing age. Its marginal distribution is characterized in a form which is amenable to explicit calculations and the transition densities of the associated K-allele models follow readily from this representation.

Alleles

Line-of-descent and genealogical processes, and their applications in population genetics models.

A variety of results for genealogical and line-of-descent processes that arise in connection with the theory of some classical selectively neutral population genetics models are reviewed. While some new results and derivations are included, the principle aim is to demonstrate the central importance and simplicity of genealogical Markov chains in this theory. Considerable attention is given to "diffusion time scale" approximations of such genealogical processes. A wide variety of results pertinent to (diffusion approximations of) the classical multiallele single-locus Wright-Fisher model and its relatives are simplified and unified by this approach. Other examples where such genealogical processes play an explicit role, such as the infinite sites and infinite alleles models, are discussed.

Alleles

Some stochastic models for plasmid copy number.

Some stochastic models for the copy number of plasmids in a cell line are studied. When considering the behavior of copy number in the whole cell line, the theory of multitype branching processes is appropriate. Attention is paid to the cure rate in the cell line, and the asymptotic fractions of cells containing a given number of plasmids. These quantities are used to compare the models numerically.

DNA Replication