PubMed Health⌕ Search

Biomedical subjects

C Wiuf

Publications and source records attributed to C Wiuf.

7 recordsLinked to original sources

A coalescence approach to gene conversion.

In this paper we develop a coalescent model with intralocus gene conversion. Such models are of increasing importance in the analysis of intralocus variability and linkage disequilibrium. We derive the distribution of the waiting time until a gene conversion event occurs in a sample in terms of the distribution of the length of the transferred segment, zeta. We do not assume any specific form of the distribution of zeta. Further, given that a gene conversion event occurs we find the distribution of (sigma, tau), the end points of the transferred segment and derive results on correlations between local trees in positions chi(1) and chi(2). Among other results we show that the correlation between the branch lengths of two local trees in the coalescent with gene conversion (and no recombination) decreases toward a nonzero constant when the distance between chi(1) and chi(2) increases. Finally, we show that a model including both recombination and gene conversion might account for the lack of intralocus associations found in, e.g., Drosophila melanogaster.

Gene Conversion↗

The coalescent with gene conversion.

In this article we develop a coalescent model with intralocus gene conversion. The distribution of the tract length is geometric in concordance with results published in the literature. We derive a simulation scheme and deduce a number of analytical results for this coalescent with gene conversion. We compare patterns of variability in samples simulated according to the coalescent with recombination with similar patterns simulated according to the coalescent with gene conversion alone. Further, an expression for the expected number of topology shifts in a sample of present-day sequences caused by gene conversion events is derived.

Gene Conversion↗

Recombination as a point process along sequences.

Histories of sequences in the coalescent model with recombination can be simulated using an algorithm that takes as input a sample of extant sequences. The algorithm traces the history of the sequences going back in time, encountering recombinations and coalescence (duplications) until the ancestral material is located on one sequence for homologous positions in the present sequences. Here an alternative algorithm is formulated not as going back in time and operating on sequences, but by moving spatially along the sequences, updating the history of the sequences as recombination points are encountered. This algorithm focuses on spatial aspects of the coalescent with recombination rather than on temporal aspects as is the case of familiar algorithms. Mathematical results related to spatial aspects of the coalescent with recombination are derived.

Algorithms↗

Conditional genealogies and the age of a neutral mutant.

This paper is concerned with the structure of the genealogy of a sample in which it is observed that some subset of chromosomes carries a particular mutation, assumed to have arisen uniquely in the history of the population. A rigorous theoretical study of this conditional genealogy is given using coalescent methods. Particular results include the mean, variance, and density of the age of the mutation conditional on its frequency in the sample. Most of the development relates to populations of constant size, but we discuss the extension to populations which have grown exponentially to their present size.

Alleles↗

The ancestry of a sample of sequences subject to recombination.

In this article we discuss the ancestry of sequences sampled from the coalescent with recombination with constant population size 2N. We have studied a number of variables based on simulations of sample histories, and some analytical results are derived. Consider the leftmost nucleotide in the sequences. We show that the number of nucleotides sharing a most recent common ancestor (MRCA) with the leftmost nucleotide is approximately log(1 + 4N Lr)/4Nr when two sequences are compared, where L denotes sequence length in nucleotides, and r the recombination rate between any two neighboring nucleotides per generation. For larger samples, the number of nucleotides sharing MRCA with the leftmost nucleotide decreases and becomes almost independent of 4N Lr. Further, we show that a segment of the sequences sharing a MRCA consists in mean of 3/8Nr nucleotides, when two sequences are compared, and that this decreases toward 1/4Nr nucleotides when the whole population is sampled. A measure of the correlation between the genealogies of two nucleotides on two sequences is introduced. We show analytically that even when the nucleotides are separated by a large genetic distance, but share MRCA, the genealogies will show only little correlation. This is surprising, because the time until the two nucleotides shared MRCA is reciprocal to the genetic distance. Using simulations, the mean time until all positions in the sample have found a MRCA increases logarithmically with increasing sequence length and is considerably lower than a theoretically predicted upper bound. On the basis of simulations, it turns out that important properties of the coalescent with recombinations of the whole population are reflected in the properties of a sample of low size.

Genome, Human↗

A codon-based model designed to describe lentiviral evolution.

A codon-based model designed to describe lentiviral evolution is developed. The model incorporates unequal base compositions in the three codon positions and selection against the CpG dinucleotide within codons to account for a deficit of this dinucleotide exhibited by lentiviral genes. The model is, to a large extent, able to account for the pattern of codon usage exhibited by the HIV1 genes gag, pol, and env, in spite of its parameter paucity. The model is extended to a similar model which operates on pentets (codons and their neighboring bases). The results obtained by the pentet model establish the importance of depression of CpGs across codon boundaries as well as within codons. The goodness of fit of the CpG depression model to the observed evolution in pairwise alignments of HIV1 sequences is assessed. The model provides a significantly better description of the observed evolution than the simpler models examined. The parameter estimates indicate that part of the unusually large biases in nucleotide frequencies observed in HIV1 genes is caused by selection against CpGs. We find that the estimates of expected numbers of substitutions, of transitions to transversions, and of synonymous to nonsynonymous substitution rates are robust to CpG depression, whereas the ratio of CpG-generating substitutions to other substitutions is strongly influenced by the choice of model.

Base Composition↗

On the number of ancestors to a DNA sequence.

If homologous sequences in a population are not subject to recombination, they can all be traced back to one ancestral sequence. However, the rest of our genome is subject to recombination and will be spread out on a series of individuals. The distribution of ancestral material to an extant chromosome is here investigated by the coalescent with recombination, and the results are discussed relative to humans. In an ancestral population of actual size 1.3 million a minority of <6.4% will carry material ancestral to any present human. The estimated actual population size can be even higher, 5 million, reducing the percentage to 1.7%.

DNA↗