Long-term persistence of bacterial DNA.
Explore the source record for details and available documents.
Biomedical subjects
Publications and source records attributed to Carsten Wiuf.
Explore the source record for details and available documents.
In this study compatibility with a tree for unphased genotype data is discussed. If the data are compatible with a tree, the data are consistent with an assumption of no recombination in its evolutionary history. Further, it is said that there is a solution to the perfect phylogeny problem; i.e., for each individual a pair of haplotypes can be defined and the set of all haplotypes can be explained without invoking recombination. A new algorithm to decide whether or not a sample is compatible with a tree is derived. The new algorithm relies on an equivalence relation between sites that mutually determine the phase of each other. (The previous algorithm was based on advanced graph theoretical tools.) The equivalence relation is used to derive the number of solutions to the perfect phylogeny problem. Further, a series of statistics, R ( j ) ( M ), j >or= 2, are defined. These can be used to detect recombination events in the sample's history and to divide the sample into regions that are compatible with a tree. The new statistics are applied to real data from human genes. The results from this application are discussed with reference to recent suggestions that recombination in the human genome is highly heterogeneous.
Genetic analyses of permafrost and temperate sediments reveal that plant and animal DNA may be preserved for long periods, even in the absence of obvious macrofossils. In Siberia, five permafrost cores ranging from 400,000 to 10,000 years old contained at least 19 different plant taxa, including the oldest authenticated ancient DNA sequences known, and megafaunal sequences including mammoth, bison, and horse. The genetic data record a number of dramatic changes in the taxonomic diversity and composition of Beringian vegetation and fauna. Temperate cave sediments in New Zealand also yielded DNA sequences of extinct biota, including two species of ratite moa, and 29 plant taxa characteristic of the prehuman environment. Therefore, many sedimentary deposits may contain unique, and widespread, genetic records of paleoenvironments.
SUMMARY: A bioinformatic tool was written to simulate haplotypes and SNPs under a modified coalescent with recombination. The most important feature of this program is that it allows for the specification of non-homogeneous recombination rates, which results in the formation of the so-called 'haplotype blocks' of the human genome. The program also implements different mutation models and flexible demographic histories. The samples generated can be very useful to better understand the architecture of the human genome and to investigate its impact in association studies searching for disease genes. AVAILABILITY: The SNPsim package is available at http://www.evolgenics.com/software
Inference about population history from DNA sequence data has become increasingly popular. For human populations, questions about whether a population has been expanding and when expansion began are often the focus of attention. For viral populations, questions about the epidemiological history of a virus, e.g., HIV-1 and Hepatitis C, are often of interest. In this paper I address the following question: Can population history be accurately inferred from single locus DNA data? An idealised world is considered in which the tree relating a sample of n non-recombining and selectively neutral DNA sequences is observed, rather than just the sequences themselves. This approach provides an upper limit to the information that possibly can be extracted from a sample. It is shown, based on Kingman's (1982a) coalescent process, that consistent estimation of parameters describing population history (e.g., a growth rate) cannot be achieved for increasing sample size, n. This is worse than often found for estimators of genetic parameters, e.g., the mutation rate typically converges at rate under the assumption that all historical mutations can be observed in the sample. In addition, various results for the distribution of maximum likelihood estimators are presented.
Tagging haplotypes with a small number of genetic markers is becoming an increasingly interesting and important problem. Surprisingly little work has been done to characterize the mathematical framework of this problem. In this paper we present a mathematical frame, based on Boolean algebras, that adequately describe the structure of a set of genetic bi-allelic markers and the corresponding set of haplotypes. We derive a number of results that relate the number of markers required to tag a set of haplotypes to the set of markers themselves.
Recent experimental findings suggest that the assumption of a homogeneous recombination rate along the human genome is too naive. These findings point to block-structured recombination rates; certain regions (called hotspots) are more prone than other regions to recombination. In this report a coalescent model incorporating hotspot or block-structured recombination is developed and investigated analytically as well as by simulation. Our main results can be summarized as follows: (1) The expected number of recombination events is much lower in a model with pure hotspot recombination than in a model with pure homogeneous recombination, (2) hotspots give rise to large variation in recombination rates along the genome as well as in the number of historical recombination events, and (3) the size of a (nonrecombining) block in the hotspot model is likely to be overestimated grossly when estimated from SNP data. The results are discussed with reference to the current debate about block-structured recombination and, in addition, the results are compared to genome-wide variation in recombination rates. A number of new analytical results about the model are derived.
In this article I derive an alternative algorithm to Hudson and Kaplan's (Genetics 111, 147-165) algorithm that gives a lower bound to the number of recombination events in a sample's history. It is shown that the number, T(M), found by the algorithm is the least number of topologies required to explain a set of DNA sequences sampled under the infinite-site assumption. Let Tao = (T(1),...,T(r)) be a list of topologies compatible with the sequences, i.e., T(k) is compatible with an interval, I(k), of sites in the alignment. A characterization of all lists having T(M) topologies is given and it is shown that T(M) relates to specific patterns in the alignment, here called chain series. Further, a number of theorems relating general lists of topologies to the number T(M) is presented. The results are discussed in relation to the true minimum number of recombination events required to explain an alignment.