PubMed Health⌕ Search

Biomedical subjects

Yuguo Chen

Publications and source records attributed to Yuguo Chen.

4 recordsLinked to original sources

Volume measures for linkage disequilibrium.

BACKGROUND: Defining measures of linkage disequilibrium (LD) that have good small sample properties and are applicable to multiallelic markers poses some challenges. The potential of volume measures in this context has been noted before, but their use has been hampered by computational challenges. RESULTS: We design a sequential importance sampling algorithm to evaluate volume measures on I x J tables. The algorithm is implemented in a C routine as a complement to exhaustive enumeration. We make the C code available as open source. We achieve fast and accurate evaluation of volume measures in two dimensional tables. CONCLUSION: Applying our code to simulated and real datasets reinforces the belief that volume measures are a very useful tool for LD evaluation: they are not inflated in small samples, their definition encompasses multiallelic markers, and they can be computed with appreciable speed.

Algorithms↗

Linkage disequilibrium and haplotype homozygosity in population samples genotyped at a high marker density.

OBJECTIVE: Analyze the information contained in homozygous haplotypes detected with high density genotyping. METHODS: We analyze the genotypes of approximately 2,500 markers on chr 22 in 12 population samples, each including 200 individuals. We develop a measure of disequilibrium based on haplotype homozygosity and an algorithm to identify genomic segments characterized by non-random homozygosity (NRH), taking into account allele frequencies, missing data, genotyping error, and linkage disequilibrium. RESULTS: We show how our measure of linkage disequilibrium based on homozygosity leads to results comparable to those of R(2), as well as the importance of correcting for small sample variation when evaluating D'. We observe that the regions that harbor NRH segments tend to be consistent across populations, are gene rich, and are characterized by lower recombination. CONCLUSIONS: It is crucial to take into account LD patterns when interpreting long stretches of homozygous markers.

Chromosomes, Human, Pair 22↗

Monte Carlo algorithms for Hardy-Weinberg proportions.

The Hardy-Weinberg law is among the most important principles in the study of biological systems. Given its importance, many tests have been devised to determine whether a finite population follows Hardy-Weinberg proportions. Because asymptotic tests can fail, Guo and Thompson developed an exact test; unfortunately, the Monte Carlo method they proposed to evaluate their test has a running time that grows linearly in the size of the population N. Here, we propose a new algorithm whose expected running time is linear in the size of the table produced, and completely independent of N. In practice, this new algorithm can be considerably faster than the original method.

Algorithms↗

Likelihoods from summary statistics: recent divergence between species.

We describe an importance-sampling method for approximating likelihoods of population parameters based on multiple summary statistics. In this first application, we address the demographic history of closely related members of the Drosophila pseudoobscura group. We base the maximum-likelihood estimation of the time since speciation and the effective population sizes of the extant and ancestral populations on the pattern of nucleotide variation at DPS2002, a noncoding region tightly linked to a paracentric inversion that strongly contributes to reproductive isolation. Consideration of summary statistics rather than entire nucleotide sequences permits a compact description of the genealogy of the sample. We use importance sampling first to propose a genealogical and mutational history consistent with the observed array of summary statistics and then to correct the likelihood with the exact probability of the history determined from a system of recursions. Analysis of a subset of the data, for which recursive computation of the exact likelihood was feasible, indicated close agreement between the approximate and exact likelihoods. Our results for the complete data set also compare well with those obtained through Metropolis-Hastings sampling of fully resolved genealogies of entire nucleotide sequences.

Animals↗