PubMed Health⌕ Search

Biomedical subjects

Yun-Xin Fu

Publications and source records attributed to Yun-Xin Fu.

8 recordsLinked to original sources

Investigating single nucleotide polymorphism (SNP) density in the human genome and its implications for molecular evolution.

We investigated the single nucleotide polymorphism (SNP) density across the human genome and in different genic categories using two SNP databases: Celera's CgsSNP, which includes SNPs identified by comparing genomic sequences, and Celera's RefSNP, which includes SNPs from a variety of sources and is biased toward disease-associated genes. Based on CgsSNP, the average numbers of SNPs per 10 kb was 8.33, 8.44, and 8.09 in the human genome, in intergenic regions, and in genic regions, respectively. In genic regions, the SNP density in intronic, exonic and adjoining untranslated regions was 8.21, 5.28, and 7.51 SNPs per 10 kb, respectively. The pattern of SNP density based on RefSNP was different from that based on CgsSNP, emphasizing its utility for genotype-phenotype association studies but not for most population genetic studies. The number of SNPs per chromosome was correlated with chromosome length, but the density of SNPs estimated by CgsSNP was not significantly correlated with the GC content of the chromosome. Based on CgsSNP, the ratio of nonsense to missense mutations (0.027), the ratio of missense to silent mutations (1.15), and the ratio of non-synonymous to synonymous mutations (1.18) was less than half of that expected in a human protein coding sequence under the neutral mutation theory, reflecting a role for natural selection, especially purifying selection.

DNA, Intergenic↗

Comparison of strategies for selecting single nucleotide polymorphisms for case/control association studies.

It is widely believed that a subset of single nucleotide polymorphisms (SNPs) is able to capture the majority of the information for genotype-phenotype association studies that is contained in the complete compliment of genetic variations. The question remains, how does one select that particular subset of SNPs in order to maximize the power of detecting a significant association? In this study, we have used a simulation approach to compare three competing methods of site selection: random selection, selection based on pair-wise linkage disequilibrium, and selection based on maximizing haplotype diversity. The results indicate that site selection based on maximizing haplotype diversity is preferred over random selection and selection based on pair-wise linkage disequilibrium. The results also indicate that it is more prudent to increase the sample size to improve a study's power than to continuously increase the number of SNPs. These results have direct implications for designing gene-based and genome-wide association studies.

Case-Control Studies↗

Neutrality tests using DNA polymorphism from multiple samples.

The polymorphism of a gene or a locus is studied with increasing frequency by multiple laboratories or the same group at different times. Such practice results in polymorphism being revealed by different samples at different regions of the locus. Tests of neutrality have been widely conducted for polymorphism data but commonly used statistical tests cannot be applied directly to such data. This article provides a procedure to conduct a neutrality test and details are given for two commonly used tests. Applying the two new tests to the chemokine-receptor gene (CCR5) in humans, we found that the hypothesis that all mutations are selectively neutral cannot explain the observed pattern of DNA polymorphism.

Chromosome Mapping↗

Genetic diversity and population history of golden monkeys (Rhinopithecus roxellana).

Golden monkey (Rhinopithecus roxellana), namely the snub-nosed monkey, is a well-known endangered primate, which distributes only in the central part of mainland China. As an effort to understand the current genetic status as well as population history of this species, we collected a sample of 32 individuals from four different regions, which cover the major habitat of this species. Forty-four allozyme loci were surveyed in our study by allozyme electrophoresis, none of which was found to be polymorphic. The void of polymorphism compared with that of other nonhuman primates is surprising particularly considering that the current population size is many times larger than that of some other endangered species. Since many independent loci are surveyed in this study, the most plausible explanation for our observation is that the population has experienced a recent bottleneck. We used a coalescent approach to explore various scenarios of population bottleneck and concluded that the most recent bottleneck could have happened within the last 15,000 years. Moreover, the proposed simulation approach could be useful to researchers who need to analyze the non- or low-polymorphism data.

Animals↗

Estimating mutation rate: how to count mutations?

Mutation rate is an essential parameter in genetic research. Counting the number of mutant individuals provides information for a direct estimate of mutation rate. However, mutant individuals in the same family can share the same mutations due to premeiotic mutation events, so that the number of mutant individuals can be significantly larger than the number of mutation events observed. Since mutation rate is more closely related to the number of mutation events, whether one should count only independent mutation events or the number of mutants remains controversial. We show in this article that counting mutant individuals is a correct approach for estimating mutation rate, while counting only mutation events will result in underestimation. We also derived the variance of the mutation-rate estimate, which allows us to examine a number of important issues about the design of such experiments. The general strategy of such an experiment should be to sample as many families as possible and not to sample much more offspring per family than the reciprocal of the pairwise correlation coefficient within each family. To obtain a reasonably accurate estimate of mutation rate, the number of sampled families needs to be in the same or higher order of magnitude as the reciprocal of the mutation rate.

Humans↗

Genetic relationship of Chinese ethnic populations revealed by mtDNA sequence diversity.

The origin and demographic history of the ethnic populations of China have not been clearly resolved. In this study, we examined the hypervariable segment I sequences (HVSI) of the mitochondrial DNA control region in 372 individuals from nine Chinese populations and one northern Thai population. A relatively high percentage of individuals was found to share sequences with those from other populations of the same ethnogenesis. In general, the populations of southern or Pai-Yuei tribal origin showed high haplotype diversity and nucleotide diversity compared with the populations of northern or Di-Qiang tribal origin. Mismatch distributions from these populations showed concordant features. All except the northern groups Nu, Lisu, Tibetan, and Mongolian showed typical signatures of ancient population expansions in the mismatch distributions and neutrality tests. Episodes of extreme size reduction in the past are one of the likely explanations for the absence of evidence of expansion in northern populations. Small sample sizes as well as samples from isolated subpopulations contributed to the bumpy mismatch distributions observed. Phylogenetic analysis and haplotype sharing among populations suggest that current mtDNA variation in these ethnic populations could reveal their ethnohistory to some extent, but in general, linguistic and geographic classifications of the populations did not agree well with classification by mtDNA variation.

Adult↗

DNA polymorphism in a worldwide sample of human X chromosomes.

DNA sequence data from humans can provide insight into the history of modern humans and the genetic variability in human populations. We report here a study of human DNA sequence variation at an X-linked noncoding region of 10,346 bp. The sample consists of 62 X chromosomes from Africa, Europe, and Asia. Forty-four polymorphic sites were found among the 62 sequences, resulting in 23 different haplotypes. Statistical analyses of the data led to the following inferences. (1) There is strong evidence of human population expansion in the relatively recent past, and this population expansion has had a significant effect on the pattern of polymorphism at this locus. (2) Non-African populations were unlikely to have been derived from a very small number of African lineages. (3) There was considerable geographic subdivision in the ancient human population, which could be an important reason why many studies failed to detect population expansion. (4) The long-term effective population size of humans is between 12,000 and 15,000. And (5) a non-African specific variant was found at a frequency of 35% in non-Africans, an estimate supported by the genotyping of additional 80 non-African and 106 African X chromosomes. This variant could have arisen in Eurasia more than 140,000 years ago, predating the emergence of modern humans. Moreover, this haplotype and all other haplotypes coalesced to the most recent common ancestor of the sample, which was estimated to be older than 490,000 years. Therefore, this region may have a long history in Eurasia.

Chromosomes, Human, X↗