PubMed Health⌕ Search

Biomedical subjects

Fengzhu Sun

Publications and source records attributed to Fengzhu Sun.

10 recordsLinked to original sources

Microsatellite mutations during the polymerase chain reaction: mean field approximations and their applications.

We develop a novel mathematical model for microsatellite mutations during polymerase chain reaction (PCR). Based on the model, we study the first- and second-order moments of the number of repeat units in a randomly chosen molecule after n PCR cycles and their corresponding mean field approximations. We give upper bounds for the approximation errors and show that the approximation errors are small when the mutation rate is low. Based on the theoretical results, we develop a moment estimation method to estimate the mutation rate per-repeat-unit per PCR cycle and the probability of expansion when mutations occur. Simulation studies show that the moment estimation method can accurately recover the true mutation rate and probability of expansion. Finally, the method is applied to experimental data from single-molecule PCR experiments.

Humans↗

The relationship between microsatellite slippage mutation rate and the number of repeat units.

Microsatellite markers are widely used for genetic studies, but the relationship between microsatellite slippage mutation rate and the number of repeat units remains unclear. In this study, microsatellite distributions in the human genome are collected from public sequence databases. We observe that there is a threshold size for slippage mutations. We consider a model of microsatellite mutation consisting of point mutations and single stepwise slippage mutations. From two sets of equations based on two stochastic processes and equilibrium assumptions, we estimate microsatellite slippage mutation rates without assuming any relationship between microsatellite slippage mutation rate and the number of repeat units. We use the least squares method with constraints to estimate expansion and contraction mutation rates. The estimated slippage mutation rate increases exponentially as the number of repeat units increases. When slippage mutations happen, expansion occurs more frequently for short microsatellites and contraction occurs more frequently for long microsatellites. Our results agree with the length-dependent mutation pattern observed from experimental data, and they explain the scarcity of long microsatellites.

Algorithms↗

Haplotype block partition with limited resources and applications to human chromosome 21 haplotype data.

Recent studies have shown that the human genome has a haplotype block structure such that it can be decomposed into large blocks with high linkage disequilibrium (LD) and relatively limited haplotype diversity, separated by short regions of low LD. One of the practical implications of this observation is that only a small fraction of all the single-nucleotide polymorphisms (SNPs) (referred as "tag SNPs") can be chosen for mapping genes responsible for human complex diseases, which can significantly reduce genotyping effort, without much loss of power. Algorithms have been developed to partition haplotypes into blocks with the minimum number of tag SNPs for an entire chromosome. In practice, investigators may have limited resources, and only a certain number of SNPs can be genotyped. In the present article, we first formulate this problem as finding a block partition with a fixed number of tag SNPs that can cover the maximal percentage of the whole genome, and we then develop two dynamic programming algorithms to solve this problem. The algorithms are sufficiently flexible to permit knowledge of functional polymorphisms to be considered. We apply the algorithms to a data set of SNPs on human chromosome 21, combining the information of coding and noncoding regions. We study the density of SNPs in intergenic regions, introns, and exons, and we find that the SNP density in intergenic regions is similar to that in introns and is higher than that in exons, results that are consistent with previous studies. We also calculate the distribution of block break points in intergenic regions, genes, exons, and coding regions and do not find any significant differences.

Chromosomes, Human, Pair 21↗

A novel class of tests for the detection of mitochondrial DNA-mutation involvement in diseases.

We develop a novel class of tests to detect mitochondrial DNA (mtDNA)-mutation involvement in complex diseases by the study of affected pedigree members. For a pedigree, affected individuals are first considered and are then connected through their relatives. We construct a reduced pedigree from an original pedigree. Each configuration of a reduced pedigree is given a score, with high scores given to configurations that are consistent with mtDNA-mutation involvement and low scores given to configurations that are not consistent with mtDNA-mutation involvement. For many pedigrees, the weighted sum of scores of the pedigrees is calculated. The tests are formed by comparing the observed score with the expected score under the null hypothesis that only nuclear autosomal mutations are involved. We study the optimality of score functions and weights under the heterogeneity model without phenocopies. We also develop a method to estimate the contribution that mtDNA mutations make if they are involved under a heterogeneity model. Finally, we apply our methods to three data sets: Leber hereditary optic neuropathy, a disease that has been proved to be caused by mtDNA mutations; non-insulin-dependent diabetes mellitus (NIDDM); and hypertension (HTN). We find evidence of mtDNA-mutation involvement in all three diseases. The estimated fraction of patients with NIDDM due to mtDNA-mutation involvement is 22% (95% confidence interval [CI] 6%-38%). The fraction of patients with HTN potentially due to mtDNA-mutation involvement is estimated at 55% (95% CI 45%-65%).

DNA Mutational Analysis↗

Taq DNA polymerase slippage mutation rates measured by PCR and quasi-likelihood analysis: (CA/GT)n and (A/T)n microsatellites.

During microsatellite polymerase chain reaction (PCR), insertion-deletion mutations produce stutter products differing from the original template by multiples of the repeat unit length. We analyzed the PCR slippage products of (CA)n and (A)n tracts cloned in a pUC18 vector. Repeat numbers varied from two to 14 (CA)n and four to 12 (A)n. Data was generated on approximately 10 single molecules for each clone type using two rounds of nested PCR. The size and peak areas of the products were obtained by capillary electrophoresis. A quasi- likelihood approach to the analysis of the data estimated the mutation rate/repeat/PCR cycle. The rate for (CA)n tracts was 3.6 x 10(-3) with contractions 14 times greater than expansions. For (A)n tracts the rate was 1.5 x 10(-2) and contractions outnumbered expansions by 5-fold. The threshold for detecting 'stutter' products was computed to be four repeats for (CA)n and eight repeats for (A)n or approximately 8 bp in both cases. A comparison was made between the computationally and experimentally derived threshold values. The threshold and expansion to contraction ratios are explained on the basis of the active site structure of Taq DNA polymerase and models of the energetics of slippage events, respectively.

Base Sequence↗

The mutation process of microsatellites during the polymerase chain reaction.

We build a mathematical model for the mutation process of microsatellites during polymerase chain reaction (PCR) using the theory of branching processes. Based on the model, we develop a method to estimate the mutation rate of microsatellites per PCR cycle and the probability of expansion by maximizing a quasi-likelihood of the observed data. We show by simulations that the proposed estimation method can accurately recover the relationship between the mutation rate and number of repeat units. The theoretical basis for the proposed method is also given. We apply the method to experimental data on poly-A and poly-CA repeats.

Animals↗

Assessment of the reliability of protein-protein interactions and protein function prediction.

As more and more high-throughput protein-protein interaction data are collected, the task of estimating the reliability of different data sets becomes increasingly important. In this paper, we present our study of two groups of protein-protein interaction data, the physical interaction data and the protein complex data, and estimate the reliability of these data sets using three different measurements: (1) the distribution of gene expression correlation coefficients, (2) the reliability based on gene expression correlation coefficients, and (3) the accuracy of protein function predictions. We develop a maximum likelihood method to estimate the reliability of protein interaction data sets according to the distribution of correlation coefficients of gene expression profiles of putative interacting protein pairs. The results of the three measurements are consistent with each other. The MIPS protein complex data have the highest mean gene expression correlation coefficients (0.256) and the highest accuracy in predicting protein functions (70% sensitivity and specificity), while Ito's Yeast two-hybrid data have the lowest mean (0.041) and the lowest accuracy (15% sensitivity and specificity). Uetz's data are more reliable than Ito's data in all three measurements, and the TAP protein complex data are more reliable than the HMS-PCI data in all three measurements as well. The complex data sets generally perform better in function predictions than do the physical interaction data sets. Proteins in complexes are shown to be more highly correlated in gene expression. The results confirm that the components of a protein complex can be assigned to functions that the complex carries out within a cell. There are three interaction data sets different from the above two groups: the genetic interaction data, the in-silico data and the syn-express data. Their capability of predicting protein functions generally falls between that of the Y2H data and that of the MIPS protein complex data. The supplementary information is available at the following Web site: http://www-hto.usc.edu/-msms/AssessInteraction/.

Computational Biology↗

Haplotype block structure and its applications to association studies: power and study designs.

Recent studies have shown that the human genome has a haplotype block structure, such that it can be divided into discrete blocks of limited haplotype diversity. In each block, a small fraction of single-nucleotide polymorphisms (SNPs), referred to as "tag SNPs," can be used to distinguish a large fraction of the haplotypes. These tag SNPs can potentially be extremely useful for association studies, in that it may not be necessary to genotype all SNPs; however, this depends on how much power is lost. Here we develop a simulation study to quantitatively assess the power loss for a variety of study designs, including case-control designs and case-parental control designs. First, a number of data sets containing case-parental or case-control samples are generated on the basis of a disease model. Second, a small fraction of case and control individuals in each data set are genotyped at all the loci, and a dynamic programming algorithm is used to determine the haplotype blocks and the tag SNPs based on the genotypes of the sampled individuals. Third, the statistical power of tests was evaluated on the basis of three kinds of data: (1) all of the SNPs and the corresponding haplotypes, (2) the tag SNPs and the corresponding haplotypes, and (3) the same number of randomly chosen SNPs as the number of tag SNPs and the corresponding haplotypes. We study the power of different association tests with a variety of disease models and block-partitioning criteria. Our study indicates that the genotyping efforts can be significantly reduced by the tag SNPs, without much loss of power. Depending on the specific haplotype block-partitioning algorithm and the disease model, when the identified tag SNPs are only 25% of all the SNPs, the power is reduced by only 4%, on average, compared with a power loss of approximately 12% when the same number of randomly chosen SNPs is used in a two-locus haplotype analysis. When the identified tag SNPs are approximately 14% of all the SNPs, the power is reduced by approximately 9%, compared with a power loss of approximately 21% when the same number of randomly chosen SNPs is used in a two-locus haplotype analysis. Our study also indicates that haplotype-based analysis can be much more powerful than marker-by-marker analysis.

Algorithms↗

A dynamic programming algorithm for haplotype block partitioning.

We develop a dynamic programming algorithm for haplotype block partitioning to minimize the number of representative single nucleotide polymorphisms (SNPs) required to account for most of the common haplotypes in each block. Any measure of haplotype quality can be used in the algorithm and of course the measure should depend on the specific application. The dynamic programming algorithm is applied to analyze the chromosome 21 haplotype data of Patil et al. [Patil, N., Berno, A. J., Hinds, D. A., Barrett, W. A., Doshi, J. M., Hacker, C. R., Kautzer, C. R., Lee, D. H., Marjoribanks, C., McDonough, D. P., et al. (2001) Science 294, 1719-1723], who searched for blocks of limited haplotype diversity. Using the same criteria as in Patil et al., we identify a total of 3,582 representative SNPs and 2,575 blocks that are 21.5% and 37.7% smaller, respectively, than those identified using a greedy algorithm of Patil et al. We also apply the dynamic programming algorithm to the same data set based on haplotype diversity. A total of 3,982 representative SNPs and 1,884 blocks are identified to account for 95% of the haplotype diversity in each block.

Algorithms↗

Inferring domain-domain interactions from protein-protein interactions.

The interaction between proteins is one of the most important features of protein functions. Behind protein-protein interactions there are protein domains interacting physically with one another to perform the necessary functions. Therefore, understanding protein interactions at the domain level gives a global view of the protein interaction network, and possibly of protein functions. Two research groups used yeast two-hybrid assays to generate 5719 interactions between proteins of the yeast Saccharomyces cerevisiae. This allows us to study the large-scale conserved patterns of interactions between protein domains. Using evolutionarily conserved domains defined in a protein-domain database called PFAM (http://PFAM.wustl.edu), we apply a Maximum Likelihood Estimation method to infer interacting domains that are consistent with the observed protein-protein interactions. We estimate the probabilities of interactions between every pair of domains and measure the accuracies of our predictions at the protein level. Using the inferred domain-domain interactions, we predict interactions between proteins. Our predicted protein-protein interactions have a significant overlap with the protein-protein interactions (MIPS: http://mips.gfs.de) obtained by methods other than the two-hybrid assays. The mean correlation coefficient of the gene expression profiles for our predicted interaction pairs is significantly higher than that for random pairs. Our method has shown robustness in analyzing incomplete data sets and dealing with various experimental errors. We found several novel protein-protein interactions such as RPS0A interacting with APG17 and TAF40 interacting with SPT3, which are consistent with the functions of the proteins.

Computational Biology↗