PubMed Health⌕ Search

Biomedical subjects

M K Kuhner

Publications and source records attributed to M K Kuhner.

13 recordsLinked to original sources

Sampling among haplotype resolutions in a coalescent-based genealogy sampler.

Analysis of the coalescent structure of a population may provide information useful in mapping disease loci. Current coalescent-based genealogy samplers require haplotyped data, but haplotypes are not always available, and it is not practical to sum over all haplotype assignments for large data sets. We describe a method of adding haplotype re-evaluation to the sampler, so that it samples not only among genealogies explaining a given haplotype configuration, but also among different haplotype configurations. Several different haplotype-rearrangement strategies are considered, but the simplest-inverting the phase of a single site in a single individual-appears to be the most successful. The straightforward haplotype sampler does not mix well; heating approaches can greatly improve its performance.

Computer Simulation↗

Usefulness of single nucleotide polymorphism data for estimating population parameters.

Single nucleotide polymorphism (SNP) data can be used for parameter estimation via maximum likelihood methods as long as the way in which the SNPs were determined is known, so that an appropriate likelihood formula can be constructed. We present such likelihoods for several sampling methods. As a test of these approaches, we consider use of SNPs to estimate the parameter Theta = 4N(e)micro (the scaled product of effective population size and per-site mutation rate), which is related to the branch lengths of the reconstructed genealogy. With infinite amounts of data, ML models using SNP data are expected to produce consistent estimates of Theta. With finite amounts of data the estimates are accurate when Theta is high, but tend to be biased upward when Theta is low. If recombination is present and not allowed for in the analysis, the results are additionally biased upward, but this effect can be removed by incorporating recombination into the analysis. SNPs defined as sites that are polymorphic in the actual sample under consideration (sample SNPs) are somewhat more accurate for estimation of Theta than SNPs defined by their polymorphism in a panel chosen from the same population (panel SNPs). Misrepresenting panel SNPs as sample SNPs leads to large errors in the maximum likelihood estimate of Theta. Researchers collecting SNPs should collect and preserve information about the method of ascertainment so that the data can be accurately analyzed.

Computer Simulation↗

Maximum likelihood estimation of recombination rates from population data.

We describe a method for co-estimating r = C/mu (where C is the per-site recombination rate and mu is the per-site neutral mutation rate) and Theta = 4N(e)mu (where N(e) is the effective population size) from a population sample of molecular data. The technique is Metropolis-Hastings sampling: we explore a large number of possible reconstructions of the recombinant genealogy, weighting according to their posterior probability with regard to the data and working values of the parameters. Different relative rates of recombination at different locations can be accommodated if they are known from external evidence, but the algorithm cannot itself estimate rate differences. The estimates of Theta are accurate and apparently unbiased for a wide range of parameter values. However, when both Theta and r are relatively low, very long sequences are needed to estimate r accurately, and the estimates tend to be biased upward. We apply this method to data from the human lipoprotein lipase locus.

Genetics, Population↗

Maximum likelihood estimation of population growth rates based on the coalescent.

We describe a method for co-estimating 4Nemu (four times the product of effective population size and neutral mutation rate) and population growth rate from sequence samples using Metropolis-Hastings sampling. Population growth (or decline) is assumed to be exponential. The estimates of growth rate are biased upwards, especially when 4Nemu is low; there is also a slight upwards bias in the estimate of 4Nemu itself due to correlation between the parameters. This bias cannot be attributed solely to Metropolis-Hastings sampling but appears to be an inherent property of the estimator and is expected to appear in any approach which estimates growth rate from genealogy structure. Sampling additional unlinked loci is much more effective in reducing the bias than increasing the number or length of sequences from the same locus.

Forecasting↗

Estimating effective population size and mutation rate from sequence data using Metropolis-Hastings sampling.

We present a new way to make a maximum likelihood estimate of the parameter 4N mu (effective population size times mutation rate per site, or theta) based on a population sample of molecular sequences. We use a Metropolis-Hastings Markov chain Monte Carlo method to sample genealogies in proportion to the product of their likelihood with respect to the data and their prior probability with respect to a coalescent distribution. A specific value of theta must be chosen to generate the coalescent distribution, but the resulting trees can be used to evaluate the likelihood at other values of theta, generating a likelihood curve. This procedure concentrates sampling on those genealogies that contribute most of the likelihood, allowing estimation of meaningful likelihood curves based on relatively small samples. The method can potentially be extended to cases involving varying population size, recombination, and migration.

Base Sequence↗

A simulation comparison of phylogeny algorithms under equal and unequal evolutionary rates.

Using simulated data, we compared five methods of phylogenetic tree estimation: parsimony, compatibility, maximum likelihood, Fitch-Margoliash, and neighbor joining. For each combination of substitution rates and sequence length, 100 data sets were generated for each of 50 trees, for a total of 5,000 replications per condition. Accuracy was measured by two measures of the distance between the true tree and the estimate of the tree, one measure sensitive to accuracy of branch lengths and the other not. The distance-matrix methods (Fitch-Margoliash and neighbor joining) performed best when they were constrained from estimating negative branch lengths; all comparisons with other methods used this constraint. Parsimony and compatibility had similar results, with compatibility generally inferior; Fitch-Margoliash and neighbor joining had similar results, with neighbor joining generally slightly inferior. Maximum likelihood was the most successful method overall, although for short sequences Fitch-Margoliash and neighbor joining were sometimes better. Bias of the estimates was inferred by measuring whether the independent estimates of a tree for different data sets were closer to the true tree than to each other. Parsimony and compatibility had particular difficulty with inaccuracy and bias when substitution rates varied among different branches. When rates of evolution varied among different sites, all methods showed signs of inaccuracy and bias.

Algorithms↗

Genetic exchange in the evolution of the human MHC class II loci.

A total of 61 DNA sequences from human major histocompatibility class II loci were searched for statistical evidence of past genetic exchange (gene conversion or recombination). Among the 12 A-locus sequences (derived from DPA1 and DQA1), 4 clusters indicating potential exchange events were found. Among the 49 B-locus sequences (derived from DOB, DPB1, DPB2, DQB1, DRB1, DRB3, DRB4 and DRB5), 15 clusters were found. The clusters suggested short exchanges (less than 100 bp) within and between loci, and were concentrated in exon 2 (coding for the antigen binding site). The most striking feature of the results was the presence of an approximately 200-bp region in the middle of B-locus exon 2 which contained almost no locus-specific substitutions, which were abundant elsewhere. This suggests either strong selection for locus specificity in the other regions of the gene or a history of frequent between-locus exchange in this part of exon 2, which is involved in forming the antigen binding site.

Algorithms↗

Gene conversion in the evolution of the human and chimpanzee MHC class I loci.

Sixty-five DNA sequences from human and chimpanzee major histocompatibility complex class I loci were searched for statistical evidence of past gene conversion. Twenty-four potential conversions were detected; they were distributed across both variable and conserved portions of the gene, and involved both classical and non-classical loci. The majority spanned less than 100 bp, comparable in length to the conversions observed in spontaneous mutations in mice. Both within-locus and between-locus conversions were observed. Certain areas of the antigen recognition site appear to have been the target for multiple conversion events. The implications of these findings for the evolution of the class I multigene family are discussed.

Alleles↗

Clues to IDDM pathogenesis from genetic and serological traits in multiply affected families.

A scheme is outlined for analyzing the genotypic contributions of two unlinked loci in producing a disease, using DR and the 5' insulin locus (INS) in insulin-dependent diabetes mellitus (IDDM) as examples. Although genotypes of both DR and INS play roles in IDDM susceptibility, both the relatively small size of the Genetic Analysis Workshop 5 (GAW5) data set and the apparently limited magnitudes of the contributory effects prevent the identification of the exact nature of the association of these two loci in disease causation. The Gm allotypes showed no association with IDDM, either alone or in combination with other variables. Association of reactivity among the six strains of Coxsackie B virus is described, with no evidence of associations with DR type and IDDM found. The unaffected offspring segregated DR alleles according to expectations, while the segregation of affected alleles revealed the various contributions of DR alleles to IDDM pathogenesis, with the suggestion that DR4 from fathers is more diabetogenic than that from mothers. Lastly, a method is described for revealing the accuracy of typing in family data, and applied to RFLP variants subdividing DR3.

Cohort Studies↗

HLA and insulin gene associations with IDDM.

The HLA DR genotype frequencies in insulin-dependent diabetes mellitus (IDDM) patients and the frequencies of DR alleles transmitted from affected parent to affected child both indicate that the DR3-associated predisposition is more "recessive" and the DR4-associated predisposition more "dominant" in inheritance after allowing for the DR3/DR4 synergistic effect. B locus distributions on patient haplotypes indicate that only subsets of both DR3 and DR4 are predisposing. Heterogeneity is detected for both the DR3 and DR4 predisposing haplotypes based on DR genotypic class. With appropriate use of the family structure of the data a control population of "unaffected" alleles can be defined. Application of this method confirms the predisposing effect associated with the class 1 allele of the polymorphic region 5' to the insulin gene.

Child↗

Genetic heterogeneity, modes of inheritance, and risk estimates for a joint study of Caucasians with insulin-dependent diabetes mellitus.

From 11 studies, a total of 1,792 Caucasian probands with insulin-dependent diabetes mellitus (IDDM) are analyzed. Antigen genotype frequencies in patients, transmission from affected parents to affected children, and the relative frequencies of HLA-DR3 and -DR4 homozygous patients all indicate that DR3 predisposes in a "recessive"-like and DR4 in a "dominant"-like or "intermediate" fashion, after allowing for the DR3/DR4 synergistic effect. Removal of DR3 and DR4 reveals an overall protective effect of DR2, predisposing effects of DR1 and DRw8, and a slight protective effect of DR5 and a predisposing effect of DRw6. Analysis of affected-parent-to-affected-child data indicates that a subset of DR2 may predispose. The non-DR3, non-DR4 antigens are not independently associated with DR3 and DR4; the largest effect is a deficiency of DR2, followed by excesses of DR1, DRw8, and DRw6, in DR4 individuals, as compared with DR3 individuals. HLA-B locus distributions on patient haplotypes indicate that only subsets of both DR3 and DR4 are predisposing. The presence or absence of Asp at position 57 of the DQ beta gene, recently implicated in IDDM predisposition, is not by itself sufficient to explain the inheritance of IDDM. At a minimum, the distinguishing features of the DR3-associated and DR4-associated predisposition remain to be identified at the molecular level. Risk estimates for sibs of probands are calculated based on an overall sibling risk of 6%; estimates for those sharing two, one, or zero haplotypes are 12.9%, 4.5%, and 1.8%, respectively. Risk estimates subdivided by the DR type of the proband are also calculated, the highest being 19.2% for sibs sharing two haplotypes with a DR3/DR4 proband.

Adolescent↗