PubMed HealthSearch

Biomedical subjects

N L Kaplan

Publications and source records attributed to N L Kaplan.

At least 19 recordsLinked to original sources

Tests for linkage and association in nuclear families.

The transmission/disequilibrium test (TDT) originally was introduced to test for linkage between a genetic marker and a disease-susceptibility locus, in the presence of association. Recently, the TDT has been used to test for association in the presence of linkage. The motivation for this is that linkage analysis typically identifies large candidate regions, and further refinement is necessary before a search for the disease gene is begun, on the molecular level. Evidence of association and linkage may indicate which markers in the region are closest to a disease locus. As a test of linkage, transmissions from heterozygous parents to all of their affected children can be included in the TDT; however, the TDT is a valid chi2 test of association only if transmissions to unrelated affected children are used in the analysis. If the sample contains independent nuclear families with multiple affected children, then one procedure that has been used to test for association is to select randomly a single affected child from each sibship and to apply the TDT to those data. As an alternative, we propose two statistics that use data from all of the affected children. The statistics give valid chi2 tests of the null hypothesis of no association or no linkage and generally are more powerful than the TDT with a single, randomly chosen, affected child from each family.

Alleles

Power studies for the transmission/disequilibrium tests with multiple alleles.

Case-control studies compare marker-allele distributions in affected and unaffected individuals, and significant results suggest linkage but may simply reflect population structure. For markers with m alleles (m > or = 2), a McNemar-like statistic, I, estimates the level of population association between marker and disease loci. To test for linkage after significant case-control tests, within-family tests are performed. These operate on the contingency table, with i, jth element equal to the number of parents that transmit marker allele Mi and do not transmit marker allele Mi to an affected offspring. The dimension of the table is the number of alleles at the marker locus. Three test statistics have recently been proposed in the literature: Tc compares symmetric pairs of cells (i, j) and (j, i), Tm compares row and column totals for the same marker allele, and a likelihood ratio statistic Tl uses all the cells in the table. In addition, we consider a new statistic, Tmhet, that uses only the heterozygous parents and is approximately chi2 with (m - 1) df. We use a Monte Carlo test to guarantee valid tests and to demonstrate the inferiority of Tc and the equality of Tm and Tl in terms of power. The power of the Tmhet test is close but not always equal to the power of the Tm test. We also show that under the alternative hypothesis of linkage, Tm is approximately noncentral chi2 with (m - 1) df and noncentrality parameter 2NT(1 - 2theta)2I*, when data on single affecteds in NT families are used. If the disease has a low population frequency, then I* is estimated using the case-control statistic I. This offers a basis for choosing sample size, or choosing a marker system.

Alleles

The coalescent process and background selection.

Some statistical properties of gene trees are described for a model with background deleterious mutations. It is argued that the history of a small sample of genes under this model is a continuous time Markov chain that quickly reaches stationarity. This observation leads to simple expressions for the expected nucleotide diversity and suggests that the frequency spectrum in small samples should be approximately the same as under a strict neutral model. The results concerning expected nucleotide diversity are compared with observed variation on the third chromosome of Drosophila melanogaster.

Animals

The hitchhiking effect on the site frequency spectrum of DNA polymorphisms.

The level of DNA sequence variation is reduced in regions of the Drosophila melanogaster genome where the rate of crossing over per physical distance is also reduced. This observation has been interpreted as support for the simple model of genetic hitchhiking, in which directional selection on rare variants, e.g., newly arising advantageous mutants, sweeps linked neutral alleles to fixation, thus eliminating polymorphisms near the selected site. However, the frequency spectra of segregating sites of several loci from some populations exhibiting reduced levels of nucleotide diversity and reduced numbers of segregating sites did not appear different from what would be expected under a neutral equilibrium model. Specifically, a skew toward an excess of rare sites was not observed in these samples, as measured by Tajima's D. Because this skew was predicted by a simple hitchhiking model, yet it had never been expressed quantitatively and compared directly to DNA polymorphism data, this paper investigates the hitchhiking effect on the site frequency spectrum, as measured by Tajima's D and several other statistics, using a computer simulation model based on the coalescent process and recurrent hitchhiking events. The results presented here demonstrate that under the simple hitchhiking model (1) the expected value of Tajima's D is large and negative (indicating a skew toward rare variants), (2) that Tajima's test has reasonable power to detect a skew in the frequency spectrum for parameters comparable to those from actual data sets, and (3) that the Tajima's Ds observed in several data sets are very unlikely to have been the result of simple hitchhiking. Consequently, the simple hitchhiking model is not a sufficient explanation for the DNA polymorphism at those loci exhibiting a decreased number of segregating sites yet not exhibiting a skew in the frequency spectrum.

Animals

Deleterious background selection with recombination.

An analytic expression for the expected nucleotide diversity is obtained for a neutral locus in a region with deleterious mutation and recombination. Our analytic results are used to predict levels of variation for the entire third chromosome of Drosophila melanogaster. The predictions are consistent with the low levels of variation that have been observed at loci near the centromeres of the third chromosome of D. melanogaster. However, the low levels of variation observed near the tips of this chromosome are not predicted using currently available estimates of the deleterious mutation rate and of selection coefficients. If considerably smaller selection coefficients are assumed, the low observed levels of variation at the tips of the third chromosome are consistent with the background selection model.

Animals

Likelihood methods for locating disease genes in nonequilibrium populations.

Until recently, attempts to map disease genes on the basis of population associations with linked markers have been based on expected values of linkage disequilibrium. These methods suffer from the large variances imposed on disequilibrium measures by the evolutionary process, but a more serious problem for many diseases is that they assume an equilibrium population. For diseases that arose only a few hundred generations ago, it is more appropriate to concentrate on the initial growth phase of the disease. We invoke a Poisson branching process for this early growth, and estimate the likelihood for the recombination fraction between marker and disease loci, on the basis of simulated disease populations. The limits of the resulting support intervals for the recombination fraction vary inversely with the age of the disease in generations. We illustrate the procedure with data on cystic fibrosis and diastrophic dysplasia, for which the method appears appropriate, and for Friedreich ataxia and Huntington disease, for which it does not. A valuable aspect of the method is the ability in some cases to compare likelihoods of the three orders for a disease locus and two linked marker loci.

Chromosome Mapping

Are moment bounds on the recombination fraction between a marker and a disease locus too good to be true? Allelic association mapping revisited for simple genetic diseases in the Finnish population.

In the past several years, allelic association has helped map a number of rare genetic diseases in the human genome. A commonly used upper bound on the recombination fraction between the disease gene and an associated marker is known to be biased downward, so there is the possibility that an investigator could be misled. This upper bound is based on a moment equation that can be derived within the context of a Poisson branching process, so its performance can be compared with a recently proposed likelihood bound. We show that the confidence level of the moment upper bound is much lower than expected, while the confidence level of the likelihood bound is in line with expectation. The effects of mutation at either the marker or disease locus on the upper bounds are also investigated. Results indicate that mutation is not an important force for typical mutation rates, unless the recombination fraction between the marker and disease locus is very small or the disease allele is very rare in the general population. Finally, the impact of sample size on the likelihood bound is investigated. The results are illustrated with data on 10 simple genetic diseased in the Finnish population.

Alleles

A statistical test for detecting geographic subdivision.

A statistical test for detecting genetic differentiation of subpopulations is described that uses molecular variation in samples of DNA sequences from two or more localities. The statistical significance of the test is determined with Monte Carlo simulations. The power of the test to detect genetic differentiation in a selectively neutral Wright-Fisher island model depends on both sample size and the rates of migration, mutation, and recombination. It is found that the power of the test is substantial with samples of size 50, when 4Nm less than 10, where N is the subpopulation size and m is the fraction of migrants in each subpopulation each generation. More powerful tests are obtained with genes with recombination than with genes without recombination.

Alcohol Dehydrogenase

Biologically based models for risk assessment.

The modelling problems associated with the estimation of risks from long-term chemical exposures at low dose levels represent a statistical and mathematical challenge with special relevance to environmental research. Determining an adequate model for estimating the relationship between dose and response is critical to reducing potential bias in the risk estimation process. This paper discusses the various assumptions and models used in carcinogenic risk assessment. The emphasis is on our ability to accurately determine the magnitude of the carcinogenic risk, the shape of the dose-response relationship and the overall variability of the risk estimates.

Animals

A numerical method for calculating moments of coalescent times in finite populations with selection.

A numerical method is developed for solving a nonstandard singular system of second-order differential equations arising from a problem in population genetics concerning the coalescent process for a sample from a population undergoing selection. The nonstandard feature of the system is that there are terms in the equations that approach infinity as one approaches the boundary. The numerical recipe is patterned after the LU decomposition for tridiagonal matrices. Although there is no analytic proof that this method leads to the correct solution, various examples are presented that suggest that the method works. This method allows one to calculate the expected number of segregating sites in a random sample of n genes from a population whose evolution is described by a model which is not selectively neutral.

Genetics, Population

Variability of safe dose estimates when using complicated models of the carcinogenic process. A case study: methylene chloride.

Advances in understanding carcinogenesis have led to the development of mathematical models that have biologically interpretable parameters. These models utilize more of the available scientific data than the empirical models routinely employed for quantifying carcinogenic risk. They also require consideration of sources of uncertainty in risk estimates that were previously ignored, such as animal-to-animal variability of physiological and pharmacological constants. A numerical technique is proposed for studying the consequences of incorporating the intrapopulation variability of biologically interpretable parameters into the risk assessment process. To demonstrate the technique, the variability of safe dose estimates for exposure to methylene chloride is considered. The results suggest that intrapopulation variability of the model parameters can increase the variability of safe dose estimates an appreciable amount.

Animals

The "hitchhiking effect" revisited.

The number of selectively neutral polymorphic sites in a random sample of genes can be affected by ancestral selectively favored substitutions at linked loci. The degree to which this happens depends on when in the history of the sample the selected substitutions happen, the strength of selection and the amount of crossing over between the sampled locus and the loci at which the selected substitutions occur. This phenomenon is commonly called hitchhiking. Using the coalescent process for a random sample of genes from a selectively neutral locus that is linked to a locus at which selection is taking place, a stochastic, finite population model is developed that describes the steady state effect of hitchhiking on the distribution of the number of selectively neutral polymorphic sites in a random sample. A prediction of the model is that, in regions of low crossing over, strongly selected substitutions in the history of the sample can substantially reduce the number of polymorphic sites in a random sample of genes from that expected under a neutral model.

Base Sequence

The coalescent process in models with selection.

Statistical properties of the process describing the genealogical history of a random sample of genes are obtained for a class of population genetics models with selection. For models with selection, in contrast to models without selection, the distribution of this process, the coalescent process, depends on the distribution of the frequencies of alleles in the ancestral generations. If the ancestral frequency process can be approximated by a diffusion, then the mean and the variance of the number of segregating sites due to selectively neutral mutations in random samples can be numerically calculated. The calculations are greatly simplified if the frequencies of the alleles are tightly regulated. If the mutation rates between alleles maintained by balancing selection are low, then the number of selectively neutral segregating sites in a random sample of genes is expected to substantially exceed the number predicted under a neutral model.

Genealogy and Heraldry

The coalescent process in models with selection and recombination.

The statistical properties of the process describing the genealogical history of a random sample of genes at a selectively neutral locus which is linked to a locus at which natural selection operates are investigated. It is found that the equations describing this process are simple modifications of the equations describing the process assuming that the two loci are completely linked. Thus, the statistical properties of the genealogical process for a random sample at a neutral locus linked to a locus with selection follow from the results obtained for the selected locus. Sequence data from the alcohol dehydrogenase (Adh) region of Drosophila melanogaster are examined and compared to predictions based on the theory. It is found that the spatial distribution of nucleotide differences between Fast and Slow alleles of Adh is very similar to the spatial distribution predicted if balancing selection operates to maintain the allozyme variation at the Adh locus. The spatial distribution of nucleotide differences between different Slow alleles of Adh do not match the predictions of this simple model very well.

Alcohol Dehydrogenase

On the divergence of genes in multigene families.

Statistical properties of the amount of divergence of genes in multigene families are studied. The model considered is an infinite-site neutral model with unbiased intrachromosomal conversion, unbiased interchromosomal conversion, and recombination. By considering the time back to the most recent common ancestor of two genes, both the probability of identity and the moments of S, the number of sites that differ between two sampled genes, are obtained. We find that if recombination rates are large or conversion is always interchromosomal, then the expectation of S is 4N mu n where N is the population size, mu is the rate of mutation per generation per gene and n is the number of genes in the gene family, as the conversion rates approach zero, the moments of divergence do not approach the moments of divergence with conversion rates equal to zero, and it is possible for a decrease in the rate of intrachromosomal conversion to result in a higher probability of identity, but a greater mean divergence of the two genes.

Gene Frequency

On the divergence of members of a transposable element family.

Statistical properties of the amount of divergence of members of a transposable element family are studied. The analysis is based on the model proposed by Langley et al. describing the evolution of a family of selectively neutral transposable elements in a finite haploid population of size 2N. By considering the time back to the most recent common ancestor of two copies, both the probability of identity and the moments of the number of sites that differ between two sampled copies are obtained. Our analytic results are consistent with the numerical results of Ohta for a similar model. The effects of gene conversion are also examined. In agreement with Slatkin, we find that gene conversion has a minimal effect on the probability of identity providing that the rate of deletion is sufficiently large.

Animals

On the divergence of alleles in nested subsamples from finite populations.

Within-population variation at the DNA level will rarely be studied by sequencing of loci of randomly chosen individuals. Instead, individuals will usually be chosen for sequencing based on some knowledge of their genotype. Data collected in this way require new sampling theory. Motivated by these observations, we have examined the sampling properties of a finite population model with two mutation processes and with no selection or recombination. One mutation process generates new alleles according to an infinite-alleles model, and the other generates polymorphisms at sites according to an infinite-sites model. A sample of n genes is considered. The stationary distribution of the number of segregating sites in a subsample from one of the allelic classes in the sample conditional on the allelic configuration of the sample is studied. A recursive scheme is developed to compute the moments of this distribution, and it is shown that the distribution is functionally independent of the number of additional alleles in the sample and their respective frequencies in the sample. For the case in which the sample contains only two alleles, the distribution of the number of segregating sites in a subsample containing both alleles conditional on the sample frequencies of the alleles is studied. The results are applied to the analysis of DNA sequences of two alleles found at the Adh locus of Drosophila melanogaster. No significant departure from the neutral model is detected.

Alleles