PubMed HealthSearch

Biomedical subjects

B S Weir

Publications and source records attributed to B S Weir.

17 recordsLinked to original sources

Population genetics in the forensic DNA debate.

The use of matching variable number of tandem repeat (VNTR) profiles to link suspects with crimes is potentially very powerful, but it has been quite controversial. Initial debate over laboratory procedures has largely given way to debate over the statistical and population genetic issues involved in calculating the frequency of a profile for a random member of a population. This frequency is used to weight the evidence of a match between suspect and crime scene material when the suspect denies responsibility for that material. A recent report from the National Research Council, intended to put to rest some of the issues, has instead raised further debate by advocating a procedure based on maximum frequencies of profile components over several different populations.

Alleles

Haplotype analysis at the low density lipoprotein receptor locus: application to the study of familial hypercholesterolemia in Israel.

Familial hypercholesterolemia (FH) results from mutations in the low density lipoprotein (LDL) receptor gene. It has been shown that restriction fragment length polymorphisms (RFLPs) associated with this gene may be used for family and population studies. The present investigation is a population-based study of 19 Jewish families with hypercholesterolemia representing 9 different countries of origin. Ten RFLP sites were used to construct 24 different haplotypes from 112 chromosomes. These haplotypes vary in frequency from 0.9% to 28.6%. Five previously undescribed haplotypes, which comprise 8.1% of the sample, are reported here. The six most common haplotypes account for 70% of the sample. Segregation analysis reveals that, in Israel, distinct LDL receptor haplotypes are associated with hypercholesterolemia in 12 (63%) out of the 19 Jewish families. Five LDL receptor haplotypes co-segregate with hypercholesterolemia. Two of these haplotypes seem to be unique to specific population groups in Israel and may therefore represent founder mutations.

Apolipoproteins B

Independence of VNTR alleles defined as fixed bins.

An analysis is presented of data collected by the Federal Bureau of Investigation at six unlinked variable number of tandem repeats (VNTR) loci for the United States population. Databases have been constructed of VNTR profiles of Caucasians, Blacks and Hispanics from Florida, Texas and California. There was very little evidence for correlations between lengths for pairs of VNTR fragments, within or between loci. When the fragment lengths were amalgamated into discrete bins, there was also little evidence for disequilibrium over all genotypes, within or between loci, for the Caucasian database, although some disequilibrium was found for the Black and Hispanic databases. No disequilibrium was found for the Caucasian or Black databases when tests were confined to heterozygous individuals. In cases of global disequilibrium, local tests can be applied to specific genotypes. The results suggest that, at the bin level, frequencies of VNTR profiles can generally be estimated as the products of the frequencies of the constituent elements. This overcomes the problem of estimating population frequencies when any particular profile does not exist in the database. There is some evidence for different frequencies, at the individual bin level, between geographic samples within each of the Caucasian, Black and Hispanic databases, and considerable evidence for differences between the three databases. These differences are less evident for the frequencies of four-locus profiles.

Alleles

Testing for equality of evolutionary rates.

A likelihood ratio test is presented for comparing rates of evolutionary change in the paths of descent leading to two species. The test is compared to previous relative rate tests based on variances of estimated numbers of base substitutions. The likelihood approach allows for different transversion and transition rates, and when these rates are actually different, the likelihood ratio test can be much more powerful than the variance-based tests. For single-parameter mutation models, however, the two tests have similar power. The tests are applied to a set of chloroplast sequences from several species of grasses, and additional indications of significantly different rates leading to barley were found with the likelihood ratio test.

Base Sequence

Expected behavior of conditional linkage disequilibrium.

The ubiquitousness of RFLPs in the human genome has greatly helped the mapping of human disease genes, and it has been suggested that population measures of association between disease and marker loci could help with this mapping. For rare diseases, random samples are taken from within disease genotypes in order to obtain reasonable sample sizes, but this sampling strategy requires a modification of the usual measures of association. We present theoretical predictions for the mean and variance of such a modified measure, under the assumption that the disease gene is maintained at a constant low frequency in the population. The coefficient of variation of this modified measure is large enough that caution is needed in using the measure to locate disease genes, and, furthermore, the coefficient of variation cannot be made arbitrarily small by increasing sample size. The modified association measure is calculated for recently published data on cystic fibrosis.

Chromosome Mapping

Independence of VNTR alleles defined as floating bins.

Data bases of VNTR fragments determined for Caucasians and blacks by Cellmark Diagnostics and Lifecodes Corporation are analyzed for independence of variants within and between loci. Floating bins are constructed around specific fragment lengths and are used to define discrete genotypes. Simple chi 2 test statistics for independence of bins within and between loci are described and applied to large sets of randomly generated four-locus profiles. The proportions of significant test statistics were about as expected under the hypotheses of independence, suggesting the absence of both Hardy-Weinberg and linkage disequilibrium. In any particular forensic application, however, these tests need to be performed on the fragments in question.

Black People

Whose DNA?

Explore the source record for details and available documents.

DNA Fingerprinting

Effect of gene conversion on variances of digenic identity measures.

The variances and covariances of digenic descent measures are studied for a two-locus model incorporating mutation, gene conversion, recombination, drift, and finite sampling. Gene conversion can occur between allelic pairs of genes or between non-allelic pairs on the same or different gametes within individuals. Most interest therefore centers on pairs of genes, and five digenic identity measures are required. The behavior over time of these measures is studied, with an emphasis on the effects of gene conversion. Because of the stochastic nature of the forces of drift, recombination, mutation, and conversion, the actual identity status of gene pairs can vary from expectation among replicate populations. To study this variation we compute the expected variances and covariances of the measures, and show that this requires the introduction of trigenic and quadrigenic measures. Allowing for conversion between genes on different gametes requires a large number of these higher-order measures.

Gene Conversion

The variance of sample heterozygosity.

The variance of sample heterozygosity, averaged over several loci, is studied in a variety of situations. The variance depends on the sampling implicit in the mating system as well as on that explicit in the loci scored and individuals sampled. There are also effects of allelic distributions over loci and of linkage or linkage disequilibrium between pairs of loci. Results are obtained for populations in drift and mutation balance, for infinite populations undergoing mixed self and random mating, and for finite monoecious populations with or without selfing. For unlinked loci in drift/mutation balance, variances appear to be lessened more by increasing the number of loci scored than by increasing the number of individuals sampled. For infinite populations under the mixed self and random mating system, however, the reverse is true. Methods for estimating the variance of sample heterozygosity are discussed, with attention being paid to unbalanced data where not all loci are scored in all individuals.

Alleles

Extensive linkage disequilibrium in the achaete-scute complex of Drosophila melanogaster.

We have analyzed the level of gametic association between restriction map variants in a sample of 44 X chromosomes from a natural population of Drosophila melanogaster. Of 21 pairwise tests involving 7 restriction map polymorphisms in the yellow-achaete-scute complex, 17 were found to be significant, including some between restriction sites over 80 kb apart. Three-way linkage disequilibria and their variances were also estimated for all 35 three-way comparisons between these loci. Twelve such tests were found to be significant, again spanning distances of up to 80 kb on the restriction map. Only 9 of a possible 128 haplotypes were represented in the sample and 8 of these could be linked together by changes at a single site. The strength of these associations at y-ac-sc is unusual by comparison with studies on other regions of the genome of D. melanogaster, and is consistent with the very low level of recombination which has been reported for the complex. However, our estimate of nucleotide diversity in the region is not significantly different from those made for some other loci in this species.

Animals

Sampling strategies for distances between DNA sequences.

An international effort is now underway to obtain the DNA sequence for the entire human genome (Watson and Jordan, 1989, Genomics 5, 654-656; Barnhart, 1989, Genomics 5, 657-660). This Human Genome Initiative will generate sequence data from several species other than humans, and will result in several copies per species of at least some regions of the genome. Although the project has generated much interest, it is but one aspect of the widespread effort to generate DNA sequence data. Published sequences are collected in common databases, and release 63 of GenBank in March 1990 contained 40,127,752 bases from 33,337 reported sequences (News from GenBank 3; Mountain View, California: Intelligenetics, Inc., 1990). Large though this database is, it is only about 1% of the number of bases in the human genome. Interpretations of data of such magnitude are going to require the collaborative efforts of biometricians and molecular biologists, and an aim of this paper is to show that there is also a role for readers of this journal in the design of surveys of DNA sequences. Discussion here will center on the use of sequence data in evolutionary studies, where some region of DNA is sequenced in several different species. The object is to infer the evolutionary history of that particular region, or of the species themselves. Statistical issues in the very important studies on sequences to locate and characterize regions responsible for human diseases will not be addressed here. We will discuss appropriate ways of measuring distances between DNA sequences and of predicting the sampling properties of the distances. There are procedures for inferring evolutionary histories for a set of elements that depend on a matrix of distances between each pair of elements, and the precision of resulting trees must be influenced by the precision of the distances. We will show that account needs to be taken of two sampling processes--the sampling of sequences by the investigator ("statistical sampling"), and the sampling of genetic material involved in the formation of offspring from a parental population ("genetic sampling").

Analysis of Variance

Inferences about linkage disequilibrium.

Existing theory for inferences about linkage disequilibrium is restricted to a measure defined on gametic frequencies. Unless gametic frequencies are directly observable, they are inferred from genotypic frequencies under the assumption of random union of gametes. Primary emphasis in this paper is given to genotypic data, and disequilibrium coefficients are defined for all subsets of two or more of the four genes, two at each of two loci, carried by an individual. Linkage disequilibrium coefficients are defined for genes within and between gametes, and methods of estimating and testing these coefficients are given for gametic data. For genotypic data, when coupling and repulsion double heterozygotes cannot be distinguished. Burrows' composite measure of linkage disequilibrium is discussed. In particular, the estimate for this measure and hypothesis tests based on it are compared to the usual maximum likelihood estimate of gametic linkage disequilibrium, and corresponding likelihood ratio or contingency chi-square tests. General use of the composite measure, whether or not random union of gametes is an appropriate assumption, is recommended. Attention is given to small samples, where the non-normality of gene frequencies will have greatest effect on methods of inference based on normal theory. Even tools such as Fisher's z-transformation for the correlation of gene frequencies are found to perform quite satisfactorily.

Gene Frequency

Quadratic analyses of reciprocal crosses.

Three different models, a two-way factorial model for familiarity, an orthogonalizing transform of this model to a diallel model, and a bio model more representative of the biological situation, are interrelated in terms of their components of variance and covariance. It is clarified that there are five components that can be reckoned with in the analysis of reciprocal crosses, including distinct maternal and paternal variances. Estimation of the components and tests of hypotheses concerning them are outlined for two types of mating designs with reciprocals. One deisgn involves a factorial mating design between two distinct sets of parents or parental lines and the other a diallel of all crosses from a single set of parents or parental lines. Both designs provide the same types of information and similar tests of hypotheses. At least some parts of the analyses corresponding to the factorial model are required to separate the maternal and paternal variances. A least squares partitioning of the sums of squares according to the diallel model, but with expectations expressed in terms of the bio model, provides most of the tests of hypotheses of interest. Worked examples are given.

Crosses, Genetic

Testing for selective neutrality of electrophoretically detectable protein polymorphisms.

The statistical assessment of gene-frequency data on protein polymorphisms in natural populations remains a contentious issue. Here we formulate a test of whether polymorphisms detected by electrophoresis are in accordance with the stepwise, or charge-state, model of mutation in finite populations in the absence of selection. First, estimates of the model parameters are derived by minimizing chi-square deviations of the observed frequencies of genotypes with alleles (0,1,2...) units apart from their theoretical expected values. Then the remaining deviation is tested under the null hypothesis of neutrality. The procedure was found to be conservative for false rejections in simulation data. We applied the test to Ayala and Tracey 's data on 27 allozymic loci in six populations of Drosophila willistoni . About one-quarter of polymorphic loci showed significant departure from the neutral theory predictions in virtually all populations. A further quarter showed significant departure in some populations. The remaining data showed an acceptable fit to the charge state model. A predominating mode of selection was selection against alleles associated with extreme electrophoretic mobilities. The advantageous properties and the difficulties of the procedure are discussed.

Animals

Population differentiation under the charge state model.

The extent of divergence between partially isolated sub-populations for electrophoretically detectable alleles was formulated assuming the island model of migration and the charge state model of mutation. At equilibrium the ratio of the variance of charge between the means of k different islands to the average within-island variance of charge was shown to be approximately 4Nemk2/(k-1)2 where Ne is the effective size of each island population and m is the migration rate. This ratio was calculated from published data for eight polymorphic loci in six island populations of Drosophila willistoni. Under the assumption that all variants are selectively neutral, migration rates of greater than 10 adults per generation per island are required to explain the observed similarity of the allelic profiles in D. willistoni. Since the islands studied appear to be virtually completely isolated it was concluded either that the observed protein variants are adaptive and maintained in populations by some form of balancing selection or that the observed variants themselves are neutral but natural selection acts to restrict the appearance of more extreme variants in the charge carried.

Drosophila