PubMed HealthSearch

Biomedical subjects

L Excoffier

Publications and source records attributed to L Excoffier.

At least 19 recordsLinked to original sources

Human genetic affinities for Y-chromosome P49a,f/TaqI haplotypes show strong correspondence with linguistics.

Numerous population samples from around the world have been tested for Y chromosome-specific p49a,f/TaqI restriction polymorphisms. Here we review the literature as well as unpublished data on Y-chromosome p49a,f/TaqI haplotypes and provide a new nomenclature unifying the notations used by different laboratories. We use this large data set to study worldwide genetic variability of human populations for this paternally transmitted chromosome segment. We observe, for the Y chromosome, an important level of population genetics structure among human populations (FST = .230, P < .001), mainly due to genetic differences among distinct linguistic groups of populations (FCT = .246, P < .001). A multivariate analysis based on genetic distances between populations shows that human population structure inferred from the Y chromosome corresponds broadly to language families (r = .567, P < .001), in agreement with autosomal and mitochondrial data. Times of divergence of linguistic families, estimated from their internal level of genetic differentiation, are fairly concordant with current archaeological and linguistic hypotheses. Variability of the p49a,f/TaqI polymorphic marker is also significantly correlated with the geographic location of the populations (r = .613, P < .001), reflecting the fact that distinct linguistic groups generally also occupy distinct geographic areas. Comparison of Y-chromosome and mtDNA RFLPs in a restricted set of populations shows a globally high level of congruence, but it also allows identification of unequal maternal and paternal contributions to the gene pool of several populations.

DNA Probes

Testing for linkage disequilibrium in genotypic data using the Expectation-Maximization algorithm.

We generalize an approach suggested by Hill (Heredity, 33, 229-239, 1974) for testing for significant association among alleles at two loci when only genotype and not haplotype frequencies are available. The principle is to use the Expectation-Maximization (EM) algorithm to resolve double heterozygotes into haplotypes and then apply a likelihood ratio test in order to determine whether the resolutions of haplotypes are significantly nonrandom, which is equivalent to testing whether there is statistically significant linkage disequilibrium between loci. The EM algorithm in this case relies on the assumption that genotype frequencies at each locus are in Hardy-Weinberg proportions. This method can accommodate X-linked loci and samples from haplodiploid species. We use three methods for testing significance of the likelihood ratio: the empirical distribution in a large number of randomized data sets, the X2 approximation for the distribution of likelihood ratios, and the Z2 test. The performance of each method is evaluated by applying it to simulated data sets and comparing the tail probability with the tail probability from Fisher's exact test applied to the actual haplotype data. For realistic sample sizes (50-150 individuals) all three methods perform well with two or three alleles per locus, but only the empirical distribution is adequate when there are five to eight alleles per locus, as is typical of hypervariable loci such as microsatellites. The method is applied to a data set of 32 microsatellite loci in a Finnish population and the results confirm the theoretical predictions. We conclude that with highly polymorphic loci, the EM algorithm does lead to a useful test for linkage disequilibrium, but that it is necessary to find the empirical distribution of likelihood ratios in order to perform a test of significance correctly.

Algorithms

A generic estimation of population subdivision using distances between alleles with special reference for microsatellite loci.

Several estimators of population differentiation have been proposed in the recent past to deal with various types of genetic markers (i.e., allozymes, nucleotide sequences, restriction fragment length polymorphisms, or microsatellites). We discuss the relationships among these estimators and show how a single analysis of variance framework can accomodate these qualitatively different data types.

Alleles

The impact of population expansion and mutation rate heterogeneity on DNA sequence polymorphism.

In order to study the effect of mutation rate heterogeneity on patterns of DNA polymorphism, we simulated samples of DNA sequences with gamma-distributed nucleotide substitution rates in stationary and expanding populations. We find that recent population expansions and mutation rate heterogeneity have similar effects on several polymorphism indicators, like the shape and the mean of the observed pairwise difference distribution, or the number of segregating sites. The inferred size of population expansion thus appears overestimated if nucleotides have dissimilar substitution rates. Interestingly, population expansion and uneven mutation rates have contrasting effects on Tajima's D statistic when acting separately, and the consequence on the associated test of selective neutrality is investigated. The patterns of polymorphism of several human populations analyzed for the mitochondrial control region are examined, mainly showing the difficulty in quantifying the respective contribution of past demographic history and uneven mutation rates from a single sampled evolutionary process. However, substitution rates appear more heterogeneous in the second hypervariable segment of the control region than in the first segment.

Animals

Mitochondrial DNA sequence variation across linguistic and geographic boundaries in Italy.

A previous investigation demonstrated the existence of extensive allele frequency diversity within an area of northern Italy crossed by a linguistic (dialect) boundary and by the Po River, either of them or both presumably constraining gene flow. We obtained hair samples from 45 school pupils from 9 localities in that area and sequenced a 255-bp segment of the mtDNA D loop. Estimates of the minimum number of migration events from gene genealogies suggest that the linguistic barrier impaired gene flow more than the river did. However, an analysis of molecular variance (AMOVA) showed that most sequence diversity occurs within rather than between populations and that the differences between groups of populations, defined either by linguistic or geographic criteria, do not reach significance. Three areas of rapid genetic variation were identified; their locations suggest that populations of the western part of the study area evolved in relative isolation. Therefore mtDNA sequence variation does not seem to reflect the same processes--drift and presence of dispersal barriers--that led to the observed distributions of nuclear allele frequencies.

Adolescent

Evolutionary correlation between control region sequence and restriction polymorphisms in the mitochondrial genome of a large Senegalese Mandenka sample.

We present here the first comparative analysis at the population level between Restriction Fragment Length Polymorphism (RFLP) and control region sequence polymorphism in a large and homogeneous Senegalese Mandenka sample. Eleven RFLP haplotypes and 60 different sequences are found in 119 individuals, revealing that a very high level of mtDNA diversity can be maintained in a small population. A sequence neighbor-joining tree and an analysis of molecular variance show that sequences associated with a given restriction haplotype are evolutionarily highly correlated: sequencing generally leads to the subtyping of RFLP haplotypes. Evolutionary relationships among RFLP haplotypes inferred from restriction site differences are in good agreement with those inferred from sequence data. A single difference is observed and is likely due to a single restriction homoplasy having occurred in the control region. Selective neutrality tests on both RFLP and sequence data accept the hypotheses of mtDNA neutrality and population equilibrium. The deep coalescence times (exceeding 50,000 yr) of sequences associated with the two most frequent restriction haplotypes confirm that the Niokolo Mandenka population has not passed through a recent bottleneck and that gene flow is maintained among West African populations despite ethnic differences.

Base Sequence

Maximum-likelihood estimation of molecular haplotype frequencies in a diploid population.

Molecular techniques allow the survey of a large number of linked polymorphic loci in random samples from diploid populations. However, the gametic phase of haplotypes is usually unknown when diploid individuals are heterozygous at more than one locus. To overcome this difficulty, we implement an expectation-maximization (EM) algorithm leading to maximum-likelihood estimates of molecular haplotype frequencies under the assumption of Hardy-Weinberg proportions. The performance of the algorithm is evaluated for simulated data representing both DNA sequences and highly polymorphic loci with different levels of recombination. As expected, the EM algorithm is found to perform best for large samples, regardless of recombination rates among loci. To ensure finding the global maximum likelihood estimate, the EM algorithm should be started from several initial conditions. The present approach appears to be useful for the analysis of nuclear DNA sequences or highly variable loci. Although the algorithm, in principle, can accommodate an arbitrary number of loci, there are practical limitations because the computing time grows exponentially with the number of polymorphic loci. Although the algorithm, in principle, can accommodate an arbitrary number of loci, there are practical limitations because the computing time grows exponentially with the number of polymorphic loci.

Algorithms

Nuclear DNA polymorphism in a Mandenka population from Senegal: comparison with eight other human populations.

A large and ethnically well defined Mandenka sample from Senegal is analysed for 80 nuclear DNA RFLPs, and compared with eight previously studied human populations. A high level of genetic diversity is found in this sample, comparable to that observed in two African Pygmy samples, but lower than that of a European sample. High population variation is observed for most markers. A neutrality test reveals that the markers used in this study can be considered as neutral. A high correlation is found between genetic and geographic distances (r = 0.62), suggesting that geography does also affect long range population genetic relationships and is an important factor behind differentiation among human populations.

Alleles

High diversity of alpha-globin haplotypes in a Senegalese population, including many previously unreported variants.

RFLP haplotypes at the alpha-globin gene complex have been examined in 190 individuals from the Niokolo Mandenka population of Senegal: haplotypes were assigned unambiguously for 210 chromosomes. The Mandenka share with other African populations a sample size-independent haplotype diversity that is much greater than that in any non-African population: the number of haplotypes observed in the Mandenka is typically twice that seen in the non-African populations sampled to date. Of these haplotypes, 17.3% had not been observed in any previous surveys, and a further 19.1% have previously been reported only in African populations. The haplotype distribution shows clear differences between African and non-African peoples, but this is on the basis of population-specific haplotypes combined with haplotypes common to all. The relationship of the newly reported haplotypes to those previously recorded suggests that several mutation processes, particularly recombination as homologous exchange or gene conversion, have been involved in their production. A computer program based on the expectation-maximization (EM) algorithm was used to obtain maximum-likelihood estimates of haplotype frequencies for the entire data set: good concordance between the unambiguous and EM-derived sets was seen for the overall haplotype frequencies. Some of the low-frequency haplotypes reported by the estimation algorithm differ greatly, in structure, from those haplotypes known to be present in human populations, and they may not represent haplotypes actually present in the sample.

Genetic Variation

Analysis of mtDNA variation in African populations reveals the most ancient of all human continent-specific haplogroups.

mtDNA sequence variation was examined in 140 Africans, including Pygmies from Zaire and Central African Republic (C.A.R.) and Mandenkalu, Wolof, and Pular from Senegal. More than 76% of the African mtDNAs (100% of the Pygmies and 67.3% of the Senegalese) formed one major mtDNA cluster (haplogroup L) defined by an African-specific HpaI site gain at nucleotide pair (np) 3592. Additional mutations subdivided haplogroup L into two subhaplogroups, each encompassing both Pygmy and Senegalese mtDNAs. A novel 12-bp homoplasmic insertion in the intergenic region between tRNA(Tyr) and cytochrome oxidase I (COI) genes was also observed in 17.6% of the Pygmies from C.A.R. This insertion is one of the largest observed in human mtDNAs. Another 25% of the Pygmy mtDNAs harbored a 9-bp deletion between the cytochrome oxidase II (COII) and tRNA(Lys) genes, a length polymorphism previously reported in non-African populations. In addition to haplogroup L, other haplogroups were observed in the Senegalese. These haplogroups were more similar to those observed in Europeans and Asians than to haplogroup L mtDNAs, suggesting that the African mtDNAs without the HpaI np 3592 site could be the ancestral types from which European and Asian mtDNAs were derived. Comparison of the intrapopulation sequence divergence in African and non-African populations confirms that African populations exhibit the largest extent of mtDNA variation, a result that further supports the hypothesis that Africans represent the most ancient human group and that all modern humans have a common and recent African origin. The age of the total African variation was estimated to be 101,000-133,000 years before present (YBP), while the age of haplogroup L was estimated at 98,000-130,000 YBP. These values substantially exceed the ages of all Asian- and European-specific mtDNA haplogroups.

Africa

Using allele frequencies and geographic subdivision to reconstruct gene trees within a species: molecular variance parsimony.

We formalize the use of allele frequency and geographic information for the construction of gene trees at the intraspecific level and extend the concept of evolutionary parsimony to molecular variance parsimony. The central principle is to consider a particular gene tree as a variable to be optimized in the estimation of a given population statistic. We propose three population statistics that are related to variance components and that are explicit functions of phylogenetic information. The methodology is applied in the context of minimum spanning trees (MSTs) and human mitochondrial DNA restriction data, but could be extended to accommodate other tree-making procedures, as well as other data types. We pursue optimal trees by heuristic optimization over a search space of more than 1.29 billion MSTs. This very large number of equally parsimonious trees underlines the lack of resolution of conventional parsimony procedures. This lack of resolution is highlighted by the observation that equally parsimonious trees yield very different estimates of population genetic diversity and genetic structure, as shown by null distributions of the population statistics, obtained by evaluation of 10,000 random MSTs. We propose a non-parametric test for the similarity between any two trees, based on the distribution of a weighted coevolutionary correlation. The ability to test for tree relatedness leads to the definition of a class of solutions instead of a single solution. Members of the class share virtually all of the critical internal structure of the tree but differ in the placement of singleton branch tips.

Alleles

HLA-DPB1 DNA polymorphism in the Swiss population: linkage disequilibrium with other HLA loci and population genetic affinities.

Allelic diversity at the HLA-DPB1 locus was determined by PCR-oligotyping in a sample of 125 healthy Swiss individuals. A total of 17 alleles were detected among which four main alleles (DPB1*0401, *0201, *0301, *0402) reached a cumulative frequency of 74.8%. HLA-A and -B (by serology) and HLA-DRB1 (by oligotyping) allelic polymorphisms were analysed also. HLA-B and HLA-DRB1 loci were highly polymorphic with 25 and 28 alleles respectively and similar heterozygosity levels of 0.93 and 0.92. These two loci were found to be more polymorphic than expected under neutrality, while lower heterozygosity levels were found for HLA-A (0.87) and DPB1 (0.81) loci. This paper presents also a global comparison of DPB1 allelic frequencies among 15 populations from four continents. As opposed to the DRB1 locus, overall DPB1 is shown to have a lower level of polymorphism and may be considered as neutral in all tested populations. DPB1 genetic diversity is correlated significantly with geography also, as found previously for DRB1. Two- and four-locus haplotype frequencies were determined and the significance of their linkage disequilibrium tested by an original non-parametric method. A significant positive linkage disequilibrium was found for 11 A-B, 16 B-DRB1, 7 DRB1-DPB1 and 3 A-B-DRB1-DPB1 haplotypes. The overall linkage disequilibrium between DRB1 and DPB1 was much lower than expected from the physical distance and lower than for A-B and B-DRB1 pairs. The implications of these results for bone marrow transplantation and for the evolution of HLA loci are discussed.

Alleles

New data for AG haplotype frequencies in Caucasoid populations and selective neutrality of the AG polymorphism.

We present the results of AG antigen typings of three Caucasoid population samples: Lebanese, Tunisians, and Finns. AG haplotype frequencies estimated by maximum-likelihood methods are compared with the frequencies observed in 13 world populations previously tested for AG specificities by computing a genetic distance matrix used in a multivariate analysis. A high degree of polymorphism characterizes the three samples, with 10 haplotypes detected in the Lebanese and 11 haplotypes detected in the Tunisians and Finns; high heterozygosity levels are also present in the three populations. The genetic distance analysis shows that the three populations possess a genetic structure intermediate between those observed in sub-Saharan Africans and in Caucasoids from the Near East and India. This tight correspondence between AG differentiation and geography is confirmed by a highly significant correlation coefficient found between genetic and geographic distances computed worldwide, suggesting that an isolation by distance model of evolution applies to the AG system. The Ewens-Watterson test for selective neutrality on all world populations tested for AG specificities also supports the hypothesis that the AG system behaves like a neutral polymorphism. Overall, the AG differentiation pattern appears to be close to the patterns observed for other serological polymorphisms, such as RH, GM, and HLA, whose evolutionary mechanisms are also discussed.

Antigens, Differentiation

[Genetic studies of relationship between Mediterranean and Atlantic populations of loggerhead turtle Caretta caretta with mitochondrial marker].

The loggerhead turtle Caretta caretta is an endangered species in the Mediterranean. Therefore, the definition of the Mediterranean population, and their relationships to the Atlantic population is of fundamental importance. For this purpose, we have sequenced a portion of the mitochondrial cytochrome b gene to generate genetic markers. Results indicate that the Mediterranean nesting female population is genetically isolated from the Atlantic nesting female population, but loggerhead turtles of Atlantic origin were found in the West Mediterranean basin. This entry of Atlantic loggerheads in the Mediterranean confirms earlier speculations and presents special conservation problems. The Spanish swordfish longline fishery which incidentally captures large numbers of loggerheads in the West Mediterranean basin has therefore an impact on the Atlantic population. These data demonstrate the international nature of marine turtle conservation.

Animals

Analysis of molecular variance inferred from metric distances among DNA haplotypes: application to human mitochondrial DNA restriction data.

We present here a framework for the study of molecular variation within a single species. Information on DNA haplotype divergence is incorporated into an analysis of variance format, derived from a matrix of squared-distances among all pairs of haplotypes. This analysis of molecular variance (AMOVA) produces estimates of variance components and F-statistic analogs, designated here as phi-statistics, reflecting the correlation of haplotypic diversity at different levels of hierarchical subdivision. The method is flexible enough to accommodate several alternative input matrices, corresponding to different types of molecular data, as well as different types of evolutionary assumptions, without modifying the basic structure of the analysis. The significance of the variance components and phi-statistics is tested using a permutational approach, eliminating the normality assumption that is conventional for analysis of variance but inappropriate for molecular data. Application of AMOVA to human mitochondrial DNA haplotype data shows that population subdivisions are better resolved when some measure of molecular differences among haplotypes is introduced into the analysis. At the intraspecific level, however, the additional information provided by knowing the exact phylogenetic relations among haplotypes or by a nonlinear translation of restriction-site change into nucleotide diversity does not significantly modify the inferred population genetic structure. Monte Carlo studies show that site sampling does not fundamentally affect the significance of the molecular variance components. The AMOVA treatment is easily extended in several different directions and it constitutes a coherent and flexible framework for the statistical analysis of molecular data.

Analysis of Variance

[Polymorphism of HLA class I loci HLA-A, -B, -C, in the Mandenka population from eastern Senegal].

A sample of 162 Mandenkalu from Eastern Senegal has been typed for three HLA class I loci: HLA-A, -B and -C. The Mandenka population presents a very high genetic variability with 15 alleles for locus A, 24 alleles for locus B, and at least 8 alleles for locus C. The calculated heterozygosities for the three loci A, B, and C are respectively 0.884, 0.944 and 0.829. The Mandenkalu allelic frequencies are close to that found in other sub-Saharan populations. They show, however, some peculiarities like the occurrence of the Bw 56 allele and the high frequencies of both B5 and B35.

Alleles

HLA-DR polymorphism in a Senegalese Mandenka population: DNA oligotyping and population genetics of DRB1 specificities.

HLA class II loci are useful markers in human population genetics, because they are extremely variable and because new molecular techniques allow large-scale analysis of DNA allele frequencies. Direct DNA typing by hybridization with sequence-specific oligonucleotide probes (HLA oligotyping) after enzymatic in vitro PCR amplification detects HLA allelic polymorphisms for all class II loci. A detailed HLA-DR oligotyping analysis of 191 individuals from a geographically, culturally, and genetically well-defined western African population, the Mandenkalu, reveals a high degree of polymorphism, with at least 24 alleles and a heterozygosity level of .884 for the DRB1 locus. The allele DRB1*1304, defined by DNA sequencing of the DRB1 first-domain exon, is the most frequent allele (27.1%). It accounts for an unusually high DR13 frequency, which is nevertheless within the neutral frequency range. The next most frequent specificities are DR11, DR3, and DR8. Among DRB3-encoded alleles, DR52b (DRB3*02) represents as much as 80.7% of all DR52 haplotypes. A survey of HLA-DR specificities in populations from different continents shows a significant positive correlation between genetic and geographic differentiation patterns. A homozygosity test for selective neutrality of DR specificities is not significant for the Mandenka population but is rejected for 20 of 24 populations. Observed high heterozygosity levels in tested populations are compatible with an overdominant model with a small selective advantage for heterozygotes.

Base Sequence

Spatial differentiation of RH and GM haplotype frequencies in Sub-Saharan Africa and its relation to linguistic affinities.

This study analyzes patterns of variation in eight GM and seven Rhesus (RH) haplotypes across sub-Saharan Africa. We examine the concordance with genetic patterns of both geographic and language-family relationships by spatial analysis and ordination techniques. The genetic variation has significant spatial structure, but positive autocorrelation declines neither asymptotically nor proportionally with increasing distance. Evidently, neither isolation by distance with increasing distance. Evidently, neither isolation by distance nor clinical migration-selection models account for the observed genetic structure. Language-family relationship is the best predictor of genetic relationship and may reflect historic migrations and expansions of ethnically different peoples within sub-Saharan Africa. Yet the greatest part of the genetic variance remains unexplained by the models we have tested.

Africa, Central