PubMed Health⌕ Search

Biomedical subjects

Martin Kreitman

Publications and source records attributed to Martin Kreitman.

17 recordsLinked to original sources

A genome-wide survey of R gene polymorphisms in Arabidopsis.

We used polymorphism analysis to study the evolutionary dynamics of 27 disease resistance (R) genes by resequencing the leucine-rich repeat (LRR) region in 96 Arabidopsis thaliana accessions. We compared single nucleotide polymorphisms (SNPs) in these R genes to an empirical distribution of SNP in the same sample based on 876 fragments selected to sample the entire genome. LRR regions are highly polymorphic for protein variants but not for synonymous changes, suggesting that they generate many alleles maintained for short time periods. Recombination is also relatively common and important for generating protein variants. Although none of the genes is nearly as polymorphic as RPP13, a locus previously shown to have strong signatures of balancing selection, seven genes show weaker indications of balancing selection. Five R genes are relatively invariant, indicating young alleles, but all contain segregating protein variants. Polymorphism analysis in neighboring fragments yielded inconclusive evidence for recent selective sweeps at these loci. In addition, few alleles are candidates for rapid increases in frequency expected under directional selection. Haplotype sharing analysis revealed significant underrepresentation of R gene alleles with extended haplotypes compared with 1102 random genomic fragments. Lack of convincing evidence for directional selection or selective sweeps argues against an arms race driving R gene evolution. Instead, the data support transient or frequency-dependent selection maintaining protein variants at a locus for variable time periods.

Arabidopsis↗

Presence/absence polymorphism for alternative pathogenicity islands in Pseudomonas viridiflava, a pathogen of Arabidopsis.

The contribution of arms race dynamics to plant-pathogen coevolution has been called into question by the presence of balanced polymorphisms in resistance genes of Arabidopsis thaliana, but less is known about the pathogen side of the interaction. Here we investigate structural polymorphism in pathogenicity islands (PAIs) in Pseudomonas viridiflava, a prevalent bacterial pathogen of A. thaliana. PAIs encode the type III secretion system along with its effectors and are essential for pathogen recognition in plants. P. viridiflava harbors two structurally distinct and highly diverged PAI paralogs (T- and S-PAI) that are integrated in different chromosome locations in the P. viridiflava genome. Both PAIs are segregating as presence/absence polymorphisms such that only one PAI ([T-PAI, nablaS-PAI] and [nablaT-PAI, S-PAI]) is present in any individual cell. A worldwide population survey identified no isolate with neither or both PAI. T-PAI and S-PAI genotypes exhibit virulence differences and a host-specificity tradeoff. Orthologs of each PAI can be found in conserved syntenic locations in other Pseudomonas species, indicating vertical phylogenetic transmission in this genus. Molecular evolutionary analysis of PAI sequences also argues against "recent" horizontal transfer. Spikes in nucleotide divergence in flanking regions of PAI and nabla-PAI alleles suggest that the dual PAI polymorphism has been maintained in this species under some form of balancing selection. Virulence differences and host specificities are hypothesized to be responsible for the maintenance of the dual PAI system in this bacterial pathogen.

Arabidopsis↗

The pattern of polymorphism in Arabidopsis thaliana.

We resequenced 876 short fragments in a sample of 96 individuals of Arabidopsis thaliana that included stock center accessions as well as a hierarchical sample from natural populations. Although A. thaliana is a selfing weed, the pattern of polymorphism in general agrees with what is expected for a widely distributed, sexually reproducing species. Linkage disequilibrium decays rapidly, within 50 kb. Variation is shared worldwide, although population structure and isolation by distance are evident. The data fail to fit standard neutral models in several ways. There is a genome-wide excess of rare alleles, at least partially due to selection. There is too much variation between genomic regions in the level of polymorphism. The local level of polymorphism is negatively correlated with gene density and positively correlated with segmental duplications. Because the data do not fit theoretical null distributions, attempts to infer natural selection from polymorphism data will require genome-wide surveys of polymorphism in order to identify anomalous regions. Despite this, our data support the utility of A. thaliana as a model for evolutionary functional genomics.

Arabidopsis↗

Functional evolution of a cis-regulatory module.

Lack of knowledge about how regulatory regions evolve in relation to their structure-function may limit the utility of comparative sequence analysis in deciphering cis-regulatory sequences. To address this we applied reverse genetics to carry out a functional genetic complementation analysis of a eukaryotic cis-regulatory module-the even-skipped stripe 2 enhancer-from four Drosophila species. The evolution of this enhancer is non-clock-like, with important functional differences between closely related species and functional convergence between distantly related species. Functional divergence is attributable to differences in activation levels rather than spatiotemporal control of gene expression. Our findings have implications for understanding enhancer structure-function, mechanisms of speciation and computational identification of regulatory modules.

Animals↗

Genetic diversity, recombination and cryptic clades in Pseudomonas viridiflava infecting natural populations of Arabidopsis thaliana.

Species-level genetic diversity and recombination in bacterial pathogens of wild plant populations have been nearly unexplored. Pseudomonas viridiflava is a common natural bacterial pathogen of Arabidopsis thaliana, for which pathogen defense genes and mechanisms are becoming increasing well known. The genetic variation contained within a worldwide sample of P. viridiflava collected from wild populations of A. thaliana was investigated using five genomic sequence fragments totaling 2.3 kb. Two distinct and deeply diverged clades were found within the P. viridiflava sample and in close proximity in multiple populations, each genetically diverse with synonymous variation as high as 9.3% in one of these clades. Within clades, there is evidence of frequent recombination within and between each sequenced locus and little geographic differentiation. Isolates from both clades were also found in a small sample of other herbaceous species in Midwest populations, indicating a possibly broad host range for P. viridiflava. The high levels of genetic variation and recombination together with a lack of geographic differentiation in this pathogen distinguish it from other bacterial plant pathogens for which intraspecific variation has been examined.

Arabidopsis↗

Intragenic spatial patterns of codon usage bias in prokaryotic and eukaryotic genomes.

To study the roles of translational accuracy, translational efficiency, and the Hill-Robertson effect in codon usage bias, we studied the intragenic spatial distribution of synonymous codon usage bias in four prokaryotic (Escherichia coli, Bacillus subtilis, Sulfolobus tokodaii, and Thermotoga maritima) and two eukaryotic (Saccharomyces cerevisiae and Drosophila melanogaster) genomes. We generated supersequences at each codon position across genes in a genome and computed the overall bias at each codon position. By quantitatively evaluating the trend of spatial patterns using isotonic regression, we show that in yeast and prokaryotic genomes, codon usage bias increases along translational direction, which is consistent with purifying selection against nonsense errors. Fruit fly genes show a nearly symmetric M-shaped spatial pattern of codon usage bias, with less bias in the middle and both ends. The low codon usage bias in the middle region is best explained by interference (the Hill-Robertson effect) between selections at different codon positions. In both yeast and fruit fly, spatial patterns of codon usage bias are characteristically different from patterns of GC-content variations. Effect of expression level on the strength of codon usage bias is more conspicuous than its effect on the shape of the spatial distribution.

Animals↗

Natural selection for polymorphism in the disease resistance gene Rps2 of Arabidopsis thaliana.

Pathogen resistance is an ecologically important phenotype increasingly well understood at the molecular genetic level. In this article, we examine levels of avrRpt2-dependent resistance and Rps2 locus DNA sequence variability in a worldwide sample of 27 accessions of Arabidopsis thaliana. The rooted parsimony tree of Rps2 sequences drawn from a diverse set of ecotypes includes a deep bifurcation separating major resistance and susceptibility clades of alleles. We find evidence for selection maintaining these alleles and identify the N-terminal part of the leucine-rich repeat region as a probable target of selection. Additional protein variants are found within the two major clades and correlate well with measurable differences among ecotypes in resistance to the avirulence gene avrRpt2 of the pathogen Pseudomonas syringae. Long-lived polymorphisms have been observed for other resistance genes of A. thaliana; the Rps2 data suggest that the long-term maintenance of phenotypic variation in resistance genes may be a general phenomenon and are consistent with diversifying selection acting in concert with selection to maintain variation.

Arabidopsis↗

A method for detecting recent selection in the human genome from allele age estimates.

Mutations that have recently increased in frequency by positive natural selection are an important component of naturally occurring variation that affects fitness. To identify such variants, we developed a method to test for recent selection by estimating the age of an allele from the extent of haplotype sharing at linked sites. Neutral coalescent simulations are then used to determine the likelihood of this age given the allele's observed frequency. We applied this method to a common disease allele, the hemochromatosis-associated HFE C282Y mutation. Our results allow us to reject neutral models incorporating plausible human demographic histories for HFE C282Y and one other young but common allele, indicating positive selection at HFE or a linked locus. This method will be useful for scanning the human genome for alleles under selection using the haplotype map now being constructed.

Alleles↗

On the power to detect SNP/phenotype association in candidate quantitative trait loci genomic regions: a simulation study.

We use coalescent methods to investigate the ability of linked neutral "markers" to reveal in simulated population samples the presence of one or more single nucleotide polymorphisms that is contributing to a trait having a complex genetic basis (QTN: quantitative trait nucleotide). Realistic mutation and recombination rates in our simulations allow us to generate SNP data appropriate for analyzing human variation across short chromosomal intervals corresponding to approximately 100 kilobases. We investigate the performance of both single marker and multiple-marker (haplotype) data for several ad hoc procedures. Our results with single SNP markers indicate that (1) the density of SNP markers need not be much higher than 10% in order to achieve near-maximal detection of a QTN; (2) a higher density of markers does not improve much on the ability to localize a QTN within an interval unless the recombination rate is high. Haplotype-based tests were investigated for the case in which more than one QTN is present in the studied interval. Larger sample sizes improve both the probability of detecting the haplotype with the largest number of QTNs, as well as the ability to infer correct haplotypes from genotypic data. Testing a series of short haplotypes across a longer interval can also be beneficial. The rate of false positives (i.e., when the most significant haplotype does not contain the greatest number of QTNs in the sample) can be very high when the contribution of individual QTNs to a trait is small. The elimination of low-frequency haplotypes does not substantially reduce the probability of detecting the haplotype with the largest number of QTNs but it can reduce the rate of false positives.

Computational Biology↗

Signature of balancing selection in Arabidopsis.

Natural selection and genetic linkage cause DNA segments to have genealogical histories resembling those of the selected sites. When a polymorphism maintained by selection is old, it will have an island of enhanced sequence variability surrounding it, which represents a detectable "signature of selection." We investigate the structure of single-nucleotide polymorphisms (SNPs) in a 20-kb interval containing the Arabidopsis thaliana disease resistance gene RPS5, a locus containing common alleles for the presence/absence of the entire locus. The alleles are considerably diverged at surrounding sites, indicative of an old polymorphism maintained by selection. The island of "enhanced" variability extends several kilobases to either side of the RPS5 deletion junction, and these SNPs are in nearly complete linkage disequilibrium with the RPS5 insertion/deletion. At a distance of 10 kb to either side of the locus, however, we find low levels of polymorphism and the absence of linkage disequilibrium between individual SNPs and RPS5 alleles. Our results show that the interval of enhanced variability surrounding this balanced polymorphism in Arabidopsis is large enough to be readily detected, but small enough to span the focal gene and few others. For this species it should be possible to identify the complete set of genes with long-lived polymorphisms, a potentially important subset of genes segregating for functional variants.

Arabidopsis↗

The extent of linkage disequilibrium in Arabidopsis thaliana.

Linkage disequilibrium (LD), the nonrandom occurrence of alleles in haplotypes, has long been of interest to population geneticists. Recently, the rapidly increasing availability of genomic polymorphism data has fueled interest in LD as a tool for fine-scale mapping, in particular for human disease loci. The chromosomal extent of LD is crucial in this context, because it determines how dense a map must be for associations to be detected and, conversely, limits how finely loci may be mapped. Arabidopsis thaliana is expected to harbor unusually extensive LD because of its high degree of selfing. Several polymorphism studies have found very strong LD within individual loci, but also evidence of some recombination. Here we investigate the pattern of LD on a genomic scale and show that in global samples, LD decays within approximately 1 cM, or 250 kb. We also show that LD in local populations may be much stronger than that of global populations, presumably as a result of founder events. The combination of a relatively high level of polymorphism and extensive haplotype structure bodes well for developing a genome-wide LD map in A. thaliana.

Arabidopsis↗

Zuckerkandl Prize.

Explore the source record for details and available documents.

Awards and Prizes↗

Patterns of genetic variation at a chromosome 4 locus of Drosophila melanogaster and D. simulans.

DNA sequence surveys of Drosophila melanogaster populations show a strong positive correlation between the recombination rate experienced by a locus and its level of nucleotide polymorphism. In particular, surveys of the fourth chromosome gene ci(D) show greatly reduced levels of nucleotide variation; this observation was originally interpreted in terms of selective sweeps occurring on the nonrecombining fourth chromosome. Subsequent theoretical work has, however, uncovered several other selective processes that can reduce variation. In this study, we revisit the Drosophila fourth chromosome, investigating variation in 5-6 kb of the gene ankyrin in D. melanogaster and D. simulans. Silent nucleotide site diversity is approximately 5 x 10(-4) for both species, consistent with the previous observations of low variation at ci(D). Given the observed frequency spectra at ankyrin, coalescent simulations indicate that reduced diversity in the region is unlikely to be due to a selective sweep alone. We find evidence for recombinational exchange at this locus, and both species appear to be fixed for an insertion of the transposable element HB in an intron of ankyrin.

Animals↗

Population, evolutionary and genomic consequences of interference selection.

Weakly selected mutations are most likely to be physically clustered across genomes and, when sufficiently linked, they alter each others' fixation probability, a process we call interference selection (IS). Here we study population genetics and evolutionary consequences of IS on the selected mutations themselves and on adjacent selectively neutral variation. We show that IS reduces levels of polymorphism and increases low-frequency variants and linkage disequilibrium, in both selected and adjacent neutral mutations. IS can account for several well-documented patterns of variation and composition in genomic regions with low rates of crossing over in Drosophila. IS cannot be described simply as a reduction in the efficacy of selection and effective population size in standard models of selection and drift. Rather, IS can be better understood with models that incorporate a constant "traffic" of competing alleles. Our simulations also allow us to make genome-wide predictions that are specific to IS. We show that IS will be more severe at sites in the center of a region containing weakly selected mutations than at sites located close to the edge of the region. Drosophila melanogaster genomic data strongly support this prediction, with genes without introns showing significantly reduced codon bias in the center of coding regions. As expected, if introns relieve IS, genes with centrally located introns do not show reduced codon bias in the center of the coding region. We also show that reasonably small differences in the length of intermediate "neutral" sequences embedded in a region under selection increase the effectiveness of selection on the adjacent selected sequences. Hence, the presence and length of sequences such as introns or intergenic regions can be a trait subject to selection in recombining genomes. In support of this prediction, intron presence is positively correlated with a gene's codon bias in D. melanogaster. Finally, the study of temporal dynamics of IS after a change of recombination rate shows that nonequilibrium codon usage may be the norm rather than the exception.

Animals↗

Sequence variation and haplotype structure at the human HFE locus.

The HFE locus encodes an HLA class-I-type protein important in iron regulation and segregates replacement mutations that give rise to the most common form of genetic hemochromatosis. The high frequency of one disease-associated mutation, C282Y, and the nature of this disease have led some to suggest a selective advantage for this mutation. To investigate the context in which this mutation arose and gain a better understanding of HFE genetic variation, we surveyed nucleotide variability in 11.2 kb encompassing the HFE locus and experimentally determined haplotypes. We fully resequenced 60 chromosomes of African, Asian, or European ancestry as well as one chimpanzee, revealing 41 variable sites and a nucleotide diversity of 0.08%. This indicates that linkage to the HLA region has not substantially increased the level of HFE variation. Although several haplotypes are shared between populations, one haplotype predominates in Asia but is nearly absent elsewhere, causing higher than average genetic differentiation among the three major populations. Our samples show evidence of intragenic recombination, so the scarcity of recombination events within the C282Y allele class is consistent with selection increasing the frequency of a young allele. Otherwise, the pattern of variability in this region does not clearly indicate the action of positive selection at this or linked loci.

Animals↗

Pseudomonas viridiflava and P. syringae--natural pathogens of Arabidopsis thaliana.

We report the isolation and identification of two natural pathogens of Arabidopsis thaliana, Pseudomonas viridiflava and Pseudomonas syringae, in the midwestern United States. P. viridiflava was found in six of seven surveyed Arabidopsis thaliana populations. We confirmed the presence in the isolates of the critical pathogenicity genes hrpS and hrpL. The pathogenicity of these isolates was verified by estimating in planta bacterial growth rates and by testing for disease symptoms and hypersensitive responses to A. thaliana. Infection of 21 A. thaliana ecotypes with six locally collected P. viridiflava isolates and with one P. syringae isolate showed both compatible (disease) and incompatible (resistance) responses. Significant variation in response to infection was evident among Arabidopsis ecotypes, both in terms of symptom development and in planta bacterial growth. The ability to grow and cause disease symptoms on particular ecotypes also varied for some P. viridiflava isolates. We believe that these pathogens will provide a powerful system for exploring coevolution in natural plant-pathogen interactions.

Arabidopsis↗