PubMed Health⌕ Search

Biomedical subjects

Hideki Innan

Publications and source records attributed to Hideki Innan.

At least 19 recordsLinked to original sources

Selection for more of the same product as a force to enhance concerted evolution of duplicated genes.

The duration of concerted evolution after gene duplication is highly variable across genes. To identify the cause of the variation, we analyzed of duplicated genes in yeast that originate from a whole genome duplication event. There appears to be a strong positive correlation between the duration of concerted evolution and the gene expression level. This observation can be explained by selection favoring more of the same product, which could enhance concerted evolution in dosage-sensitive genes.

Evolution, Molecular↗

Modified Hudson-Kreitman-Aguade test and two-dimensional evaluation of neutrality tests.

There are a number of polymorphism-based statistical tests of neutrality, but most of them focus on either the amount or the pattern of polymorphism. In this article, a new test called the two-dimensional (2D) test is developed. This test evaluates a pair of summary statistics in a two-dimensional field. One statistic should summarize the pattern of polymorphism, while the other could be a measure of the level of polymorphism. For the latter summary statistic, the polymorphism-divergence ratio is used following the idea of the Hudson-Kreitman-Aguadé (HKA) test. To incorporate the HKA test in the 2D test, a summary statistic-based version of the HKA test is developed such that the polymorphism-divergence ratio at a particular region of interest is examined if it is consistent with the average of those in other independent regions.

Computer Simulation↗

The effect of gene flow on the coalescent time in the human-chimpanzee ancestral population.

The coalescent process in the human-chimpanzee ancestral population is investigated using a model, which incorporates a certain time period of gene flow during the speciation process. a is a parameter to represent the degree and time of gene flow, and the model is identical to the null model with an instantaneous species split when a=infinity. A maximum likelihood (ML) method is developed to estimate a, and its power and reliability is investigated by coalescent simulations. The ML method is applied to nucleotide divergence data between human and chimpanzee. It is found that the null model with an instantaneous species split explains the data best, and no strong evidence for gene flow is detected. The result is discussed in the view of the mode of speciation. Another ML method is developed to estimate the male-female ratio (alpha) of mutation rate, in which the coalescent process in the ancestral population is taken into account.

Animals↗

Estimating the time to the whole-genome duplication and the duration of concerted evolution via gene conversion in yeast.

A maximum-likelihood (ML) method is developed to estimate the duration of concerted evolution and the time to the whole-genome duplication (WGD) event in baker's yeast (Saccharomyces cerevisiae). The models with concerted evolution fit the data significantly better than the molecular clock model, indicating a crucial role of concerted evolution via gene conversion after gene duplication in yeast. Our ML estimate of the time to the WGD is nearly identical to the time to the speciation event between S. cerevisiae and Kluyveromyces waltii, suggesting that the WGD occurred in very early stages after speciation or the WGD might have been involved in the speciation event.

Algorithms↗

Statistical tests of the coalescent model based on the haplotype frequency distribution and the number of segregating sites.

Several tests of neutral evolution employ the observed number of segregating sites and properties of the haplotype frequency distribution as summary statistics and use simulations to obtain rejection probabilities. Here we develop a "haplotype configuration test" of neutrality (HCT) based on the full haplotype frequency distribution. To enable exact computation of rejection probabilities for small samples, we derive a recursion under the standard coalescent model for the joint distribution of the haplotype frequencies and the number of segregating sites. For larger samples, we consider simulation-based approaches. The utility of the HCT is demonstrated in simulations of alternative models and in application to data from Drosophila melanogaster.

Chromosome Segregation↗

Very low gene duplication rate in the yeast genome.

The gene duplication rate in the yeast genome is estimated without assuming the molecular clock model to be approximately 0.01 to 0.06 per gene per billion years; this rate is two orders of magnitude lower than a previous estimate based on the molecular clock model. This difference is explained by extensive concerted evolution via gene conversion between duplicated genes, which violates the assumption of the molecular clock in the analyses of duplicated genes. The average length of the period of concerted evolution and the gene conversion rate are estimated to be approximately 25 million years and approximately 28 times the mutation rate, respectively.

Evolution, Molecular↗

The probability and chromosomal extent of trans-specific polymorphism.

Balancing selection may result in trans-specific polymorphism: the maintenance of allelic classes that transcend species boundaries by virtue of being more ancient than the species themselves. At the selected site, gene genealogies are expected not to reflect the species tree. Because of linkage, the same will be true for part of the surrounding chromosomal region. Here we obtain various approximations for the distribution of the length of this region and discuss the practical implications of our results. Our main finding is that the trans-specific region surrounding a single-locus balanced polymorphism is expected to be quite short, probably too short to be readily detectable. Thus lack of obvious trans-specific polymorphism should not be taken as evidence against balancing selection. When trans-specific polymorphism is obvious, on the other hand, it may be reasonable to argue that selection must be acting on multiple sites or that recombination is suppressed in the surrounding region.

Chromosomes↗

Pattern of polymorphism after strong artificial selection in a domestication event.

The process of strong artificial selection during a domestication event is modeled, and its effect on the pattern of DNA polymorphism is investigated. The model also considers population bottleneck during domestication. Artificial selection during domestication is different from a regular selective sweep because artificial selection acts on alleles that may have been neutral variants before domestication. Therefore, the fixation of such a beneficial allele does not always wipe out DNA variation in the surrounding region. The amount by which variation is reduced largely depends on the initial frequency of the beneficial allele, p. As a consequence, p has a strong effect on the likelihood of detecting the signature of selection during domestication from patterns of polymorphism. These theoretical results are discussed in light of data collected from maize. Although the main focus of this article is on domestication, this model can also be generalized to describe selective sweeps from standing genetic variation.

Alleles↗

Theories for analyzing polymorphism data in duplicated genes.

A simple model for the evolutionary process of a pair of duplicated genes under concerted evolution is developed. The model considers mutation, recombination and gene conversion between two genes in a finite population. Based on diffusion theory, the expected amount of DNA variation within and between two genes are obtained. To investigate the pattern of DNA polymorphism, a coalescent tool to simulate patterns of polymorphism is developed. The theoretical results are well in agreement with polymorphism data in duplicated genes. The effect of selection on the pattern of polymorphism is also considered.

Alleles↗

The effect of gene conversion on the divergence between duplicated genes.

Nonindependent evolution of duplicated genes is called concerted evolution. In this article, we study the evolutionary process of duplicated regions that involves concerted evolution. The model incorporates mutation and gene conversion: the former increases d, the divergence between two duplicated regions, while the latter decreases d. It is demonstrated that the process consists of three phases. Phase I is the time until d reaches its equilibrium value, d(0). In phase II d fluctuates around d(0), and d increases again in phase III. Our simulation results demonstrate that the length of concerted evolution (i.e., phase II) is highly variable, while the lengths of the other two phases are relatively constant. It is also demonstrated that the length of phase II approximately follows an exponential distribution with mean tau, which is a function of many parameters including gene conversion rate and the length of gene conversion tract. On the basis of these findings, we obtain the probability distribution of the level of divergence between a pair of duplicated regions as a function of time, mutation rate, and tau. Finally, we discuss potential problems in genomic data analysis of duplicated genes when it is based on the molecular clock but concerted evolution is common.

Animals↗

A two-locus gene conversion model with selection and its application to the human RHCE and RHD genes.

A two-locus gene conversion model with selection is developed. Under the joint action of selection, mutation, gene conversion, recombination, and random genetic drift, approximate formulas for the expectations of the moments of allele frequencies and the expected amounts of variation within and between two loci are obtained by a diffusion method assuming relatively strong selection. It is shown that the pattern of allelic variation is mainly determined by the balance between gene conversion and selection, because these two mechanisms act in opposite directions. As an application of the theoretical results, the human RHCE and RHD genes are considered. The very high level of amino acid divergence between the two genes is observed only in a short region around exon 7. It is known that exon 7 encodes amino acids that characterize the difference between the RHCE and RHD antigens. The observed pattern of DNA variation in this region is consistent with the selection model developed in this article, suggesting that strong selection might be working to maintain the RHCE/RHD antigen variation in the two-locus system. The selection intensity is estimated on the basis of the theoretical result.

Alleles↗

Relaxed selective pressure on an essential component of pheromone transduction in primate evolution.

The vomeronasal organ (VNO) detects pheromones in many vertebrate species but is likely to be vestigial in humans. TRPC2(TRP2), a gene that is essential for VNO function in the mouse, is a pseudogene in humans. Because TRPC2 is expressed only in the VNO, the loss of selective pressure on this gene can serve as a molecular marker for the time at which the VNO became vestigial. By analyzing sequence data from the TRPC2 gene of 15 extant primate species, we provide evidence that the VNO was most likely functional in the common ancestor of New World monkeys and Old World monkeys and apes, but then became vestigial in the common ancestor of Old World monkeys and apes. We propose that, at this point in evolution, other modalities, notably the development of color vision, may have largely replaced signaling by pheromones.

Amino Acid Sequence↗

The coalescent and infinite-site model of a small multigene family.

The infinite-site model of a small multigene family with two duplicated genes is studied. The expectations of the amounts of nucleotide variation within and between two genes and linkage disequilibrium are obtained, and a coalescent-based method for simulating patterns of polymorphism in a small multigene family is developed. The pattern of DNA variation is much more complicated than that in a single-copy gene, which can be simulated by the standard coalescent. Using the coalescent simulation of duplicated genes, the applicability of statistical tests of neutrality to multigene families is considered.

Data Interpretation, Statistical↗

The genealogy of sequences containing multiple sites subject to strong selection in a subdivided population.

A stochastic model for the genealogy of a sample of recombining sequences containing one or more sites subject to selection in a subdivided population is described. Selection is incorporated by dividing the population into allelic classes and then conditioning on the past sizes of these classes. The past allele frequencies at the selected sites are thus treated as parameters rather than as random variables. The purpose of the model is not to investigate the dynamics of selection, but to investigate effects of linkage to the selected sites on the genealogy of the surrounding chromosomal region. This approach is useful for modeling strong selection, when it is natural to parameterize the past allele frequencies at the selected sites. Several models of strong balancing selection are used as examples, and the effects on the pattern of neutral polymorphism in the chromosomal region are discussed. We focus in particular on the statistical power to detect balancing selection when it is present.

Animals↗

The extent of linkage disequilibrium and haplotype sharing around a polymorphic site.

Various expressions related to the length of a conserved haplotype around a polymorphism of known frequency are derived. We obtain exact expressions for the probability that no recombination has occurred in a sample or subsample. We obtain an approximation for the probability that no recombination that could give rise to a detectable recombination event (through the four-gamete test) has occurred. The probabilities can be used to obtain approximate distributions for the length of variously defined haplotypes around a polymorphic site. The implications of our results for data analysis, and in particular for detecting selection, are discussed.

Data Interpretation, Statistical↗

Distinguishing the hitchhiking and background selection models.

A simple method to distinguish hitchhiking and background selection is proposed. It is based on the observation that these models make different predictions about the average level of nucleotide diversity in regions of low recombination. The method is applied to data from Drosophila melanogaster and two highly selfing tomato species.

Animals↗

The pattern of polymorphism on human chromosome 21.

Polymorphism data from 20 partially resequenced copies of human chromosome 21-more than 20,000 polymorphic sites-were analyzed. The allele-frequency distribution shows no deviation from the simplest population genetic model with a constant population size (although we show that our analysis has no power to detect population growth). The average rate of recombination per site is estimated to be roughly one-half of the rate of mutation per site, again in agreement with simple model predictions. However, sliding-window analyses of the amount of polymorphism and the extent of linkage disequilibrium (LD) show significant deviations from standard models. This could be due to the history of selection or demographic change, but it is impossible to draw strong conclusions without much better knowledge of variation in the relationship between genetic and physical distance along the chromosome.

Alleles↗

Molecular population genetics.

Molecular population genetics is entering a new era dominated by studies of genomic polymorphism. Some of the theory that will be needed to analyze data generated by such studies is already available, but much more work is needed. Furthermore, population genetics is becoming increasingly relevant to other fields of biology, for example to genetic epidemiology, because of disease gene mapping in general populations.

Evolution, Molecular↗