PubMed Health⌕ Search

Biomedical subjects

Masatoshi Nei

Publications and source records attributed to Masatoshi Nei.

At least 19 recordsLinked to original sources

Selectionism and neutralism in molecular evolution.

Charles Darwin proposed that evolution occurs primarily by natural selection, but this view has been controversial from the beginning. Two of the major opposing views have been mutationism and neutralism. Early molecular studies suggested that most amino acid substitutions in proteins are neutral or nearly neutral and the functional change of proteins occurs by a few key amino acid substitutions. This suggestion generated an intense controversy over selectionism and neutralism. This controversy is partially caused by Kimura's definition of neutrality, which was too strict (|2Ns|< or =1). If we define neutral mutations as the mutations that do not change the function of gene products appreciably, many controversies disappear because slightly deleterious and slightly advantageous mutations are engulfed by neutral mutations. The ratio of the rate of nonsynonymous nucleotide substitution to that of synonymous substitution is a useful quantity to study positive Darwinian selection operating at highly variable genetic loci, but it does not necessarily detect adaptively important codons. Previously, multigene families were thought to evolve following the model of concerted evolution, but new evidence indicates that most of them evolve by a birth-and-death process of duplicate genes. It is now clear that most phenotypic characters or genetic systems such as the adaptive immune system in vertebrates are controlled by the interaction of a number of multigene families, which are often evolutionarily related and are subject to birth-and-death evolution. Therefore, it is important to study the mechanisms of gene family interaction for understanding phenotypic evolution. Because gene duplication occurs more or less at random, phenotypic evolution contains some fortuitous elements, though the environmental factors also play an important role. The randomness of phenotypic evolution is qualitatively different from allele frequency changes by random genetic drift. However, there is some similarity between phenotypic and molecular evolution with respect to functional or environmental constraints and evolutionary rate. It appears that mutation (including gene duplication and other DNA changes) is the driving force of evolution at both the genic and the phenotypic levels.

Animals↗

Evolutionary change of the numbers of homeobox genes in bilateral animals.

It has been known that the conservation or diversity of homeobox genes is responsible for the similarity and variability of some of the morphological or physiological characters among different organisms. To gain some insights into the evolutionary pattern of homeobox genes in bilateral animals, we studied the change of the numbers of these genes during the evolution of bilateral animals. We analyzed 2,031 homeodomain sequences compiled from 11 species of bilateral animals ranging from Caenorhabditis elegans to humans. Our phylogenetic analysis using a modified reconciled-tree method suggested that there were at least about 88 homeobox genes in the common ancestor of bilateral animals. About 50-60 genes of them have left at least one descendant gene in each of the 11 species studied, suggesting that about 30-40 genes were lost in a lineage-specific manner. Although similar numbers of ancestral genes have survived in each species, vertebrate lineages gained many more genes by duplication than invertebrate lineages, resulting in more than 200 homeobox genes in vertebrates and about 100 in invertebrates. After these gene duplications, a substantial number of old duplicate genes have also been lost in each lineage. Because many old duplicate genes were lost, it is likely that lost genes had already been differentiated from other groups of genes at the time of gene loss. We conclude that both gain and loss of homeobox genes were important for the evolutionary change of phenotypic characters in bilateral animals.

Algorithms↗

Evolutionary dynamics of olfactory receptor genes in fishes and tetrapods.

Olfaction, which is an important physiological function for the survival of mammals, is controlled by a large multigene family of olfactory receptor (OR) genes. Fishes also have this gene family, but the number of genes is known to be substantially smaller than in mammals. To understand the evolutionary dynamics of OR genes, we conducted a phylogenetic analysis of all functional genes identified from the genome sequences of zebrafish, pufferfish, frogs, chickens, humans, and mice. The results suggested that the most recent common ancestor between fishes and tetrapods had at least nine ancestral OR genes, and all OR genes identified were classified into nine groups, each of which originated from one ancestral gene. Eight of the nine group genes are still observed in current fish species, whereas only two group genes were found from mammalian genomes, showing that the OR gene family in fishes is much more diverse than in mammals. In mammals, however, one group of genes, gamma, expanded enormously, containing approximately 90% of the entire gene family. Interestingly, the gene groups observed in mammals or birds are nearly absent in fishes. The OR gene repertoire in frogs is as diverse as that in fishes, but the expansion of group gamma genes also occurred, indicating that the frog OR gene family has both mammal- and fish-like characters. All of these observations can be explained by the environmental change that organisms have experienced from the time of the common ancestor of all vertebrates to the present.

Animals↗

Origin and evolution of the chicken leukocyte receptor complex.

In mammals, the cell surface receptors encoded by the leukocyte receptor complex (LRC) regulate the activity of T lymphocytes and B lymphocytes, as well as that of natural killer cells, and thus provide protection against pathogens and parasites. The chicken genome encodes many Ig-like receptors that are homologous to the LRC receptors. The chicken Ig-like receptor (CHIR) genes are members of a large monophyletic gene family and are organized into genomic clusters, which are in conserved synteny with the mammalian LRC. One-third of CHIR genes encode polypeptide molecules that contain both activating and inhibitory motifs. These genes are present in different phylogenetic groups, suggesting that the primordial CHIR gene could have encoded both types of motifs in a single molecule. In contrast to the mammalian LRC genes, the CHIR genes with similar function (inhibition or activation) are evolutionarily closely related. We propose that, in addition to recombination, single nucleotide substitutions played an important role in the generation of receptors with different functions. Structural models and amino acid analyses of the CHIR proteins reveal the presence of different types of Ig-like domains in the same phylogenetic groups, as well as sharing of conserved residues and conserved changes of residues between different CHIR groups and between CHIRs and LRCs. Our data support the notion that the CHIR gene clusters are regions homologous to the mammalian LRC gene cluster and favor a model of evolution by repeated processes of birth and death (expansion-contraction) of the Ig-like receptor genes.

Animals↗

Rapid expansion of killer cell immunoglobulin-like receptor genes in primates and their coevolution with MHC Class I genes.

The gene family of killer cell immunoglobulin-like receptors (KIRs) in primates provides the first line of defense against virus infection and tumor transformation. Interacting with MHC class I molecules, KIRs can regulate the cytotoxic activity of natural killer (NK) cells and distinguish the tumor and virus infected cells from normal body cells. Phylogenetic analysis and comparison of domain structures identified three major groups of KIR genes (group I, II, and III genes). These groups of KIR genes, generated by a series of gene duplications, have acquired different MHC-binding specificity. Inference of ancestral KIR sequences suggested that the functional divergence of group I genes from group II genes occurred by positive selection at the MHC-binding sites after duplication. Our evolutionary study has shown that group I genes diverged from group II genes about 17 million years ago (Mya) apparently after separation of hominoids from Old World (OW) monkeys. Around the same time, gene duplication generating the class I MHC-C locus appears to have occurred. These findings suggest that KIR and MHC class I genes have coevolved as an interacting system. The KIR gene family has experienced a rapid expansion in primate species. The rate of expansion of this gene family seems to be one of the highest among all hominoid gene families. The KIR gene family is also subject to birth-and-death evolution.

Amino Acid Sequence↗

Eighty percent of proteins are different between humans and chimpanzees.

The chimpanzee is our closest living relative. The morphological differences between the two species are so large that there is no problem in distinguishing between them. However, the nucleotide difference between the two species is surprisingly small. The early genome comparison by DNA hybridization techniques suggested a nucleotide difference of 1-2%. Recently, direct nucleotide sequencing confirmed this estimate. These findings generated the common belief that the human is extremely close to the chimpanzee at the genetic level. However, if one looks at proteins, which are mainly responsible for phenotypic differences, the picture is quite different, and about 80% of proteins are different between the two species. Still, the number of proteins responsible for the phenotypic differences may be smaller since not all genes are directly responsible for phenotypic characters.

Animals↗

Origin and evolution of the Ig-like domains present in mammalian leukocyte receptors: insights from chicken, frog, and fish homologues.

In mammals many natural killer (NK) cell receptors, encoded by the leukocyte receptor complex (LRC), regulate the cytotoxic activity of NK cells and provide protection against virus-infected and tumor cells. To investigate the origin of the Ig-like domains encoded by the LRC genes, a subset of C2-type Ig-like domain sequences was compiled from mammals, birds, amphibians, and fish. Phylogenetic analysis of these sequences generated seven monophyletic groups in mammals (MI, MII, and FcI, FcIIa, FcIIb, FcIII, FcIV), two in chicken (CI, CII), four in frog (FI-FIV), and five in zebrafish (ZI-ZV). The analysis of the major groups supported the following order of divergence: ZI [or a common ancestor of ZI and F (a cluster composed of the FcIII and FIII groups)], F, CII (or a common ancestor of CII and MII), MII, and MI-CI. The relationships of the remaining groups were unclear, since the phylogenetic positions of these groups were not supported by high bootstrap values. Two main conclusions can be drawn from this analysis. First, the two groups of mammalian LRC sequences must diverged before the separation of the avian and mammalian lineages. Second, the mammalian LRC sequences are most closely related to the Fc receptor sequences and these two groups diverged before the separation of birds and mammals.

Animals↗

A simple method for predicting the functional differentiation of duplicate genes and its application to MIKC-type MADS-box genes.

A simple statistical method for predicting the functional differentiation of duplicate genes was developed. This method is based on the premise that the extent of functional differentiation between duplicate genes is reflected in the difference in evolutionary rate because the functional change of genes is often caused by relaxation or intensification of functional constraints. With this idea in mind, we developed a window analysis of protein sequences to identify the protein regions in which the significant rate difference exists. We applied this method to MIKC-type MADS-box proteins that control flower development in plants. We examined 23 pairs of sequences of floral MADS-box proteins from petunia and found that the rate differences for 14 pairs are significant. The significant rate differences were observed mostly in the K domain, which is important for dimerization between MADS-box proteins. These results indicate that our statistical method may be useful for predicting protein regions that are likely to be functionally differentiated. These regions may be chosen for further experimental studies.

Data Interpretation, Statistical↗

Comparative evolutionary analysis of olfactory receptor gene clusters between humans and mice.

Olfactory receptor (OR) genes form the largest multigene family in mammalian genomes. Humans have approximately 800 OR genes, but >50% of them are pseudogenes. By contrast, mice have approximately 1400 OR genes and pseudogenes are approximately 25%. To understand the evolutionary processes that shaped the difference of OR gene families between humans and mice, we studied the genomic locations of all human and mouse OR genes and conducted a detailed phylogenetic analysis using functional genes and pseudogenes. We identified 40 phylogenetic clades with high bootstrap supports, most of which contain both human and mouse genes. Interestingly, a particular clade contains approximately 100 pseudogenes in humans, whereas the numbers of pseudogenes are <20 for most of the mouse clades. We also found that the organization of OR genomic clusters is well conserved between humans and mice in many chromosomal locations. Despite the difference in the numbers of genes, the numbers of large genomic clusters are nearly the same for humans and mice. These observations suggest that the greater OR gene repertoire in mice has been generated mainly by tandem gene duplication within each genomic cluster.

Animals↗

Evolutionary changes of the number of olfactory receptor genes in the human and mouse lineages.

The numbers of functional olfactory receptor (OR) genes are quite variable among mammalian species. Previously we have reported that humans have 388 functional OR genes and 414 pseudogenes, while mice have 1037 functional genes and 354 pseudogenes. These observations suggest either that humans lost many functional OR genes after the human-mouse divergence (HMD) or that mice gained many functional genes. To distinguish between these two hypotheses, we devised a new method of inferring the number of functional OR genes in the most recent common ancestor (MRCA) of humans and mice. An application of this method suggested that the MRCA had approximately 750 functional OR genes and that mice acquired approximately 350 new OR genes after the HMD whereas approximately 430 OR genes in the MRCA have become pseudogenes or eliminated in the human lineage. Therefore, the two evolutionary hypotheses mentioned above are not mutually exclusive and both are nearly equally responsible for the difference in the number of OR genes between humans and mice.

Animals↗

Genomic organization and evolutionary analysis of Ly49 genes encoding the rodent natural killer cell receptors: rapid evolution by repeated gene duplication.

Ly49 genes regulate the cytotoxic activity of natural killer (NK) cells in rodents and provide important protection against virus-infected or tumor cells. About 15 Ly49 genes have been identified in mice, but only a few genes have been reported to date in rats. Here we studied all Ly49 genes in the entire rat genome sequence and identified 17 putative functional and 16 putative non-functional genes together with their genomic locations in a 1.8-Mb region of chromosome 4. Phylogenetic analysis of these genes indicated that the Ly49 gene family expanded rapidly in recent years, and this expansion was mediated by both tandem and genomic block duplication. The joint phylogenetic analysis of mouse and rat genes suggested that the most recent common ancestor of the two species had at least several Ly49 genes, but that the majority of current duplicate genes were generated after divergence of the two species. In both species Ly49 genes are apparently subject to birth-and-death evolution, but the birth and death rates of Ly49 genes are higher in rats than in mice. The rate of gene expansion in the Ly49 gene family in rats is one of the highest among all mammalian multigene families so far studied. The biochemical function of Ly49 genes is essentially the same as that of KIR genes in primates, but the molecular structures of the two groups of NK cell receptors are very different. A hypothesis was presented to explain the origin of the differential use of Ly49 and KIR genes in rodents and primates.

Animals↗

Prospects for inferring very large phylogenies by using the neighbor-joining method.

Current efforts to reconstruct the tree of life and histories of multigene families demand the inference of phylogenies consisting of thousands of gene sequences. However, for such large data sets even a moderate exploration of the tree space needed to identify the optimal tree is virtually impossible. For these cases the neighbor-joining (NJ) method is frequently used because of its demonstrated accuracy for smaller data sets and its computational speed. As data sets grow, however, the fraction of the tree space examined by the NJ algorithm becomes minuscule. Here, we report the results of our computer simulation for examining the accuracy of NJ trees for inferring very large phylogenies. First we present a likelihood method for the simultaneous estimation of all pairwise distances by using biologically realistic models of nucleotide substitution. Use of this method corrects up to 60% of NJ tree errors. Our simulation results show that the accuracy of NJ trees decline only by approximately 5% when the number of sequences used increases from 32 to 4,096 (128 times) even in the presence of extensive variation in the evolutionary rate among lineages or significant biases in the nucleotide composition and transition/transversion ratio. Our results encourage the use of complex models of nucleotide substitution for estimating evolutionary distances and hint at bright prospects for the application of the NJ and related methods in inferring large phylogenies.

Computer Simulation↗

False-positive selection identified by ML-based methods: examples from the Sig1 gene of the diatom Thalassiosira weissflogii and the tax gene of a human T-cell lymphotropic virus.

Sexually induced gene 1 (Sig1) in the centric diatom Thalassiosira weissflogii is considered to encode a gamete recognition protein. Sorhannus (2003) analyzed nucleotide sequences of Sig1 using parsimony analysis and the maximum-likelihood (ML)-based Bayesian method for inferring positive selection at single amino acid sites and reported that positively selected sites were detected by the latter method but not by the former. He then concluded that for this type of study, the ML-based method is more reliable than parsimony analysis. Here we show that his results apparently represent false-positive cases of the ML-based method and that there is no solid evidence that this gene contains positively selected sites. We further demonstrate that in the tax gene of human T-cell lymphotropic virus type I (HTLV-I), all codon sites, including invariable sites, can be inferred as positively selected sites by the ML-based method. These observations indicate that the ML-based method may produce many false-positive sites. One of the main reasons for the occurrence of false positives is that in the ML-based method, codon sites are grouped into several categories, with different nonsynonymous/synonymous rate ratios (omegas), on a purely statistical basis, and positive selection is inferred indirectly by examining whether the average omega for each category is greater than 1. In parsimony analysis, however, the evolutionary change of nucleotides at each codon site is examined. For this reason, parsimony-based methods rarely produce false positives and are safer than ML-based methods for detecting positive selection at individual codon sites, although a large number of sequences are necessary.

Bayes Theorem↗

Type I MADS-box genes have experienced faster birth-and-death evolution than type II MADS-box genes in angiosperms.

Plant MADS-box genes form a large gene family for transcription factors and are involved in various aspects of developmental processes, including flower development. They are known to be subject to birth-and-death evolution, but the detailed features of this mode of evolution remain unclear. To have a deeper insight into the evolutionary pattern of this gene family, we enumerated all available functional and nonfunctional (pseudogene) MADS-box genes from the Arabidopsis and rice genomes. Plant MADS-box genes can be classified into types I and II genes on the basis of phylogenetic analysis. Conducting extensive homology search and phylogenetic analysis, we found 64 presumed functional and 37 nonfunctional type I genes and 43 presumed functional and 4 nonfunctional type II genes in Arabidopsis. We also found 24 presumed functional and 6 nonfunctional type I genes and 47 presumed functional and 1 nonfunctional type II genes in rice. Our phylogenetic analysis indicated there were at least about four to eight type I genes and approximately 15-20 type II genes in the most recent common ancestor of Arabidopsis and rice. It has also been suggested that type I genes have experienced a higher rate of birth-and-death evolution than type II genes in angiosperms. Furthermore, the higher rate of birth-and-death evolution in type I genes appeared partly due to a higher frequency of segmental gene duplication and weaker purifying selection in type I than in type II genes.

Arabidopsis↗

MEGA3: Integrated software for Molecular Evolutionary Genetics Analysis and sequence alignment.

With its theoretical basis firmly established in molecular evolutionary and population genetics, the comparative DNA and protein sequence analysis plays a central role in reconstructing the evolutionary histories of species and multigene families, estimating rates of molecular evolution, and inferring the nature and extent of selective forces shaping the evolution of genes and genomes. The scope of these investigations has now expanded greatly owing to the development of high-throughput sequencing techniques and novel statistical and computational methods. These methods require easy-to-use computer programs. One such effort has been to produce Molecular Evolutionary Genetics Analysis (MEGA) software, with its focus on facilitating the exploration and analysis of the DNA and protein sequence variation from an evolutionary perspective. Currently in its third major release, MEGA3 contains facilities for automatic and manual sequence alignment, web-based mining of databases, inference of the phylogenetic trees, estimation of evolutionary distances and testing evolutionary hypotheses. This paper provides an overview of the statistical methods, computational tools, and visual exploration modules for data input and the results obtainable in MEGA.

Databases, Genetic↗

Concerted and nonconcerted evolution of the Hsp70 gene superfamily in two sibling species of nematodes.

We have identified the Hsp70 gene superfamily of the nematode Caenorhabditis briggsae and investigated the evolution of these genes in comparison with Hsp70 genes from C. elegans, Drosophila, and yeast. The Hsp70 genes are classified into three monophyletic groups according to their subcellular localization, namely, cytoplasm (CYT), endoplasmic reticulum (ER), and mitochondria (MT). The Hsp110 genes can be classified into the polyphyletic CYT group and the monophyletic ER group. The different Hsp70 and Hsp110 groups appeared to evolve following the model of divergent evolution. This model can also explain the evolution of the ER and MT genes. On the other hand, the CYT genes are divided into heat-inducible and constitutively expressed genes. The constitutively expressed genes have evolved more or less following the birth-and-death process, and the rates of gene birth and gene death are different between the two nematode species. By contrast, some heat-inducible genes show an intraspecies phylogenetic clustering. This suggests that they are subject to sequence homogenization resulting from gene conversion-like events. In addition, the heat-inducible genes show high levels of sequence conservation in both intra-species and inter-species comparisons, and in most cases, amino acid sequence similarity is higher than nucleotide sequence similarity. This indicates that purifying selection also plays an important role in maintaining high sequence similarity among paralogous Hsp70 genes. Therefore, we suggest that the CYT heat-inducible genes have been subjected to a combination of purifying selection, birth-and-death process, and gene conversion-like events.

Animals↗

Evolution of olfactory receptor genes in the human genome.

Olfactory receptor (OR) genes form the largest known multigene family in the human genome. To obtain some insight into their evolutionary history, we have identified the complete set of OR genes and their chromosomal locations from the latest human genome sequences. We detected 388 potentially functional genes that have intact ORFs and 414 apparent pseudogenes. The number and the fraction (48%) of functional genes are considerably larger than the ones previously reported. The human OR genes can clearly be divided into class I and class II genes, as was previously noted. Our phylogenetic analysis has shown that the class II OR genes can further be classified into 19 phylogenetic clades supported by high bootstrap values. We have also found that there are many tandem arrays of OR genes that are phylogenetically closely related. These genes appear to have been generated by tandem gene duplication. However, the relationships between genomic clusters and phylogenetic clades are very complicated. There are a substantial number of cases in which the genes in the same phylogenetic clade are located on different chromosomal regions. In addition, OR genes belonging to distantly related phylogenetic clades are sometimes located very closely in a chromosomal region and form a tight genomic cluster. These observations can be explained by the assumption that several chromosomal rearrangements have occurred at the regions of OR gene clusters and the OR genes contained in different genomic clusters are shuffled.

Biological Evolution↗