PubMed Health⌕ Search

Biomedical subjects

Xun Gu

Publications and source records attributed to Xun Gu.

At least 19 recordsLinked to original sources

Functional divergence after gene duplication and sequence-structure relationship: a case study of G-protein alpha subunits.

In this article, we use animal G-protein alpha subunit family as an example to illustrate a comprehensive analytical pipeline for detecting different types of functional divergence of protein families, which is phylogeny-dependent, combined with ancestral sequence inference and available protein structure information. In particular, we focus on (i) Type-I functional divergence, or site-specific rate shift, as typically exemplified by amino acid residue highly conserved in a subset of homologous genes but highly variable in a different subset of homologous genes, and (ii) Type-II functional divergence, or the shift of cluster-specific amino acid property, as exemplified by a radical shift of amino acid property between duplicate genes, which is otherwise evolutionally conserved. We utilized the software DIVERGE2 to carry out these analyses. In the case of G-protein alpha subunit gene family, we have predicted amino acid residues that are related to either Type-I or Type-II functional divergence. The inferred ancestral sequences for these sites are helpful to explore the trends of functional divergence. Finally, these predicted residues are mapped to the protein structures to test whether these residues may have 3D structure or solvent accessibility preference.

Amino Acid Sequence↗

Stabilizing selection of protein function and distribution of selection coefficient among sites.

In this study, I take a new approach to modeling the evolutionary constraint of protein sequence, introducing the stabilizing selection of protein function into the nearly-neutral theory. In other words, protein function under stabilizing selection generates the evolutionary conservation at the sequence level. With the help of random mutational effects of nucleotides on protein function, I have derived the distribution of selection coefficient among sites, called the S-distribution whose parameters have clear biological interpretations. Moreover, I have studied the inverse relationship between the evolutionary rate and the effective population size, showing that the number of molecular phenotypes of protein function, i.e., independent components in the fitness of the organism, may play a key role for the molecular clock under the nearly-neutral theory. These results are helpful for having a better understanding of the underlying evolutionary mechanism of protein sequences, as well as human disease-related mutations.

Evolution, Molecular↗

Time squared: repeated measures on phylogenies.

Studies of gene expression profiles in response to external perturbation generate repeated measures data that generally follow nonlinear curves. To explore the evolution of such profiles across a gene family, we introduce phylogenetic repeated measures (PR) models. These models draw strength from 2 forms of correlation in the data. Through gene duplication, the family's evolutionary relatedness induces the first form. The second is the correlation across time points within taxonic units, individual genes in this example. We borrow a Brownian diffusion process along a given phylogenetic tree to account for the relatedness and co-opt a repeated measures framework to model the latter. Through simulation studies, we demonstrate that repeated measures models outperform the previously available approaches that consider the longitudinal observations or their differences as independent and identically distributed by using deviance information criteria as Bayesian model selection tools; PR models that borrow phylogenetic information also perform better than nonphylogenetic repeated measures models when appropriate. We then analyze the evolution of gene expression in the yeast kinase family using splines to estimate nonlinear behavior across 3 perturbation experiments. Again, the PR models outperform previous approaches and afford the prediction of ancestral expression profiles. To demonstrate PR model applicability more generally, we conclude with a short examination of variation in brain development across 4 primate species.

Algorithms↗

A simple statistical method for estimating type-II (cluster-specific) functional divergence of protein sequences.

Predicting functional amino acid residues in silico is important for comparative genomics. In this paper, we focus on the issue of how to statistically identify cluster-specific amino acid residues that are related to the functional divergence after gene duplication. We approach this problem using a framework based on site-specific shift of amino acid property (type-II functional divergence), as opposed to site-specific shift of evolutionary rate (type-I functional divergence). An efficient statistical procedure is implemented to facilitate the development of phylogenomic database for cluster-specific residues of large-scale protein families. Our method has the following features: 1) statistical testing of the type-II functional divergence and 2) the site-specific Bayesian profile to measure how amino acid residues contribute to type-II (cluster-specific) functional divergence. Consequently, one may obtain the posterior probability for "functional" cluster-specific residues. Case studies are presented and indicate that radical cluster-specific residues are responsible for most of inferred type-II functional divergence, whereas conserved cluster-specific residues appear less than even those imperfect radical cluster-specific residues to this type of functional divergence.

Amino Acid Sequence↗

Intron gain and loss in segmentally duplicated genes in rice.

BACKGROUND: Introns are under less selection pressure than exons, and consequently, intronic sequences have a higher rate of gain and loss than exons. In a number of plant species, a large portion of the genome has been segmentally duplicated, giving rise to a large set of duplicated genes. The recent completion of the rice genome in which segmental duplication has been documented has allowed us to investigate intron evolution within rice, a diploid monocotyledonous species. RESULTS: Analysis of segmental duplication in rice revealed that 159 Mb of the 371 Mb genome and 21,570 of the 43,719 non-transposable element-related genes were contained within a duplicated region. In these duplicated regions, 3,101 collinear paired genes were present. Using this set of segmentally duplicated genes, we investigated intron evolution from full-length cDNA-supported non-transposable element-related gene models of rice. Using gene pairs that have an ortholog in the dicotyledonous model species Arabidopsis thaliana, we identified more intron loss (49 introns within 35 gene pairs) than intron gain (5 introns within 5 gene pairs) following segmental duplication. We were unable to demonstrate preferential intron loss at the 3' end of genes as previously reported in mammalian genomes. However, we did find that the four nucleotides of exons that flank lost introns had less frequently used 4-mers. CONCLUSION: We observed that intron evolution within rice following segmental duplication is largely dominated by intron loss. In two of the five cases of intron gain within segmentally duplicated genes, the gained sequences were similar to transposable elements.

Amino Acid Sequence↗

Neurotransmitter inactivation is important for the origin of nerve system in animal early evolution: a suggestion from genomic comparison.

Metazoans possess complicated multicellular structure among the multicellularities of eukaryotes. One evolutionary pressure that permits such complexity relates to the directed and precise informational transmission performed by numerous synapses in neuron system. Neurotransmitter inactivations play essential roles in the termination of synaptic transmission and are thus crucial for precise synaptic transmission. Here, we performed a genomic comparison among 11 eukaryotic organisms including five bilaterian species and six pan-unicellular eukaryotes to search for genes related to metazoan multicellular function. The result showed that the majority of genes related to neurotransmitter inactivation in the synaptic cleft endured high and stable selective pressure and were specifically present in bilaterians, whereas genes related to transmitter release and postsynaptic transmitter receptors did not show these properties. From these data we conclude that neurotransmitter inaction may play a critical role in the origin of the nerve system encountered in the early evolution of metazoan. In addition, we suggest that neurotransmitter inactivation probably participates in the formation or refinement of the synapse, following the concept of "ontogeny recapitulates phylogeny." Further experimental evidence is needed to support the suggestion and to explain the importance of neurotransmitter inactivation to metazoan multicellular function.

Animals↗

Taking the first steps towards a standard for reporting on phylogenies: Minimum Information About a Phylogenetic Analysis (MIAPA).

In the eight years since phylogenomics was introduced as the intersection of genomics and phylogenetics, the field has provided fundamental insights into gene function, genome history and organismal relationships. The utility of phylogenomics is growing with the increase in the number and diversity of taxa for which whole genome and large transcriptome sequence sets are being generated. We assert that the synergy between genomic and phylogenetic perspectives in comparative biology would be enhanced by the development and refinement of minimal reporting standards for phylogenetic analyses. Encouraged by the development of the Minimum Information About a Microarray Experiment (MIAME) standard, we propose a similar roadmap for the development of a Minimal Information About a Phylogenetic Analysis (MIAPA) standard. Key in the successful development and implementation of such a standard will be broad participation by developers of phylogenetic analysis software, phylogenetic database developers, practitioners of phylogenomics, and journal editors.

Genomics↗

Evolution of alternative splicing after gene duplication.

Alternative splicing and gene duplication are two major sources of proteomic function diversity. Here, we study the evolutionary trend of alternative splicing after gene duplication by analyzing the alternative splicing differences between duplicate genes. We observed that duplicate genes have fewer alternative splice (AS) forms than single-copy genes, and that a negative correlation exists between the mean number of AS forms and the gene family size. Interestingly, we found that the loss of alternative splicing in duplicate genes may occur shortly after the gene duplication. These results support the subfunctionization model of alternative splicing in the early stage after gene duplication. Further analysis of the alternative splicing distribution in human duplicate pairs showed the asymmetric evolution of alternative splicing after gene duplications; i.e., the AS forms between duplicates may differ dramatically. We therefore conclude that alternative splicing and gene duplication may not evolve independently. In the early stage after gene duplication, young duplicates may take over a certain amount of protein function diversity that previously was carried out by the alternative splicing mechanism. In the late stage, the gain and loss of alternative splicing seem to be independent between duplicates.

Alternative Splicing↗

Expression divergence between duplicate genes.

A general picture of the role of expression divergence in the evolution of duplicate genes is emerging, thanks to the availability of completely sequenced genomes and functional genomic data, such as microarray data. It is now clear that expression divergence, regulatory-motif divergence and coding-sequence divergence all increase with the age of duplicate genes, although their exact interrelationships remain to be determined. It is also clear that gene duplication increases expression diversity and enables tissue or developmental specialization to evolve. However, the relative roles of subfunctionalization and neofunctionalization in the retention of duplicate genes remain to be clarified, especially for higher eukaryotes. In addition, the relationship between gene duplication and evolution of transcriptional regulatory networks is largely unexplored.

Evolution, Molecular↗

SplitTester: software to identify domains responsible for functional divergence in protein family.

BACKGROUND: Many protein families have undergone functional divergence after gene duplications such that current subgroups of the family carry out overlapping but distinct biological roles. For the protein families with known functional subtypes (a functional split), we developed the software, SplitTester, to identify potential regions that are responsible for the observed distinct functional subtypes within the same protein family. RESULTS: Our software, SplitTester, takes a multiple protein sequences alignment as input, generated from protein members of two subgroups with known functional divergence. SplitTester was designed to construct the neighbor joining tree (a split cluster) from variable-sized sliding windows across the alignment in a process called split-clustering. SplitTester identifies the regions, whose split cluster is consistent with the functional split, but may be inconsistent with the phylogeny of the protein family. We hypothesize that at least some number of these identified regions, which are not following a random mutation process, are responsible for the observed functional split. To test our method, we used reverse transcriptase from a group of Pseudoviridae retrotransposons: to identify residues specific for diverged primer recognition. Candidate regions were then mapped onto the three dimensional structures of reverse transcriptase. The locations of these amino acids within the enzyme are consistent with their biological roles. CONCLUSION: SplitTester aims to identify specific domain sequences responsible for functional divergence of subgroups within a protein family. From the analysis of retroelements reverse transcriptase family, we successfully identified the regions splitting this family according to the primer specificity, implying their functions in the specific primer selection.

Algorithms↗

Rapid evolution of expression and regulatory divergences after yeast gene duplication.

Although gene duplication is widely believed to be the major source of genetic novelty, how the expression or regulatory network of duplicate genes evolves remains poorly understood. In this article, we propose an additive expression distance between duplicate genes, so that the evolutionary rate of expression divergence after gene duplication can be estimated through phylogenomic analysis. We have analyzed yeast genome sequences, microarrays, and transcriptional regulatory networks, showing a >10-fold increase in the initial rate for both expression and regulatory network evolution after gene duplication but only an approximately 20% rate increase in the early stage for protein sequences. Based on the estimated age distribution of yeast duplicate genes, we roughly estimate that the initial rate of expression divergence shortly after gene duplication is 2.9 x 10(-9) per year, whereas the baseline rate for very ancient gene duplication is 0.14 x 10(-9) per year. Relative expression rate tests suggest that the expression of duplicate genes tends to evolve asymmetrically, that is, the expression of one copy evolves rapidly, whereas the other one largely maintains the ancestral expression profile. Our study highlights the crucial role of early rapid evolution after gene/genome duplication for continuously increasing the complexity of the yeast regulatory network.

Evolution, Molecular↗

Web-based resources for comparative genomics.

The available web-based genome data and related resources provide great opportunities for biomedical scientists to identify functional elements in a particular genome region or to explore the evolutionary pattern of genome dynamics. Comparative genomics is an indispensable tool for achieving these goals. Because of the broad scope of comparative genomics, it is difficult to address all of its aspects in this short survey. A few currently 'hot' topics have therefore been selected and a brief review of the availability of web-based databases and software is given.

Animals↗

Predicting binding sites of hydrolase-inhibitor complexes by combining several methods.

BACKGROUND: Protein-protein interactions play a critical role in protein function. Completion of many genomes is being followed rapidly by major efforts to identify interacting protein pairs experimentally in order to decipher the networks of interacting, coordinated-in-action proteins. Identification of protein-protein interaction sites and detection of specific amino acids that contribute to the specificity and the strength of protein interactions is an important problem with broad applications ranging from rational drug design to the analysis of metabolic and signal transduction networks. RESULTS: In order to increase the power of predictive methods for protein-protein interaction sites, we have developed a consensus methodology for combining four different methods. These approaches include: data mining using Support Vector Machines, threading through protein structures, prediction of conserved residues on the protein surface by analysis of phylogenetic trees, and the Conservatism of Conservatism method of Mirny and Shakhnovich. Results obtained on a dataset of hydrolase-inhibitor complexes demonstrate that the combination of all four methods yield improved predictions over the individual methods. CONCLUSIONS: We developed a consensus method for predicting protein-protein interface residues by combining sequence and structure-based methods. The success of our consensus approach suggests that similar methodologies can be developed to improve prediction accuracies for other bioinformatic problems.

Algorithms↗

GeneContent: software for whole-genome phylogenetic analysis.

UNLABELLED: GeneContent is a software system to infer the genome phylogeny based on an additive genome distance that can be estimated from the extended gene content data, which contains the genome-wide information (absence of a gene family, presence as single copy or presence as duplicates) across multiple species. GeneContent can also be used to explore the genome-wide evolutionary pattern of gene loss and proliferation. AVAILABILITY: Distribution packages of GeneContent for both Microsoft Windows and Linux operating systems are available at http://xgu.zool.iastate.edu CONTACT: xgu@iastate.edu.

Algorithms↗

Maximum likelihood for genome phylogeny on gene content.

With the rapid growth of entire genome data, reconstructing the phylogenetic relationship among different genomes has become a hot topic in comparative genomics. Maximum likelihood approach is one of the various approaches, and has been very successful. However, there is no reported study for any applications in the genome tree-making mainly due to the lack of an analytical form of a probability model and/or the complicated calculation burden. In this paper we studied the mathematical structure of the stochastic model of genome evolution, and then developed a simplified likelihood function for observing a specific phylogenetic pattern under four genome situation using gene content information. We use the maximum likelihood approach to identify phylogenetic trees. Simulation results indicate that the proposed method works well and can identify trees with a high correction rate. Real data application provides satisfied results. The approach developed in this paper can serve as the basis for reconstructing phylogenies of more than four genomes.

Journal Article↗

Identification of conserved gene structures and carboxy-terminal motifs in the Myb gene family of Arabidopsis and Oryza sativa L. ssp. indica.

BACKGROUND: Myb proteins contain a conserved DNA-binding domain composed of one to four repeat motifs (referred to as R0R1R2R3); each repeat is approximately 50 amino acids in length, with regularly spaced tryptophan residues. Although the Myb proteins comprise one of the largest families of transcription factors in plants, little is known about the functions of most Myb genes. Here we use computational techniques to classify Myb genes on the basis of sequence similarity and gene structure, and to identify possible functional relationships among subgroups of Myb genes from Arabidopsis and rice (Oryza sativa L. ssp. indica). RESULTS: This study analyzed 130 Myb genes from Arabidopsis and 85 from rice. The collected Myb proteins were clustered into subgroups based on sequence similarity and phylogeny. Interestingly, the exon-intron structure differed between subgroups, but was conserved in the same subgroup. Moreover, the Myb domains contained a significant excess of phase 1 and 2 introns, as well as an excess of nonsymmetric exons. Conserved motifs were detected in carboxy-terminal coding regions of Myb genes within subgroups. In contrast, no common regulatory motifs were identified in the noncoding regions. Additionally, some Myb genes with similar functions were clustered in the same subgroups. CONCLUSIONS: The distribution of introns in the phylogenetic tree suggests that Myb domains originally were compact in size; introns were inserted and the splicing sites conserved during evolution. Conserved motifs identified in the carboxy-terminal regions are specific for Myb genes, and the identified Myb gene subgroups may reflect functional conservation.

Amino Acid Motifs↗

Genome phylogenetic analysis based on extended gene contents.

With the rapid growth of entire genome data, whole-genome approaches such as gene content become popular for genome phylogeny inference, including the tree of life. However, the underlying model for genome evolution is unclear, and the proposed (ad hoc) genome distance measure may violate the additivity. In this article, we formulate a stochastic framework for genome evolution, which provides a basis for defining an additive genome distance. However, we show that it is difficult to utilize the typical gene content data-i.e., the presence or absence of gene families across genomes-to estimate the genome distance. We solve this problem by introducing the concept of extended gene content; that is, the status of a gene family in a given genome could be absence, presence as single copy, or presence as duplicates, any of which can be used to estimate the genome distance and phylogenetic inference. Computer simulation shows that the new tree-making method is efficient, consistent, and fairly robust. The example of 35 microbial complete genomes demonstrates that it is useful not only to study the universal tree of life but also to explore the evolutionary pattern of genomes.

Computational Biology↗

Ordered origin of the typical two- and three-repeat Myb genes.

Myb domain proteins contain a conserved DNA-binding domain composed of one to four conserved repeat motifs. In animals, Myb proteins are encoded by a small gene family and commonly contain three repeat motifs (R1R2R3); whereas, plant Myb proteins are encoded by a very large and diverse gene family in which a motif containing two repeats (R2R3) is the most common. In contrast to the conservation in the Myb domain, other regions of Myb proteins are highly variable. To explore the evolutionary origin of Myb genes, we cloned and sequenced Myb domains from maize and sorghum, and conducted a comprehensive phylogenetic analysis of Myb genes. The results indicate that the origins of individual Myb repeats are strikingly distinct, and that the R2 repeat has evolved more slowly than the R1 and R3 repeats. However, it is not clear which repeat is the most ancient one. The evidence also suggests that R2R3 and R1R2R3 Myb genes co-existed in eukaryotes before the divergence of plants and animals. Based on our results, we propose that R1R2R3 Myb genes were derived from R2R3 Myb genes by gain of the R1 repeat through an ancient intragenic duplication; this gain model is more parsimonious than the previous proposal that R2R3 Myb genes were derived from R1R2R3 Mybs by loss of the R1 repeat. A separate group of diverse non-typical Myb proteins exhibits a polyphyletic origin and a complex evolutionary pattern. Finally, a small group of ancient Myb paralogs prior to the amplification of current Myb genes is identified. Together, these results support a new model for the ordered evolution of Myb gene family.

Binding Sites↗