PubMed Health⌕ Search

Biomedical subjects

David A Liberles

Publications and source records attributed to David A Liberles.

At least 19 recordsLinked to original sources

Evolution after gene duplication: models, mechanisms, sequences, systems, and organisms.

Gene duplication is postulated to have played a major role in the evolution of biological novelty. Here, gene duplication is examined across levels of biological organization in an attempt to create a unified picture of the mechanistic process by which gene duplication can have played a role in generating biodiversity. Neofunctionalization and subfunctionalization have been proposed as important processes driving the retention of duplicate genes. These models have foundations in population genetic theory, which is now being refined by explicit consideration of the structural constraints placed upon genes encoding proteins through physical chemistry. Further, such models can be examined in the context of comparative genomics, where an integration of gene-level evolution and species-level evolution allows an assessment of the frequency of duplication and the fate of duplicate genes. This process, of course, is dependent upon the biochemical role that duplicated genes play in biological systems, which is in turn dependent upon the mechanism of duplication: whole genome duplication involving a co-duplication of interacting partners vs. single gene duplication. Lastly, the role that these processes may have played in driving speciation is examined.

Animals↗

Using evolutionary information and ancestral sequences to understand the sequence-function relationship in GLP-1 agonists.

Glucagon-like peptide-1 (GLP-1) is an incretin hormone with therapeutic potential for type 2 diabetes. A variety of GLP-1 sequences are known from amphibian species, and some of these have been tested here and found to be able to bind and activate the human GLP-1 receptor. While little difference was observed for the in vitro potency for the human GLP-1 receptor, larger differences were found in the enzymatic stability of these peptides. Two peptides showed increased enzymatic stability, and they group together phylogenetically, though they originate from Amphibia and Reptilia. We have used ancestral sequence reconstruction to analyze the evolution of these GLP-1 molecules, including the synthesis of new peptides. We find that the increased stability could not be observed in the resurrected peptides from the common ancestor of frogs, even though they maintain the ability to activate the human GLP-1 receptor. Another method, using residue mapping on evolutionary branches yielded peptides that had maintained potency towards the receptor and also showed increased stability. This represents a new approach using evolutionary data in protein engineering.

Amino Acid Sequence↗

A systematic analysis of lineage-specific evolution in metabolic pathways.

In a search for the lineage-specific evolution of pathways between human, chimpanzee, mouse, and rat, orthologous gene families were generated from genome sequences. For each family, a model-based ratio of nonsynonymous to synonymous nucleotide substitution rates was calculated. Where the free-ratio model of individual ratios on each branch was supported, these families were mapped to two databases of metabolic pathways (KEGG and BioCyc) and the lineage-specific evolution of pathways was evaluated. The most similar pathway evolution was seen between mouse and rat, while the evolutionary pattern between human and chimpanzee was less correlated. Individual pathways in the human lineage were observed to evolve in a faster, lineage-specific manner, including the pathway involving arachidonic acid metabolism (identified through the KEGG analysis) and pyrimidine metabolism (identified through both analyses).

Animals↗

Optimal gene trees from sequences and species trees using a soft interpretation of parsimony.

Gene duplication and gene loss as well as other biological events can result in multiple copies of genes in a given species. Because of these gene duplication and loss dynamics, in addition to variation in sequence evolution and other sources of uncertainty, different gene trees ultimately present different evolutionary histories. All of this together results in gene trees that give different topologies from each other, making consensus species trees ambiguous in places. Other sources of data to generate species trees are also unable to provide completely resolved binary species trees. However, in addition to gene duplication events, speciation events have provided some underlying phylogenetic signal, enabling development of algorithms to characterize these processes. Therefore, a soft parsimony algorithm has been developed that enables the mapping of gene trees onto species trees and modification of uncertain or weakly supported branches based on minimizing the number of gene duplication and loss events implied by the tree. The algorithm also allows for rooting of unrooted trees and for removal of in-paralogues (lineage-specific duplicates and redundant sequences masquerading as such). The algorithm has also been made available for download as a software package, Softparsmap.

Algorithms↗

Evaluation of models for the evolution of protein sequences and functions under structural constraint.

In the field of evolutionary structural genomics, methods are needed to evaluate why genomes evolved to contain the fold distributions that are observed. In order to study the effects of population dynamics in the evolved genomes we need fast and accurate evolutionary models which can analyze the effects of selection, drift and fixation of a protein sequence in a population that are grounded by physical parameters governing the folding and binding properties of the sequence. In this study, various knowledge-based, force field, and statistical methods for protein folding have been evaluated with four different folds: SH2 domains, SH3 domains, Globin-like, and Flavodoxin-like, to evaluate the speed and accuracy of the energy functions. Similarly, knowledge-based and force field methods have been used to predict ligand binding specificity in SH2 domain. To demonstrate the applicability of these methods, the dynamics of evolution of new binding capabilities by an SH2 domain is demonstrated.

Computational Biology↗

A systematic search for positive selection in higher plants (Embryophytes).

BACKGROUND: Previously, a database characterizing examples of Embryophyte gene family lineages showing evidence of positive selection was reported. Of the gene family trees, 138 Embryophyte branches showed Ka/Ks>>1 and are candidates for functional adaptation. The database and these examples have now been studied in further detail to better understand the molecular basis for plant genome evolution. RESULTS: Neutral modeling showed an excess of positive and/or negative selection in the database over a neutral expectation centered on the mean Ka/Ks ratio. Out of 673 families with assigned structures, 490 have at least one branch with Ka/Ks >>1 in a region of the protein, enabling a picture of selective pressures delineated by protein structure. Most gene families allowed reconstruction back to the last common ancestor of flowering plants (Magnoliophytes) without saturation of 4- fold degenerate codon position. Positive selection occurred in a wide variety of gene families with different functions, including in the self incompatibility locus, in defense against pathogens, in embryogenesis, in cold acclimation, and in electrontransport. Structurally, selective pressures were similar between alpha-helices and beta- sheets, but were less negative and more variant on the surface and away from the hydrophobic core. CONCLUSION: Positive selection was detected statistically significantly in a small and nonrandom minority of gene families in a systematic analysis of embryophyte gene families. More sensitive methods increased the level of positive selection that was detected and presented a structural basis for the role of positive selection in plant genomes.

Adaptation, Physiological↗

Characterization of hARD2, a processed hARD1 gene duplicate, encoding a human protein N-alpha-acetyltransferase.

BACKGROUND: Protein acetylation is increasingly recognized as an important mechanism regulating a variety of cellular functions. Several human protein acetyltransferases have been characterized, most of them catalyzing epsilon-acetylation of histones and transcription factors. We recently described the human protein acetyltransferase hARD1 (human Arrest Defective 1). hARD1 interacts with NATH (N-Acetyl Transferase Human) forming a complex expressing protein N-terminal alpha-acetylation activity. RESULTS: We here describe a human protein, hARD2, with 81 % sequence identity to hARD1. The gene encoding hARD2 most likely originates from a eutherian mammal specific retrotransposition event. hARD2 mRNA and protein are expressed in several human cell lines. Immunoprecipitation experiments show that hARD2 protein potentially interacts with NATH, suggesting that hARD2-NATH complexes may be responsible for protein N-alpha-acetylation in human cells. In NB4 cells undergoing retinoic acid mediated differentiation, the level of endogenous hARD1 and NATH protein decreases while the level of hARD2 protein is stable. CONCLUSION: A human protein N-alpha-acetyltransferase is herein described. ARD2 potentially complements the functions of ARD1, adding more flexibility and complexity to protein N-alpha-acetylation in human cells as compared to lower organisms which only have one ARD.

Acetylation↗

Analysis of transitions at two-fold redundant sites in mammalian genomes. Transition redundant approach-to-equilibrium (TREx) distance metrics.

BACKGROUND: The exchange of nucleotides at synonymous sites in a gene encoding a protein is believed to have little impact on the fitness of a host organism. This should be especially true for synonymous transitions, where a pyrimidine nucleotide is replaced by another pyrimidine, or a purine is replaced by another purine. This suggests that transition redundant exchange (TREx) processes at the third position of conserved two-fold codon systems might offer the best approximation for a neutral molecular clock, serving to examine, within coding regions, theories that require neutrality, determine whether transition rate constants differ within genes in a single lineage, and correlate dates of events recorded in genomes with dates in the geological and paleontological records. To date, TREx analysis of the yeast genome has recognized correlated duplications that established a new metabolic strategies in fungi, and supported analyses of functional change in aromatases in pigs. TREx dating has limitations, however. Multiple transitions at synonymous sites may cause equilibration and loss of information. Further, to be useful to correlate events in the genomic record, different genes within a genome must suffer transitions at similar rates. RESULTS: A formalism to analyze divergence at two fold redundant codon systems is presented. This formalism exploits two-state approach-to-equilibrium kinetics from chemistry. This formalism captures, in a single equation, the possibility of multiple substitutions at individual sites, avoiding any need to "correct" for these. The formalism also connects specific rate constants for transitions to specific approximations in an underlying evolutionary model, including assumptions that transition rate constants are invariant at different sites, in different genes, in different lineages, and at different times. Therefore, the formalism supports analyses that evaluate these approximations. Transitions at synonymous sites within two-fold redundant coding systems were examined in the mouse, rat, and human genomes. The key metric (f2), the fraction of those sites that holds the same nucleotide, was measured for putative ortholog pairs. A transition redundant exchange (TREx) distance was calculated from f2 for these pairs. Pyrimidine-pyrimidine transitions at these sites occur approximately 14% faster than purine-purine transitions in various lineages. Transition rate constants were similar in different genes within the same lineages; within a set of orthologs, the f2 distribution is only modest overdispersed. No correlation between disparity and overdispersion is observed. In rodents, evidence was found for greater conservation of TREx sites in genes on the X chromosome, accounting for a small part of the overdispersion, however. CONCLUSION: The TREx metric is useful to analyze the history of transition rate constants within these mammals over the past 100 million years. The TREx metric estimates the extent to which silent nucleotide substitutions accumulate in different genes, on different chromosomes, with different compositions, in different lineages, and at different times.

Animals↗

Datasets for evolutionary comparative genomics.

Many decisions about genome sequencing projects are directed by perceived gaps in the tree of life, or towards model organisms. With the goal of a better understanding of biology through the lens of evolution, however, there are additional genomes that are worth sequencing. One such rationale for whole-genome sequencing is discussed here, along with other important strategies for understanding the phenotypic divergence of species.

Animals↗

Phylogenetic reconstruction of ancestral character states for gene expression and mRNA splicing data.

BACKGROUND: As genomes evolve after speciation, gene content, coding sequence, gene expression, and splicing all diverge with time from ancestors with close relatives. A minimum evolution general method for continuous character analysis in a phylogenetic perspective is presented that allows for reconstruction of ancestral character states and for measuring along branch evolution. RESULTS: A software package for reconstruction of continuous character traits, like relative gene expression levels or alternative splice site usage data is presented and is available for download at http://www.rossnes.org/phyrex. This program was applied to a primate gene expression dataset to detect transcription factor binding sites that have undergone substitution, potentially having driven lineage-specific differences in gene expression. CONCLUSION: Systematic analysis of lineage-specific evolution is becoming the cornerstone of comparative genomics. New methods, like phyrex, extend the capabilities of comparative genomics by tracing the evolution of additional biomolecular processes.

Alternative Splicing↗

Subfunctionalization of duplicated genes as a transition state to neofunctionalization.

BACKGROUND: Gene duplication has been suggested to be an important process in the generation of evolutionary novelty. Neofunctionalization, as an adaptive process where one copy mutates into a function that was not present in the pre-duplication gene, is one mechanism that can lead to the retention of both copies. More recently, subfunctionalization, as a neutral process where the two copies partition the ancestral function, has been proposed as an alternative mechanism driving duplicate gene retention in organisms with small effective population sizes. The relative importance of these two processes is unclear. RESULTS: A set of lattice model genes that fold and bind to two peptide ligands with overlapping binding pockets, but not a third ligand present in the cell was designed. Each gene was duplicated in a model haploid species with a small constant population size and no recombination. One set of models allowed subfunctionalization of binding events following duplication, while another set did not allow subfunctionalization. Modeling under such conditions suggests that subfunctionalization plays an important role, but as a transition state to neofunctionalization rather than as a terminal fate of duplicated genes. There is no apparent selective pressure to maintain redundancy. CONCLUSION: Subfunctionalization results in an increase in the preservation of duplicated gene copies, including those that are neofunctionalized, but never represents a substantial fraction of duplicate gene copies at any evolutionary time point and ultimately leads to neofunctionalization of those preserved copies. This conclusion also may reflect changes in gene function after duplication with time in real genomes.

Animals↗

Catalysis, subcellular localization, expression and evolution of the targeting peptides degrading protease, AtPreP2.

We have previously identified a zinc metalloprotease involved in the degradation of mitochondrial and chloroplast targeting peptides, the presequence protease (PreP). In the Arabidopsis thaliana genomic database, there are two genes that correspond to the protease, the zinc metalloprotease (AAL90904) and the putative zinc metalloprotease (AAG13049). We have named the corresponding proteins AtPreP1 and AtPreP2, respectively. AtPreP1 and AtPreP2 show significant differences in their targeting peptides and the proteins are predicted to be localized in different compartments. AtPreP1 was shown to degrade both mitochondrial and chloroplast targeting peptides and to be dual targeted to both organelles using an ambiguous targeting peptide. Here, we have overexpressed, purified and characterized proteolytic and targeting properties of AtPreP2. AtPreP2 exhibits different proteolytic subsite specificity from AtPreP1 when used for degradation of organellar targeting peptides and their mutants. Interestingly, AtPreP2 precursor protein was also found to be dual targeted to both mitochondria and chloroplasts in a single and dual in vitro import system. Furthermore, targeting peptide of the AtPreP2 dually targeted green fluorescent protein (GFP) to both mitochondria and chloroplasts in tobacco protoplasts and leaves using an in vivo transient expression system. The targeting of both AtPreP1 and AtPreP2 proteases to chloroplasts in A. thaliana in vivo was confirmed via a shotgun mass spectrometric analysis of highly purified chloroplasts. Reverse transcription-polymerase chain reaction (RT-PCR) analysis revealed that AtPreP1 and AtPreP2 are differentially expressed in mature A. thaliana plants. Phylogenetic evidence indicated that AtPreP1 and AtPreP2 are recent gene duplicates that may have diverged through subfunctionalization.

Amino Acid Sequence↗

The Adaptive Evolution Database (TAED): a phylogeny based tool for comparative genomics.

From 138,662 embryophyte (higher plant) and 348,142 chordate genes, 4216 embryophyte and 15,452 chordate gene families were generated. For each of these gene families, multiple sequence alignments, phylogenetic trees, ratios of non-synonymous to synonymous nucleotide substitution rates (K(a)/K(s)), mappings from gene trees to the NCBI taxonomy and structural links to solved three-dimensional protein structures in the Protein Data Bank (PDB) with Grantham-weighted mutational factors were all calculated. Of the 'gene family trees', 173 embryophyte and 505 chordate branches show K(a)/K(s) >> 1 and are candidates for functional adaptation. The calculated information is available both as a gene family database and as a phylogenetically indexed resource, called 'The Adaptive Evolution Database' (TAED), available at http://www.bioinfo.no/tools/TAED.

Animals↗

Tertiary windowing to detect positive diversifying selection.

As a protein-encoding gene evolves, different selective pressures act on the gene temporally and spatially. An examination of the ratio of nonsynonymous-to-synonymous nucleotide substitution rate ratios (K(a)/K(s)) has proven to be a valuable method to examine selective pressures on protein encoding genes, including detecting positive diversifying selection. To gain power over averaging all sites in a gene together, examination of sites in primary sequence windows has frequently been employed. However, selection acts on folded proteins and sites that are close in tertiary space may not be close in primary sequence. A new method for the examination of K(a)/K(s) ratios based upon windows in tertiary structure is introduced and applied to the leptin gene family in mammals. Tertiary sequence windowing detects new sites under positive diversifying selection and detects positive diversifying selection with a more significant signal along various branches of the leptin gene family tree.

Animals↗

The planetary biology of cytochrome P450 aromatases.

BACKGROUND: Joining a model for the molecular evolution of a protein family to the paleontological and geological records (geobiology), and then to the chemical structures of substrates, products, and protein folds, is emerging as a broad strategy for generating hypotheses concerning function in a post-genomic world. This strategy expands systems biology to a planetary context, necessary for a notion of fitness to underlie (as it must) any discussion of function within a biomolecular system. RESULTS: Here, we report an example of such an expansion, where tools from planetary biology were used to analyze three genes from the pig Sus scrofa that encode cytochrome P450 aromatases-enzymes that convert androgens into estrogens. The evolutionary history of the vertebrate aromatase gene family was reconstructed. Transition redundant exchange silent substitution metrics were used to interpolate dates for the divergence of family members, the paleontological record was consulted to identify changes in physiology that correlated in time with the change in molecular behavior, and new aromatase sequences from peccary were obtained. Metrics that detect changing function in proteins were then applied, including KA/KS values and those that exploit structural biology. These identified specific amino acid replacements that were associated with changing substrate and product specificity during the time of presumed adaptive change. The combined analysis suggests that aromatase paralogs arose in pigs as a result of selection for Suoidea with larger litters than their ancestors, and permitted the Suoidea to survive the global climatic trauma that began in the Eocene. CONCLUSIONS: This combination of bioinformatics analysis, molecular evolution, paleontology, cladistics, global climatology, structural biology, and organic chemistry serves as a paradigm in planetary biology. As the geological, paleontological, and genomic records improve, this approach should become widely useful to make systems biology statements about high-level function for biomolecular systems.

Amino Acid Sequence↗

Visualising very large phylogenetic trees in three dimensional hyperbolic space.

BACKGROUND: Common existing phylogenetic tree visualisation tools are not able to display readable trees with more than a few thousand nodes. These existing methodologies are based in two dimensional space. RESULTS: We introduce the idea of visualising phylogenetic trees in three dimensional hyperbolic space with the Walrus graph visualisation tool and have developed a conversion tool that enables the conversion of standard phylogenetic tree formats to Walrus' format. With Walrus, it becomes possible to visualise and navigate phylogenetic trees with more than 100,000 nodes. CONCLUSION: Walrus enables desktop visualisation of very large phylogenetic trees in 3 dimensional hyperbolic space. This application is potentially useful for visualisation of the tree of life and for functional genomics derivatives, like The Adaptive Evolution Database (TAED).

Animals↗

Myostatin rapid sequence evolution in ruminants predates domestication.

Myostatin (GDF-8) is a negative regulator of skeletal muscle development. This gene has previously been implicated in the double muscling phenotype in mice and cattle. A systematic analysis of myostatin sequence evolution in ruminants was performed in a phylogenetic context. The myostatin coding sequence was determined from duiker (Sylvicapra grimmia caffra), eland (Taurotragus derbianus), gaur (Bos gaurus), ibex (Capra ibex), impala (Aepyceros melampus rednilis), pronghorn (Antilocapra americana), and tahr (Hemitragus jemlahicus). Analysis of nonsynonymous to synonymous nucleotide substitution rate ratios (Ka/Ks) indicates that positive selection may have been operating on this gene during the time of divergence of Bovinae and Antilopinae, starting from approximately 23 million years ago, a period that appears to account for most of the sequence difference between myostatin in these groups. These periods of positive selective pressure on myostatin may correlate with changes in skeletal muscle mass during the same period.

Amino Acid Sequence↗

Phylogenetic relationships of the Fox (Forkhead) gene family in the Bilateria.

The Forkhead or Fox gene family encodes putative transcription factors. There are at least four Fox genes in yeast, 16 in Drosophila melanogaster (Dm) and 42 in humans. Recently, vertebrate Fox genes have been classified into 17 groups named FoxA to FoxQ. Here, we extend this analysis to invertebrates, using available sequences from D. melanogaster, Anopheles gambiae (Ag), Caenorhabditis elegans (Ce), the sea squirt Ciona intestinalis (Ci) and amphioxus Branchiostoma floridae (Bf), from which we also cloned several Fox genes. Phylogenetic analyses lend support to the previous overall subclassification of vertebrate genes, but suggest that four subclasses (FoxJ, L, N and Q) could be further subdivided to reflect their relationships to invertebrate genes. We were unable to identify orthologs of Fox subclasses E, H, I, J, M and Q1 in D. melanogaster, A. gambiae or C. elegans, suggesting either considerable loss in ecdysozoans or the evolution of these subclasses in the deuterostome lineage. Our analyses suggest that the common ancestor of protostomes and deuterostomes had a minimum complement of 14 Fox genes.

Animals↗