PubMed Health⌕ Search

Biomedical subjects

David D Pollock

Publications and source records attributed to David D Pollock.

16 recordsLinked to original sources

EGenBio: a data management system for evolutionary genomics and biodiversity.

BACKGROUND: Evolutionary genomics requires management and filtering of large numbers of diverse genomic sequences for accurate analysis and inference on evolutionary processes of genomic and functional change. We developed Evolutionary Genomics and Biodiversity (EGenBio; http://egenbio.lsu.edu) to begin to address this. DESCRIPTION: EGenBio is a system for manipulation and filtering of large numbers of sequences, integrating curated sequence alignments and phylogenetic trees, managing evolutionary analyses, and visualizing their output. EGenBio is organized into three conceptual divisions, Evolution, Genomics, and Biodiversity. The Genomics division includes tools for selecting pre-aligned sequences from different genes and species, and for modifying and filtering these alignments for further analysis. Species searches are handled through queries that can be modified based on a tree-based navigation system and saved. The Biodiversity division contains tools for analyzing individual sequences or sequence alignments, whereas the Evolution division contains tools involving phylogenetic trees. Alignments are annotated with analytical results and modification history using our PRAED format. A miscellaneous Tools section and Help framework are also available. EGenBio was developed around our comparative genomic research and a prototype database of mtDNA genomes. It utilizes MySQL-relational databases and dynamic page generation, and calls numerous custom programs. CONCLUSION: EGenBio was designed to serve as a platform for tools and resources to ease combined analysis in evolution, genomics, and biodiversity.

Animals↗

Assessing the accuracy of ancestral protein reconstruction methods.

The phylogenetic inference of ancestral protein sequences is a powerful technique for the study of molecular evolution, but any conclusions drawn from such studies are only as good as the accuracy of the reconstruction method. Every inference method leads to errors in the ancestral protein sequence, resulting in potentially misleading estimates of the ancestral protein's properties. To assess the accuracy of ancestral protein reconstruction methods, we performed computational population evolution simulations featuring near-neutral evolution under purifying selection, speciation, and divergence using an off-lattice protein model where fitness depends on the ability to be stable in a specified target structure. We were thus able to compare the thermodynamic properties of the true ancestral sequences with the properties of "ancestral sequences" inferred by maximum parsimony, maximum likelihood, and Bayesian methods. Surprisingly, we found that methods such as maximum parsimony and maximum likelihood that reconstruct a "best guess" amino acid at each position overestimate thermostability, while a Bayesian method that sometimes chooses less-probable residues from the posterior probability distribution does not. Maximum likelihood and maximum parsimony apparently tend to eliminate variants at a position that are slightly detrimental to structural stability simply because such detrimental variants are less frequent. Other properties of ancestral proteins might be similarly overestimated. This suggests that ancestral reconstruction studies require greater care to come to credible conclusions regarding functional evolution. Inferred functional patterns that mimic reconstruction bias should be reevaluated.

Algorithms↗

Observations of amino acid gain and loss during protein evolution are explained by statistical bias.

The authors of a recent manuscript in "Nature" claim to have discovered "universal trends" of amino acid gain and loss in protein evolution. Here, we show that this universal trend can be simply explained by a bias that is unavoidable with the 3-taxon trees used in the original analysis. We demonstrate that a rigorously reversible equilibrium model, when analyzed with the same methods as the "Nature" manuscript, yields identical (and in this case, clearly erroneous) conclusions. A main source of the bias is the division of the sequence data into "informative" and "noninformative" sites, which favors the observation of certain transitions.

Algorithms↗

From DNA to fitness differences: sequences and structures of adaptive variants of Colias phosphoglucose isomerase (PGI).

Colias eurytheme butterflies display extensive allozyme polymorphism in the enzyme phosphoglucose isomerase (PGI). Earlier studies on biochemical and fitness effects of these genotypes found evidence of strong natural selection maintaining this polymorphism in the wild. Here we analyze the molecular features of this polymorphism by sequencing multiple alleles and modeling their structures. PGI is a dimer with rotational symmetry. Each monomer provides a critical residue to the other monomer's catalytic center. Sequenced alleles differ at multiple amino acid positions, including cryptic charge-neutral variation, but most consistent differences among the electromorph alleles are at the charge-changing amino acid sites. Principal candidate sites of selection, identified by structural and functional analyses and by their variants' population frequencies, occur in interpenetrating loops across the interface between monomers, where they may alter subunit interactions and catalytic center geometry. Comparison to a second (and basal) species, Colias meadii, also polymorphic for PGI under natural selection, reveals one fixed amino acid difference between their PGIs, which is located in the interpenetrating loop and accompanies functional differences among their variants. We also study nucleotide variability among the PGI alleles, comparing these data to similar data from another glycolytic enzyme gene, glyceraldehyde-3-phosphate dehydrogenase. Despite extensive nonsynonymous and synonymous polymorphism at PGI in each species, the only base changes fixed between species are the two causing the amino acid replacement; this absence of synonymous fixation yields a significant McDonald-Kreitman test. Analyses of these data suggest historical population expansion. Positive peaks of Tajima's D statistic, representing regions of neutral "hitchhiking," are found around the principal candidate sites of selection. This study provides novel views of molecular-structural mechanisms, and beginnings of historical evidence, for a long-persistent balanced enzyme polymorphism at PGI in these and perhaps other species.

Alleles↗

Context dependence and coevolution among amino acid residues in proteins.

As complete genomes accumulate and the generation of genomic biodiversity proceeds at an accelerating pace, the need to understand the interaction between sequence evolution and protein structure and function rises in prominence. The pattern and pace of substitutions in proteins can provide important clues to functional importance, functional divergence, and adaptive response. Coevolution between amino acid residues and the context dependence of the evolutionary process are often ignored, however, because of their complexity, but they are critical for the accurate interpretation of reconstructed evolutionary events. Because residues interact with one another, and because the effect of substitutions can depend on the structural and physiological environment in which they occur, an accurate science of evolutionary functional genomics and a complete understanding of selection in proteins require a better understanding of how context dependence affects protein evolution. Here, we present new evidence from vertebrate cytochrome oxidase sequences that pairwise coevolutionary interactions between protein residues are highly dependent on tertiary and secondary structure. We also discuss theoretical predictions that impinge on our expectations of how protein residues may interact over long distances because of their shared need to maintain protein stability.

Animals↗

The beetle gut: a hyperdiverse source of novel yeasts.

We isolated over 650 yeasts over a three year period from the gut of a variety of beetles and characterized them on the basis of LSU rDNA sequences and morphological and metabolic traits. Of these, at least 200 were undescribed taxa, a number equivalent to almost 30% of all currently recognized yeast species. A Bayesian analysis of species discovery rates predicts further sampling of previously sampled habitats could easily produce another 100 species. The sampled habitat is, thereby, estimated to contain well over half as many more species as are currently known worldwide. The beetle gut yeasts occur in 45 independent lineages scattered across the yeast phylogenetic tree, often in clusters. The distribution suggests that the some of the yeasts diversified by a process of horizontal transmission in the habitats and subsequent specialization in association with insect hosts. Evidence of specialization comes from consistent associations over time and broad geographical ranges of certain yeast and beetle species. The discovery of high yeast diversity in a previously unexplored habitat is a first step toward investigating the basis of the interactions and their impact in relation to ecology and evolution.

Animals↗

Evolution of base-substitution gradients in primate mitochondrial genomes.

Inferences of phylogenies and dates of divergence rely on accurate modeling of evolutionary processes; they may be confounded by variation in substitution rates among sites and changes in evolutionary processes over time. In vertebrate mitochondrial genomes, substitution rates are affected by a gradient along the genome of the time spent being single-stranded during replication, and different types of substitutions respond differently to this gradient. The gradient is controlled by biological factors including the rate of replication and functionality of repair mechanisms; little is known, however, about the consistency of the gradient over evolutionary time, or about how evolution of this gradient might affect phylogenetic analysis. Here, we evaluate the evolution of response to this gradient in complete primate mitochondrial genomes, focusing particularly on A-->G substitutions, which increase linearly with the gradient. We developed a methodology to evaluate the posterior probability densities of the response parameter space, and used likelihood ratio tests and mixture models with different numbers of classes to determine whether groups of genomes have evolved in a similar fashion. Substitution gradients usually evolve slowly in primates, but there have been at least two large evolutionary jumps: on the lineage leading to the great apes, and a convergent change on the lineage leading to baboons (Papio). There have also been possible convergences at deeper taxonomic levels, and different types of substitutions appear to evolve independently. The placements of the tarsier and the tree shrew within and in relation to primates may be incorrect because of convergence in these factors.

Animals↗

Divergence, recombination and retention of functionality during protein evolution.

We have only a vague idea of precisely how protein sequences evolve in the context of protein structure and function. This is primarily because structural and functional contexts are not easily predictable from the primary sequence, and evaluating patterns of evolution at individual residue positions is also difficult. As a result of increasing biodiversity in genomics studies, progress is being made in detecting context-dependent variation in substitution processes, but it remains unclear exactly what context-dependent patterns we should be looking for. To address this, we have been simulating protein evolution in the context of structure and function using lattice models of proteins and ligands (or substrates). These simulations include thermodynamic features of protein stability and population dynamics. We refer to this approach as 'ab initio evolution' to emphasise the fact that the equilibrium details of fitness distributions arise from the physical principles of the system and not from any preconceived notions or arbitrary mathematical distributions. Here, we present results on the retention of functionality in homologous recombinants following population divergence. A central result is that protein structure characteristics can strongly influence recombinant functionality. Exceptional structures with many sequence options evolve quickly and tend to retain functionality--even in highly diverged recombinants. By contrast, the more common structures with fewer sequence options evolve more slowly, but the fitness of recombinants drops off rapidly as homologous proteins diverge. These results have implications for understanding viral evolution, speciation and directed evolutionary experiments. Our analysis of the divergence process can also guide improved methods for accurately approximating folding probabilities in more complex but realistic systems.

Evolution, Molecular↗

Ancestral sequence reconstruction in primate mitochondrial DNA: compositional bias and effect on functional inference.

Reconstruction of ancestral DNA and amino acid sequences is an important means of inferring information about past evolutionary events. Such reconstructions suggest changes in molecular function and evolutionary processes over the course of evolution and are used to infer adaptation and convergence. Maximum likelihood (ML) is generally thought to provide relatively accurate reconstructed sequences compared to parsimony, but both methods lead to the inference of multiple directional changes in nucleotide frequencies in primate mitochondrial DNA (mtDNA). To better understand this surprising result, as well as to better understand how parsimony and ML differ, we constructed a series of computationally simple "conditional pathway" methods that differed in the number of substitutions allowed per site along each branch, and we also evaluated the entire Bayesian posterior frequency distribution of reconstructed ancestral states. We analyzed primate mitochondrial cytochrome b (Cyt-b) and cytochrome oxidase subunit I (COI) genes and found that ML reconstructs ancestral frequencies that are often more different from tip sequences than are parsimony reconstructions. In contrast, frequency reconstructions based on the posterior ensemble more closely resemble extant nucleotide frequencies. Simulations indicate that these differences in ancestral sequence inference are probably due to deterministic bias caused by high uncertainty in the optimization-based ancestral reconstruction methods (parsimony, ML, Bayesian maximum a posteriori). In contrast, ancestral nucleotide frequencies based on an average of the Bayesian set of credible ancestral sequences are much less biased. The methods involving simpler conditional pathway calculations have slightly reduced likelihood values compared to full likelihood calculations, but they can provide fairly unbiased nucleotide reconstructions and may be useful in more complex phylogenetic analyses than considered here due to their speed and flexibility. To determine whether biased reconstructions using optimization methods might affect inferences of functional properties, ancestral primate mitochondrial tRNA sequences were inferred and helix-forming propensities for conserved pairs were evaluated in silico. For ambiguously reconstructed nucleotides at sites with high base composition variability, ancestral tRNA sequences from Bayesian analyses were more compatible with canonical base pairing than were those inferred by other methods. Thus, nucleotide bias in reconstructed sequences apparently can lead to serious bias and inaccuracies in functional predictions.

Animals↗

The ambush hypothesis: hidden stop codons prevent off-frame gene reading.

Coding sequences lack stop codons, but many stops appear off-frame. Off-frame stops (stops in -1 and +1 shifted reading frames, termed hidden stops) terminate frame-shifted translation, potentially decreasing energy, and resource waste on nonfunctional proteins. Benefits may include reduced waste elimination costs and avoidance of potentially cytotoxic frame-shifted products. Our "ambush" hypothesis suggests that hidden stops are sometimes selected for. Codons of many amino acids can contribute to hidden stops, depending on the synonymous position state and adjacent codons. In vertebrate mitochondria, 31.75% of all amino acid combinations can form hidden stops. Codons with more potential to form hidden stops have greater usage frequency and bias in their favor among synonymous codons. Among primates, predicted mitochondrial rRNA secondary structure stability correlates negatively with the number of hidden stops in the mitochondrial genome. The taxonomic distribution of genetic codes suggests that +1 frameshifts might be more frequent than -1 frameshifts. This is confirmed by analyses of primate mitochondrial genomes: species with unstable rRNAs have more +1 stops, but the correlation is weak for -1 stops. High hidden stop density seems to be an adaptation in species with slippage prone ribosomes (unstable rRNAs). Hidden stops may thus compensate for reduced efficiency of some parts of the biosynthetic machinery. Some experimental data confirm our hypothesis: gene expression increases with the experimentally manipulated number of stops in the promoter region of a gene, suggesting biotechnological applications.

Animals↗

Detecting gradients of asymmetry in site-specific substitutions in mitochondrial genomes.

During mitochondrial replication, spontaneous mutations occur and accumulate asymmetrically during the time spent single stranded by the heavy strand (DssH). The predominant mutations appear to be deaminations from adenine to hypoxanthine (A --> H, which leads to an A --> G substitution) and cytosine to thymine (C --> T). Previous findings indicated that C --> T substitutions accumulate rapidly and then saturate at high DssH, suggesting protection or repair, whereas A --> G accumulates linearly with DssH. We describe here the implementation of a simple hidden Markov model (HMM) of among-site rate correlations to provide an almost continuous profile of the asymmetry in substitution response for any particular substitution type. We implement this model using a phylogeny-based Bayesian Markov chain Monte Carlo (MCMC) approach. We compare and contrast the relative asymmetries in all 12 possible substitution types, and find that the observed transition substitution responses determined using our new method agree quite well with previous predictions of a saturating curve for C --> T transition substitutions and a linear accumulation of A --> G transitions. The patterns seen in transversion substitutions show much lower among-site variation, and are nonlinear and more complex than those seen in transitions. We also find that, after accounting for the principal linear effect, some of the residual variation in A --> G/G --> A response ratios is explained by the average predicted nucleic acid secondary structure propensity at a site, possibly due to protection from mutation when secondary structure forms.

Genome↗

Estimating the degree of saturation in mutant screens.

Large-scale screens for loss-of-function mutants have played a significant role in recent advances in developmental biology and other fields. In such mutant screens, it is desirable to estimate the degree of "saturation" of the screen (i.e., what fraction of the possible target genes has been identified). We applied Bayesian and maximum-likelihood methods for estimating the number of loci remaining undetected in large-scale screens and produced credibility intervals to assess the uncertainty of these estimates. Since different loci may mutate to alleles with detectable phenotypes at different rates, we also incorporated variation in the degree of mutability among genes, using either gamma-distributed mutation rates or multiple discrete mutation rate classes. We examined eight published data sets from large-scale mutant screens and found that credibility intervals are much broader than implied by previous assumptions about the degree of saturation of screens. The likelihood methods presented here are a significantly better fit to data from published experiments than estimates based on the Poisson distribution, which implicitly assumes a single mutation rate for all loci. The results are reasonably robust to different models of variation in the mutability of genes. We tested our methods against mutant allele data from a region of the Drosophila melanogaster genome for which there is an independent genomics-based estimate of the number of undetected loci and found that the number of such loci falls within the predicted credibility interval for our models. The methods we have developed may also be useful for estimating the degree of saturation in other types of genetic screens in addition to classical screens for simple loss-of-function mutants, including genetic modifier screens and screens for protein-protein interactions using the yeast two-hybrid method.

Animals↗

Likelihood analysis of asymmetrical mutation bias gradients in vertebrate mitochondrial genomes.

Protein-coding genes in mitochondrial genomes have varying degrees of asymmetric skew in base frequencies at the third codon position. The variation in skew among genes appears to be caused by varying durations of time that the heavy strand spends in the mutagenic single-strand state during replication (D(ssH)). The primary data used to study skew have been the gene-by-gene base frequencies in individual taxa, which provide little information on exactly what kinds of mutations are responsible for the base frequency skew. To assess the contribution of individual mutation components to the ancestral vertebrate substitution pattern, here we analyze a large data set of complete vertebrate mitochondrial genomes in a phylogeny-based likelihood context. This also allows us to evaluate the change in skew continuously along the mitochondrial genome and to directly estimate relative substitution rates. Our results indicate that different types of mutation respond differently to the D(ssH) gradient. A primary role for hydrolytic deamination of cytosines in creating variance in skew among genes was not supported, but rather linearly increasing rates of mutation from adenine to hypoxanthine with D(ssH) appear to drive regional differences in skew. Substitutions due to hydrolytic deamination of cytosines, although common, appear to quickly saturate, possibly due to stabilization by the mitochondrial DNA single-strand-binding protein. These results should form the basis of more realistic models of DNA and protein evolution in mitochondria.

Animals↗

Genomic biodiversity, phylogenetics and coevolution in proteins.

Comprehensive sampling of genomic biodiversity is fast becoming a reality for some genomic regions and complete organelle genomes. Genomic biodiversity is defined as large genomic sequences from many species, and here some recent work is reviewed that demonstrates the potential benefits of genomic biodiversity for molecular evolutionary analysis and phylogenetic reconstruction. This work shows that using likelihood-based approaches, taxon addition can dramatically improve phylogenetic reconstruction. Features or dynamics of the evolutionary process are much more easily inferred with large numbers of taxa, and large numbers are essential for discriminating differences in evolutionary patterns between sites. Accurate prediction of site-specific patterns can improve phylogenetic reconstruction by an amount equivalent to quadrupling sequence length. Genomic biodiversity is particularly central to research relating patterns of evolution, adaptation and coevolution to structural and functional features of proteins. Research on detecting coevolution between amino acid residues in proteins demonstrates a clear need for much greater numbers of closely related taxa to better discriminate site-specific patterns of interaction, and to allow more detailed analysis of coevolutionary interactions between subunits in protein complexes. It is argued that parsing out coevolutionary and other context-dependent substitution probabilities is essential for discriminating between coevolution and adaptation, and for more realistically modelling the evolution of proteins. Also reviewed is research that argues for increasing the efficiency of acquiring genomic biodiversity, and suggests that this might be done by simultaneously shotgun cloning and sequencing genomic mixtures from many species. Increased efficiency is a prerequisite if genomic biodiversity levels are to rapidly increase by orders of magnitude, and thus lead to dramatically improved understanding of interactions between protein structure, function and sequence evolution.

Biodiversity↗