PubMed Health⌕ Search

Biomedical subjects

Gavin A Huttley

Publications and source records attributed to Gavin A Huttley.

4 recordsLinked to original sources

Vestige: maximum likelihood phylogenetic footprinting.

BACKGROUND: Phylogenetic footprinting is the identification of functional regions of DNA by their evolutionary conservation. This is achieved by comparing orthologous regions from multiple species and identifying the DNA regions that have diverged less than neutral DNA. Vestige is a phylogenetic footprinting package built on the PyEvolve toolkit that uses probabilistic molecular evolutionary modelling to represent aspects of sequence evolution, including the conventional divergence measure employed by other footprinting approaches. In addition to measuring the divergence, Vestige allows the expansion of the definition of a phylogenetic footprint to include variation in the distribution of any molecular evolutionary processes. This is achieved by displaying the distribution of model parameters that represent partitions of molecular evolutionary substitutions. Examination of the spatial incidence of these effects across regions of the genome can identify DNA segments that differ in the nature of the evolutionary process. RESULTS: Vestige was applied to a reference dataset of the SCL locus from four species and provided clear identification of the known conserved regions in this dataset. To demonstrate the flexibility to use diverse models of molecular evolution and dissect the nature of the evolutionary process Vestige was used to footprint the Ka/Ks ratio in primate BRCA1 with a codon model of evolution. Two regions of putative adaptive evolution were identified illustrating the ability of Vestige to represent the spatial distribution of distinct molecular evolutionary processes. CONCLUSION: Vestige provides a flexible, open platform for phylogenetic footprinting. Underpinned by the PyEvolve toolkit, Vestige provides a framework for visualising the signatures of evolutionary processes across the genome of numerous organisms simultaneously. By exploiting the maximum-likelihood statistical framework, the complex interplay between mutational processes, DNA repair and selection can be evaluated both spatially (along a sequence alignment) and temporally (for each branch of the tree) providing visual indicators to the attributes and functions of DNA sequences.

Algorithms↗

Modeling the impact of DNA methylation on the evolution of BRCA1 in mammals.

The modified base 5-methylcytosine ((m)C) plays an important functional role in the biology of mammals as an epigenetic modification and appears to exert a striking impact on the molecular evolution of mammal genomes. The collective epigenetic functions of (m)C revolve around its effect on gene transcription, while the influence of this modified base on the evolution of mammal genomes derives from the greatly elevated spontaneous mutation rate of (m)C to T. In mammals, (m)C occurs at the dinucleotides CpG, CpA, and CpT. As a step toward a comprehensive statistical examination of the role of (m)C in mammal molecular evolution, we have developed novel Markov models of codon substitution that incorporate dinucleotide-level terms relevant to (m)C mutation. We apply these models to two data sets of aligned BRCA1 exon 11 sequences from bats and primates. In all cases, terms specific to mutations that affect the dinucleotides CpG, CpA, and CpT significantly improved model fit. For the CpG-specific terms, both transition and transversion substitution rates were elevated. These rates differed between the data sets. Bats exhibited a lower relative rate of substitutions at CpG-containing codons. Transition substitutions were significantly less than 1 at CpA-containing codons but greater than 1 at CpT-containing codons. The inclusion of interaction terms in the codon models to represent possible confounding with the effect of natural selection were supported for codons that contained CpG and CpT, but not CpA. From the results, we infer that mutation of (m)C is a probable factor that affects BRCA1 codons containing the dinucleotide CpG, a possible factor for CpA-containing codons, and an unlikely factor that affects CpT-containing codons. The confounding of estimated terms with the effect of natural selection indicate this confounding must be addressed for comparisons between different coding and noncoding regions.

Animals↗

Modelling and bioinformatics studies of the human Kappa-class glutathione transferase predict a novel third glutathione transferase family with similarity to prokaryotic 2-hydroxychromene-2-carboxylate isomerases.

The Kappa class of GSTs (glutathione transferases) comprises soluble enzymes originally isolated from the mitochondrial matrix of rats. We have characterized a Kappa class cDNA from human breast. The cDNA is derived from a single gene comprising eight exons and seven introns located on chromosome 7q34-35. Recombinant hGSTK1-1 was expressed in Escherichia coli as a homodimer (subunit molecular mass approximately 25.5 kDa). Significant glutathione-conjugating activity was found only with the model substrate CDNB (1-chloro-2,4-ditnitrobenzene). Hyperbolic kinetics were obtained for GSH (parameters: K(m)app, 3.3+/-0.95 mM; V(max)app, 21.4+/-1.8 micromol/min per mg of enzyme), while sigmoidal kinetics were obtained for CDNB (parameters: S0.5app, 1.5+/-1.0 mM; V(max)app, 40.3+/-0.3 micromol/min per mg of enzyme; Hill coefficient, 1.3), reflecting low affinities for both substrates. Sequence analyses, homology modelling and secondary structure predictions show that hGSTK1 has (a) most similarity to bacterial HCCA (2-hydroxychromene-2-carboxylate) isomerases and (b) a predicted C-terminal domain structure that is almost identical to that of bacterial disulphide-bond-forming DsbA oxidoreductase (root mean square deviation 0.5-0.6 A). The structures of hGSTK1 and HCCA isomerase are predicted to possess a thioredoxin fold with a polyhelical domain (alpha(x)) embedded between the beta-strands (betaalphabetaalpha(x)betabetaalpha, where the underlined elements represent the N and C motifs of the thioredoxin fold), as occurs in the bacterial disulphide-bond-forming oxidoreductases. This is in contrast with the cytosolic GSTs, where the helical domain occurs exclusively at the C-terminus (betaalphabetaalphabetabetaalphaalpha(x)). Although hGSTK1-1 catalyses some typical GST reactions, we propose that it is structurally distinct from other classes of cytosolic GSTs. The present study suggests that the Kappa class may have arisen in prokaryotes well before the divergence of the cytosolic GSTs.

Amino Acid Sequence↗

PyEvolve: a toolkit for statistical modelling of molecular evolution.

BACKGROUND: Examining the distribution of variation has proven an extremely profitable technique in the effort to identify sequences of biological significance. Most approaches in the field, however, evaluate only the conserved portions of sequences - ignoring the biological significance of sequence differences. A suite of sophisticated likelihood based statistical models from the field of molecular evolution provides the basis for extracting the information from the full distribution of sequence variation. The number of different problems to which phylogeny-based maximum likelihood calculations can be applied is extensive. Available software packages that can perform likelihood calculations suffer from a lack of flexibility and scalability, or employ error-prone approaches to model parameterisation. RESULTS: Here we describe the implementation of PyEvolve, a toolkit for the application of existing, and development of new, statistical methods for molecular evolution. We present the object architecture and design schema of PyEvolve, which includes an adaptable multi-level parallelisation schema. The approach for defining new methods is illustrated by implementing a novel dinucleotide model of substitution that includes a parameter for mutation of methylated CpG's, which required 8 lines of standard Python code to define. Benchmarking was performed using either a dinucleotide or codon substitution model applied to an alignment of BRCA1 sequences from 20 mammals, or a 10 species subset. Up to five-fold parallel performance gains over serial were recorded. Compared to leading alternative software, PyEvolve exhibited significantly better real world performance for parameter rich models with a large data set, reducing the time required for optimisation from approximately 10 days to approximately 6 hours. CONCLUSION: PyEvolve provides flexible functionality that can be used either for statistical modelling of molecular evolution, or the development of new methods in the field. The toolkit can be used interactively or by writing and executing scripts. The toolkit uses efficient processes for specifying the parameterisation of statistical models, and implements numerous optimisations that make highly parameter rich likelihood functions solvable within hours on multi-cpu hardware. PyEvolve can be readily adapted in response to changing computational demands and hardware configurations to maximise performance. PyEvolve is released under the GPL and can be downloaded from http://cbis.anu.edu.au/software.

Animals↗