PubMed Health⌕ Search

PubMed · 14706121

PyEvolve: a toolkit for statistical modelling of molecular evolution.

Abstract

BACKGROUND: Examining the distribution of variation has proven an extremely profitable technique in the effort to identify sequences of biological significance. Most approaches in the field, however, evaluate only the conserved portions of sequences - ignoring the biological significance of sequence differences. A suite of sophisticated likelihood based statistical models from the field of molecular evolution provides the basis for extracting the information from the full distribution of sequence variation. The number of different problems to which phylogeny-based maximum likelihood calculations can be applied is extensive. Available software packages that can perform likelihood calculations suffer from a lack of flexibility and scalability, or employ error-prone approaches to model parameterisation. RESULTS: Here we describe the implementation of PyEvolve, a toolkit for the application of existing, and development of new, statistical methods for molecular evolution. We present the object architecture and design schema of PyEvolve, which includes an adaptable multi-level parallelisation schema. The approach for defining new methods is illustrated by implementing a novel dinucleotide model of substitution that includes a parameter for mutation of methylated CpG's, which required 8 lines of standard Python code to define. Benchmarking was performed using either a dinucleotide or codon substitution model applied to an alignment of BRCA1 sequences from 20 mammals, or a 10 species subset. Up to five-fold parallel performance gains over serial were recorded. Compared to leading alternative software, PyEvolve exhibited significantly better real world performance for parameter rich models with a large data set, reducing the time required for optimisation from approximately 10 days to approximately 6 hours. CONCLUSION: PyEvolve provides flexible functionality that can be used either for statistical modelling of molecular evolution, or the development of new methods in the field. The toolkit can be used interactively or by writing and executing scripts. The toolkit uses efficient processes for specifying the parameterisation of statistical models, and implements numerous optimisations that make highly parameter rich likelihood functions solvable within hours on multi-cpu hardware. PyEvolve can be readily adapted in response to changing computational demands and hardware configurations to maximise performance. PyEvolve is released under the GPL and can be downloaded from http://cbis.anu.edu.au/software.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Andrew Butterfield, Vivek Vedagiri, Edward Lang, Cath Lawrence, Matthew J Wakefield, Alexander Isaev, Gavin A Huttley. 2004-01-05. PyEvolve: a toolkit for statistical modelling of molecular evolution.. https://doi.org/10.1186/1471-2105-5-1

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Protocol for Detecting and Sequencing Chikungunya Virus from Field-Collected Mosquitoes.

Arboviral diseases represent a major public health challenge, especially in tropical regions where environmental conditions may favor the proliferation and spread of mosquito vectors. Thus, early and accurate detection of chikungunya virus (CHIKV) in mosquito populations can be a valuable tool for effective surveillance of circulating variants and for identifying new viral introductions. Given the challenges of detecting arboviruses in field-captured mosquitoes, we describe an integrated workflow for CHIKV molecular detection and whole-genome sequencing. This protocol includes mosquito homogenization using a bead-based mechanical disruptor, RNA extraction using TRIzol reagent with minor modifications, molecular screening using CHIKV-specific RT-qPCR, and whole-genome amplification followed by sequencing on Illumina platforms. Despite the protocol being optimized for individual mosquitoes, it results in high-quality RNA suitable for both entomological surveillance and genomic analysis. As this protocol allows recovery of complete CHIKV genomes from mosquito specimens, it can serve as a basis for genomic epidemiology studies, enabling monitoring of viral diversity and lineage dynamics, and facilitating early detection of emerging variants to support timely and targeted public health interventions in endemic and at-risk regions.

Animals↗

Genomic Profiling of Chromatin State Using CUT&Tag.

Alterations in chromatin state, mediated through histone modifications and the incorporation of histone variants, are fundamental to establishing transcriptional networks and cell identity. Recent advances in low-input epigenome profiling methods, such as CUT&Tag and CUT&RUN, have enabled the study of chromatin states from very limited starting materials. In this chapter, we describe procedures for generating CUT&Tag libraries to profile histone modifications and histone variants in early-developing zebrafish embryos.

Animals↗

Relaxin-2: Shaping the Proteomic Landscape of Skeletal Muscle Physiology, Glucose Trafficking, and Mitochondrial Function in Rat.

Relaxin-2 is a hormone with robust beneficial effects on the heart and blood vessels and potential as a therapy for cardiovascular (CV) disease. Considering the interorgan communication between skeletal muscle and heart, and the relation between muscle quality/composition and CV events, we hypothesize that relaxin-2 may regulate skeletal muscle physiology and metabolism. We aim to evaluate the impact of relaxin-2 on the proteome of skeletal muscle from healthy Sprague-Dawley rats. Animals were treated with 0.4 mg/kg/day of serelaxin (recombinant form of human relaxin-2) or vehicle (PBS) for 2 weeks employing subcutaneous osmotic minipumps. Skeletal muscle protein identification and quantification were performed by LC-MS/MS using a Data-Independent Acquisition (DIA)-Sequential Window Acquisition of All Theoretical Fragment Ion Spectra (SWATH) method. SWATH/MS quantitative analysis identified that relaxin-2 significantly decreased 95 proteins and significantly increased 32 proteins in rat skeletal muscle when compared to control rats. From these, 34 proteins were associated with muscle function, myogenesis, muscle differentiation and/or regeneration, 20 are mitochondrial proteins (six from the complexes of the electron transport chain), and 10 proteins participate in glucose metabolism. Qualitative data-dependent workflow analysis identified 35 proteins exclusive to the skeletal muscle of the relaxin-2-treated group: eight proteins related to processes of skeletal muscle function (size, ion homeostasis or organization of caveolae structures and cytoskeleton) and myogenesis, and two proteins involved in muscle differentiation. Our work highlighted for the first time the role of relaxin-2 in crucial processes of muscle physiology and energetic metabolism, which could influence several processes involved in myopathy and CV.

Animals↗