PubMed Health⌕ Search

Biomedical subjects

Andrew R Dalby

Publications and source records attributed to Andrew R Dalby.

5 recordsLinked to original sources

COPASAAR--a database for proteomic analysis of single amino acid repeats.

BACKGROUND: Single amino acid repeats make up a significant proportion in all of the proteomes that have currently been determined. They have been shown to be functionally and medically significant, and are associated with cancers and neuro-degenerative diseases such as Huntington's Chorea, where a poly-glutamine repeat is responsible for causing the disease. The COPASAAR database is a new tool to facilitate the rapid analysis of single amino acid repeats at a proteome level. The database aims to simplify the comparison of repeat distributions between proteomes in order to provide a better understanding of their function and evolution. RESULTS: A comparative analysis of all proteomes in the database (currently 244) shows that single amino acid repeats account for about 12-14% of the proteome of any given species. They are more common in eukaryotes (14%) than in either archaea or bacteria (both 13%). Individual analyses of proteomes show that long single amino acid repeats (6+ residues) are much more common in the Eukaryotes and that longer repeats are usually made up of hydrophilic amino acids such as glutamine, glutamic acid, asparagine, aspartic acid and serine. CONCLUSION: COPASAAR is a useful tool for comparative proteomics that provides rapid access to amino acid repeat data that can be readily data-mined. The COPASAAR database can be queried at the kingdom, proteome or individual protein level. As the amount of available proteome data increases this will be increasingly important in order to automate proteome comparison. The insights gained from these studies will give a better insight into the evolution of protein sequence and function.

Algorithms↗

One gene, two diseases and three conformations: molecular dynamics simulations of mutants of human prion protein at room temperature and elevated temperatures.

Fatal familial insomnia (FFI) and Creutzfeldt-Jakob disease (CJD) are associated to the same mutation at codon 178 but differentiate into clinicopathologically distinct diseases determined by this mutation and a naturally occurring methionine-valine polymorphism at codon 129 of the prion protein gene. It has been suggested that the clinical and pathological difference between FFI and CJD is caused by different conformations of the prion protein. Using molecular dynamics (MD), we investigated the effect of the mutation at codon 178 and the polymorphism at codon 129 on prion protein dynamics and conformation at normal and elevated temperatures. Four model structures were examined with a focus on their dynamics and conformational changes. The results showed differences in stability and dynamics between polymorphic variants. Methionine variants demonstrated a higher stability than valine variants. Elongation of existing beta-sheets and formation of new beta-sheets was found to occur more readily in valine polymorphic variants. We also discovered the inhibitory effect of proline residue on existing beta-sheet elongation.

Computer Simulation↗

Mining HIV protease cleavage data using genetic programming with a sum-product function.

MOTIVATION: In order to design effective HIV inhibitors, studying and understanding the mechanism of HIV protease cleavage specification is critical. Various methods have been developed to explore the specificity of HIV protease cleavage activity. However, success in both extracting discriminant rules and maintaining high prediction accuracy is still challenging. The earlier study had employed genetic programming with a min-max scoring function to extract discriminant rules with success. However, the decision will finally be degenerated to one residue making further improvement of the prediction accuracy difficult. The challenge of revising the min-max scoring function so as to improve the prediction accuracy motivated this study. RESULTS: This paper has designed a new scoring function called a sum-product function for extracting HIV protease cleavage discriminant rules using genetic programming methods. The experiments show that the new scoring function is superior to the min-max scoring function. AVAILABILITY: The software package can be obtained by request to Dr Zheng Rong Yang.

Algorithms↗

Reduced bio basis function neural network for identification of protein phosphorylation sites: comparison with pattern recognition algorithms.

Protein phosphorylation is a post-translational modification performed by a group of enzymes known as the protein kinases or phosphotransferases (Enzyme Commission classification 2.7). It is essential to the correct functioning of both proteins and cells, being involved with enzyme control, cell signalling and apoptosis. The major problem when attempting prediction of these sites is the broad substrate specificity of the enzymes. This study employs back-propagation neural networks (BPNNs), the decision tree algorithm C4.5 and the reduced bio-basis function neural network (rBBFNN) to predict phosphorylation sites. The aim is to compare prediction efficiency of the three algorithms for this problem, and examine knowledge extraction capability. All three algorithms are effective for phosphorylation site prediction. Results indicate that rBBFNN is the fastest and most sensitive of the algorithms. BPNN has the highest area under the ROC curve and is therefore the most robust, and C4.5 has the highest prediction accuracy. C4.5 also reveals the amino acid 2 residues upstream from the phosporylation site is important for serine/threonine phosphorylation, whilst the amino acid 3 residues upstream is important for tyrosine phosphorylation.

Algorithms↗

Predicting the phosphorylation sites using hidden Markov models and machine learning methods.

Accurately predicting phosphorylation sites in proteins is an important issue in postgenomics, for which how to efficiently extract the most predictive features from amino acid sequences for modeling is still challenging. Although both the distributed encoding method and the bio-basis function method work well, they still have some limits in use. The distributed encoding method is unable to code the biological content in sequences efficiently, whereas the bio-basis function method is a nonparametric method, which is often computationally expensive. As hidden Markov models (HMMs) can be used to generate one model for one cluster of aligned protein sequences, the aim in this study is to use HMMs to extract features from amino acid sequences, where sequence clusters are determined using available biological knowledge. In this novel method, HMMs are first constructed using functional sequences only. Both functional and nonfunctional training sequences are then inputted into the trained HMMs to generate functional and nonfunctional feature vectors. From this, a machine learning algorithm is used to construct a classifier based on these feature vectors. It is found in this work that (1) this method provides much better prediction accuracy than the use of HMMs only for prediction, and (2) the support vector machines (SVMs) algorithm outperforms decision trees and neural network algorithms when they are constructed on the features extracted using the trained HMMs.

Algorithms↗