PubMed · 8744776
Using substitution probabilities to improve position-specific scoring matrices.
Abstract
Each column of amino acids in a multiple alignment of protein sequences can be represented as a vector of 20 amino acid counts. For alignment and searching applications, the count vector is an imperfect representation of a position, because the observed sequences are an incomplete sample of the full set of related sequences. One general solution to this problem is to model unobserved sequences by adding artificial 'pseudo-counts' to the observed counts. We introduce a simple method for computing pseudo-counts that combines the diversity observed in each alignment position with amino acid substitution probabilities. In extensive empirical tests, this position-based method out-performed other pseudo-count methods and was a substantial improvement over the traditional average score method used for constructing profiles.
Explore related subjects
Keep this discovery
Explore connections, maps & timelines
J G Henikoff, S Henikoff. 1996. Using substitution probabilities to improve position-specific scoring matrices.. https://doi.org/10.1093/bioinformatics%2F12.2.135
Cite the original work for its findings. Save a collection to share your selection of sources.