PubMed Health⌕ Search

Biomedical subjects

H Mamitsuka

Publications and source records attributed to H Mamitsuka.

5 recordsLinked to original sources

Predicting peptides that bind to MHC molecules using supervised learning of hidden Markov models.

The binding of a major histocompatibility complex (MHC) molecule to a peptide originating in an antigen is essential to recognizing antigens in immune systems, and it has proved to be important to use computers to predict the peptides that will bind to an MHC molecule. The purpose of this paper is twofold: First, we propose to apply supervised learning of hidden Markov models (HMMs) to this problem, which can surpass existing methods for the problem of predicting MHC-binding peptides. Second, we generate peptides that have high probabilities to bind to a certain MHC molecule, based on our proposed method using peptides binding to MHC molecules as a set of training data. From our experiments, in a type of cross-validation test, the discrimination accuracy of our supervised learning method is usually approximately 2-15% better than those of other methods, including backpropagation neural networks, which have been regarded as the most effective approach to this problem. Furthermore, using an HMM trained for HLA-A2, we present new peptide sequences that are provided with high binding probabilities by the HMM and that are thus expected to bind to HLA-A2 proteins. Peptide sequences not shown in this paper but with rather high binding probabilities can be obtained from the author.

Algorithms↗

A learning method of hidden Markov models for sequence discrimination.

We propose a learning method for hidden Markov models (HMM) for sequence discrimination. When given an HMM, our method sets a function that corresponds to the product of a difference between the observed and the desired likelihoods for each training sequence, and using a gradient descent algorithm, trains the HMM parameters so that the function should be minimized. This method allows us to use not only the examples belonging to a class that should be represented by the HMM, but also the examples not belonging to the class, i.e., negative examples. We evaluated our method in a series of experiments based on a type of cross-validation, and compared the results with those of two existing methods. Experimental results show that our method greatly reduces the discrimination errors made by the other two methods. We conclude that both the use of negative examples and our method of using negative examples are useful for training HMMs in discriminating unknown sequences.

Algorithms↗

alpha-Helix region prediction with stochastic rule learning.

We propose a new method, based on the theory of stochastic rule learning, for predicting alpha-helix regions in a given protein sequence. Our method (hereafter referred to as the SR method) produces stochastic rules, each of which assigns, to any region in an amino acid sequence, the probability that it is an alpha-helix region. When learning a stochastic rule from a particular alpha-helix region, our method makes use of positive training examples obtained from a number of regions that are homologous to that region. Each stochastic rule is optimized using the minimum description length (MDL) principle, and such optimized stochastic rules are used to predict alpha-helix regions of any given protein sequence. In our experiments, using 25 proteins selected from the HSSP database as training examples, we applied the SR method to the problem of predicting alpha-helix regions in test examples, which consisted of > 5000 residues with 38% alpha-helix content. Each of these test examples possesses < 25% homology to any proteins in the training and other test examples. Our method achieved 81% average prediction accuracy for the test examples; this compares favorably to Qian and Sejnowski's method, which attains no more than 75% average accuracy, and further which compares to Rost and Sander's method which has proven to be one of the best secondary structure prediction methods.

Amino Acid Sequence↗

Representing inter-residue dependencies in protein sequences with probabilistic networks.

A new method for representing a local region of a protein sequence as a probabilistic network is proposed. The method produces, from a large number of examples of a local region, a network which describes dependencies that exist among amino acid residues in the region. The network is constructed using the greedy-search algorithm based on the minimum description length (MDL) principle. In our experiments, we construct a probabilistic network of the EF-hand motif domain in calcium-binding proteins. Experimental results show that our method provides a visual aid to understanding the inter-residue dependencies of those regions using a probabilistic network, and the network captures several important features which are peculiar to the motif.

Algorithms↗

Predicting location and structure of beta-sheet regions using stochastic tree grammars.

We describe and demonstrate the effectiveness of a method of predicting protein secondary structures, beta-sheet regions in particular, using a class of stochastic tree grammars as representational language for their amino acid sequence patterns. The family of stochastic tree grammars we use, the Stochastic Ranked Node Rewriting Grammars (SRNRG), is one of the rare families of stochastic grammars that are expressive enough to capture the kind of long-distance dependencies exhibited by the sequences of beta-sheet regions, and at the same time enjoy relatively efficient processing. We applied our method on real data obtained from the HSSP database and the results obtained are encouraging: Using an SRNRG trained by data of a particular protein, our method was actually able to predict the location and structure of beta-sheet regions in a number of different proteins, whose sequences are less than 25 per cent homologous to the training sequences. The learning algorithm we use is an extension of the 'Inside-Outside' algorithm for stochastic context free grammars, but with a number of significant modifications. First, we restricted the grammars used to be members of the 'linear' subclass of SRNRG, and devised simpler and faster algorithms for this subclass. Secondly, we reduced the alphabet size (i.e. the number of amino acids) by clustering them using their physicochemical properties, gradually through the iterations of the learning algorithm. Finally, we parallelized our parsing algorithm to run on a highly parallel computer, a 32-processor CM-5, and were able to obtain a nearly linear speed-up.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms↗