PubMed HealthSearch

Biomedical subjects

M Kanehisa

Publications and source records attributed to M Kanehisa.

At least 19 recordsLinked to original sources

An assessment of neural network and statistical approaches for prediction of E. coli promoter sites.

We have constructed a perceptron type neural network for E. coli promoter prediction and improved its ability to generalize with a new technique for selecting the sequence features shown during training. We have also reconstructed five previous prediction methods and compared the effectiveness of those methods and our neural network. Surprisingly, the simple statistical method of Mulligan et al. performed the best amongst the previous methods. Our neural network was comparable to Mulligan's method when false positives were kept low and better than Mulligan's method when false negatives were kept low. We also showed the correlation between the prediction rates of neural networks achieved by previous researchers and the information content of their data sets.

Base Sequence

A knowledge base for predicting protein localization sites in eukaryotic cells.

To automate examination of massive amounts of sequence data for biological function, it is important to computerize interpretation based on empirical knowledge of sequence-function relationships. For this purpose, we have been constructing a knowledge base by organizing various experimental and computational observations as a collection of if-then rules. Here we report an expert system, which utilizes this knowledge base, for predicting localization sites of proteins only from the information on the amino acid sequence and the source origin. We collected data for 401 eukaryotic proteins with known localization sites (subcellular and extracellular) and divided them into training data and testing data. Fourteen localization sites were distinguished for animal cells and 17 for plant cells. When sorting signals were not well characterized experimentally, various sequence features were computationally derived from the training data. It was found that 66% of the training data and 59% of the testing data were correctly predicted by our expert system. This artificial intelligence approach is powerful and flexible enough to be used in genome analyses.

Algorithms

Construction of a dictionary of sequence motifs that characterize groups of related proteins.

An automatic procedure is proposed to identify, from the protein sequence database, conserved amino acid patterns (or sequence motifs) that are exclusive to a group of functionally related proteins. This procedure is applied to the PIR database and a dictionary of sequence motifs that relate to specific superfamilies constructed. The motifs have a practical relevance in identifying the membership of specific superfamilies without the need to perform sequence database searches in 20% of newly determined sequences. The sequence motifs identified represent functionally important sites on protein molecules. When multiple blocks exist in a single motif they are often close together in the 3-D structure. Furthermore, occasionally these motif blocks were found to be split by introns when the correlation with exon structures was examined.

Amino Acid Sequence

Expert system for predicting protein localization sites in gram-negative bacteria.

We have developed an expert system that makes use of various kinds of knowledge organized as "if-then" rules for predicting protein localization sites in Gram-negative bacteria, given the amino acid sequence information alone. We considered four localization sites: the cytoplasm, the inner (cytoplasmic) membrane, the periplasm, and the outer membrane. Most rules were derived from experimental observations. For example, the rule to recognize an inner membrane protein is the presence of either a hydrophobic stretch in the predicted mature protein or an uncleavable N-terminal signal sequence. Lipoproteins are first recognized by a consensus pattern and then assumed present at either the inner or outer membrane. These two possibilities are further discriminated by examining an acidic residue in the mature N-terminal portion. Furthermore, we found an empirical rule that periplasmic and outer membrane proteins were successfully discriminated by their different amino acid composition. Overall, our system could predict 83% of the localization sites of proteins in our database.

Bacterial Outer Membrane Proteins

Fragment peptide library for classification and functional prediction of proteins.

From protein sequence comparison data found in the literature, a library was organized using peptide fragment sequences which are common to related proteins. Each of the fragments was then examined for its occurrence in all the protein superfamilies defined by the NBRF-PIR data base. We have selected those fragment peptides that appear exclusively in one or a few superfamilies, and thus made a library of fragment peptides that characterize specific superfamilies. Such characteristic peptides are, in general, five to seven residues long and contain unusually high proportions of glycine and cysteine. This collection is a useful resource for the classification and functional prediction of protein molecules.

Amino Acid Sequence

Discriminant analysis of promoter regions in Escherichia coli sequences.

We have previously developed a general method based on the statistical technique of discriminant analysis to predict splice junctions in eukaryotic mRNA sequences [Nakata, K., Kanehisa, M. and DeLisi, C. (1985) Nucleic Acids Res., 13, 5327-5340]. In order to evaluate further applicability of this method, we now analyze the promoter region of Escherichia coli sequences. The attributes used for discrimination include the accuracy of consensus sequence patterns measured by the perceptron algorithm, the thermal stability map, the base composition and the Calladine-Dickerson rules for helical twist angle, roll angle, torsion angle and propeller twist angle. When applied to selected E. coli sequences in the GenBank database, the method correctly identifies 75% of the true promoter regions.

Algorithms

Prediction of in-vivo modification sites of proteins from their primary structures.

In order to make better use of the information contained in rapidly expanding amino acid sequence data, a new method to predict various modification sites of proteins from their primary structures is presented. It is also applicable to the prediction of other functional sites in proteins. Here we show the examples of N-glycosylation and serine/threonine phosphorylation sites. The method is essentially an elaboration of consensus sequence pattern matching based on stepwise discriminant analysis. The occurring amino acids near a potential modification site are represented by six numerical values which reflect various properties of amino acids. Longer-range effects around these sites are also considered. The stepwise procedure enabled us to automatically select effective features for discrimination. A computer program with our method first identifies potential modification sites by a sequence pattern, NX(S/T) for N-glycosylation or (S/T) for phosphorylation, and then decides by discriminant analysis whether a potential site is likely to be a true modification site. The prediction accuracy in the second step of discrimination was about 60% for glycosylation sites and about 80% for phosphorylation sites.

Amino Acid Sequence

A multivariate analysis method for discriminating protein secondary structural segments.

Using discriminant analysis, three types of protein secondary structure segments--helices, beta-strands and coils--are discriminated by amino acid sequence information alone. A variable in the discriminant analysis is defined by the amino acid index used to represent the sequence data and by the calculation method used to extract a feature in this representation. Thus, the three types of secondary structure segments derived from a set of non-homologous proteins from the Protein Data Bank are analyzed by 888 variables, which correspond to the mean, standard deviation, 3.6-residue periodicity and 2-residue periodicity for the numerical profiles determined from 222 published amino acid indices. These variables are combined to obtain best discrimination of the three types of segments. When up to three variables are combined, the best discrimination rate was 75%. The variables selected consist of the mean of alpha propensity (or turn propensity), the mean of beta propensity, and the 3.6-residue periodicity of hydrophobicity. This variable selection procedure can also be applied to other types of discrimination problem, once groups of sequence data are properly organized.

Amino Acid Sequence

Cluster analysis of amino acid indices for prediction of protein structure and function.

The relationship among 222 published indices representing various physicochemical and biochemical properties of amino acid residues has been investigated by hierarchical cluster analysis. The clustering result is illustrated by the minimum spanning tree, which is conveniently divided into four regions: alpha and turn propensities, beta propensity, hydrophobicity and other physicochemical properties including, among others, bulkiness of amino acid residues. In addition, several subclasses of hydrophobicity scales have been identified: preference of inside and outside, accessible surface area, surrounding hydrophobicity and other mostly experimental scales including transfer free energy, partition coefficients, HPLC parameters and polarity. Representative amino acid indices are identified in each of these groups. The collection of amino acid indices is a useful resource for empirical analyses correlating sequence information with structural and functional properties of proteins. As an example, the indices that best reproduce the amino acid mutation data matrix are searched against this collection.

Amino Acids

Detection of weak sequence homology of proteins for tertiary structure prediction.

Multiple measures of similarity were employed to detect weak homologies among protein sequences (e.g., below 30% residue identity). A set of thresholds was empirically determined, by using sample proteins of known structure, so as to select only correct pairs of sequences; correct or incorrect alignment of sequences was judged by direct comparison of corresponding conformations. The empirical criterion thus set up is applicable to the prediction of a protein structure when the structure of the other protein in the pair is known. We searched all the combinations between 84 proteins of known structure and 4610 proteins stored in a sequence database, and found about 4000 pairs of sequences which satisfied the criterion. However, after excluding such pairs of proteins that belong to the same family or superfamily, the number of pairs remaining was reduced to only 19. The reliability of these data for structural prediction is discussed.

Amino Acid Sequence

L-aspartate ammonia-lyase and fumarate hydratase share extensive sequence homology.

Based on our recent determinations of the nucleotide sequences of the L-aspartate ammonia-lyase genes from Escherichia coli and Pseudomonas fluorescens, primary structures of the two L-aspartate ammonia-lyases and fumarate hydratases from Bacillus subtilis and E. coli (N-terminal partial sequence) were compared by computer analysis. These four enzymes exhibited a significant homology of at least 37%, implying that L-aspartate ammonia-lyase and fumarate hydratase share a common evolutionary origin. To authors' knowledge, this feature appears to be the first example showing that two kinds of enzymes catalyzing different types of reactions, albeit similar, share such a high degree of sequence homology.

Amino Acid Sequence

Structure of the human interleukin-2 receptor gene.

The gene encoding the human interleukin-2 (IL-2) receptor consists of 8 exons spanning more than 25 kilobases on chromosome 10. Exons 2 and 4 were derived from a gene duplication event and unexpectedly also are homologous to the recognition domain of human complement factor B. Alternative messenger RNA (mRNA) splicing may delete exon 4 sequences, resulting in a mRNA that does not encode a functional IL-2 receptor. Leukemic T cells infected with HTLV-I and normal activated T cells express IL-2 receptors with identical deduced protein sequences. Receptor gene transcription is initiated at two principal sites in normal activated T cells. Adult T cell leukemia cells infected with HTLV-I show activity at both of these sites, but also at a third transcription initiation site.

Amino Acid Sequence

Prediction of splice junctions in mRNA sequences.

A general method based on the statistical technique of discriminant analysis is developed to distinguish boundaries of coding and non-coding regions in nucleic acid sequences. In particular, the method is applied to the prediction of splicing sites in messenger RNA precursors. Information used for discrimination includes consensus sequence patterns around splice junctions, free energy of snRNA and mRNA base pairing, and statistical differences between coding and non-coding regions such as periodic appearance of specific bases in coding regions reflecting the non-random usage of degenerate codons. Given the reading frame of an exon (but not the exon/intron boundaries), the method will predict the following exon, namely, the intron to be excised out. When applied to human sequences in the GenBank database, the method correctly identified 80% of true splice junctions.

Base Sequence

The detection and classification of membrane-spanning proteins.

Discriminant analysis can be used to precisely classify membrane proteins as integral or peripheral and to estimate the odds that the classification is correct. Specifically, using 102 membrane proteins from the National Biomedical Research Foundation (NBRF) database we find that discrimination between integral and peripheral membrane proteins can be achieved with 99% reliability. Hydrophobic segments of integral membrane proteins can also be distinguished from interior segments of globular soluble proteins with better than 95% reliability. We also propose a procedure for determining boundaries of membrane-spanning segments and apply it to several integral membrane proteins. For the limited data available (such as on transplantation antigens), the residues at the boundaries of a membrane-spanning segment are predictable to within the error inherent in the concept of boundary. As a specific indication of resolution, seven membrane-spanning segments of bacteriorhodopsin are resolved with no information other than sequence, and the predicted boundary residues agree with the experimental data on proteolytic cleavage sites. Several definitive but yet to be tested predictions are also made, and the relation to other predictive methods is briefly discussed. A computer program in FORTRAN for prediction of membrane-spanning segments is available from the authors.

Animals

The GenBank nucleic acid sequence database.

The GenBank nucleic acid sequence database is a computer-based collection of all published DNA and RNA sequences; it contains over five million bases in close to six thousand sequence entries drawn from four thousand five hundred published articles. Each sequence is accompanied by relevant biological annotation. The database is available either on magnetic tape, on floppy diskettes, on-line or in hardcopy form. We discuss the structure of the database, the extent of the data and the implications of the database for research on nucleic acids.

Base Sequence