PubMed Health⌕ Search

Biomedical subjects

D Kihara

Publications and source records attributed to D Kihara.

8 recordsLinked to original sources

TOUCHSTONE: an ab initio protein structure prediction method that uses threading-based tertiary restraints.

The successful prediction of protein structure from amino acid sequence requires two features: an efficient conformational search algorithm and an energy function with a global minimum in the native state. As a step toward addressing both issues, a threading-based method of secondary and tertiary restraint prediction has been developed and applied to ab initio folding. Such restraints are derived by extracting consensus contacts and local secondary structure from at least weakly scoring structures that, in some cases, can lack any global similarity to the sequence of interest. Furthermore, to generate representative protein structures, a reduced lattice-based protein model is used with replica exchange Monte Carlo to explore conformational space. We report results on the application of this methodology, termed TOUCHSTONE, to 65 proteins whose lengths range from 39 to 146 residues. For 47 (40) proteins, a cluster centroid whose rms deviation from native is below 6.5 (5) A is found in one of the five lowest energy centroids. The number of correctly predicted proteins increases to 50 when atomic detail is added and a knowledge-based atomic potential is combined with clustered and nonclustered structures for candidate selection. The combination of the ratio of the relative number of contacts to the protein length and the number of clusters generated by the folding algorithm is a reliable indicator of the likelihood of successful fold prediction, thereby opening the way for genome-scale ab initio folding.

Algorithms↗

Generalized comparative modeling (GENECOMP): a combination of sequence comparison, threading, and lattice modeling for protein structure prediction and refinement.

An improved generalized comparative modeling method, GENECOMP, for the refinement of threading models is developed and validated on the Fischer database of 68 probe-template pairs, a standard benchmark used to evaluate threading approaches. The basic idea is to perform ab initio folding using a lattice protein model, SICHO, near the template provided by the new threading algorithm PROSPECTOR. PROSPECTOR also provides predicted contacts and secondary structure for the template-aligned regions, and possibly for the unaligned regions by garnering additional information from other top-scoring threaded structures. Since the lowest-energy structure generated by the simulations is not necessarily the best structure, we employed two structure-selection protocols: distance geometry and clustering. In general, clustering is found to generate somewhat better quality structures in 38 of 68 cases. When applied to the Fischer database, the protocol does no harm and in a significant number of cases improves upon the initial threading model, sometimes dramatically. The procedure is readily automated and can be implemented on a genomic scale.

Algorithms↗

Defrosting the frozen approximation: PROSPECTOR--a new approach to threading.

PROSPECTOR (PROtein Structure Predictor Employing Combined Threading to Optimize Results) is a new threading approach that uses sequence profiles to generate an initial probe-template alignment and then uses this "partly thawed" alignment in the evaluation of pair interactions. Two types of sequence profiles are used: the close set, composed of sequences in which sequence identity lies between 35% and 90%; and the distant set, composed of sequences with a FASTA E-score less than 10. Thus, a total of four scoring functions are used in a hierarchical method: the close (distant) sequence profiles screen a structural database to provide an initial alignment of the probe sequence in each of the templates. The same database is then screened with a scoring function composed of sequence plus secondary structure plus pair interaction profiles. This combined hierarchical threading method is called PROSPECTOR1. For the original Fischer database, 59 of 68 pairs are correctly identified in the top position. Next, the set of the top 20 scoring sequences (four scoring functions times the top five structures) is used to construct a protein-specific pair potential based on consensus side-chain contacts occurring in 25% of the structures. In subsequent threading iterations, this protein-specific pair potential, when combined in a composite manner, is found to be more sensitive in identifying the correct pairs than when the original statistical potential is used, and it increases the number of recognized structures for the combined scoring functions, termed PROSPECTOR2, to a total of 61 Fischer pairs identified in the top position. Application to a second, smaller Fischer database of 27 probe-template pairs places 18 (17) structures in the top position for PROSPECTOR1 (PROSPECTOR2). Overall, these studies show that the use of pair interactions as assessed by the improved Z-score enhances the specificity of probe-template matches. Thus, when the hierarchy of scoring functions is combined, the ability to identify correct probe-template pairs is significantly enhanced. Finally, a web server has been established for use by the academic community (http://bioinformatics.danforthcenter.org/services/threading.html).

Benchmarking↗

Ab initio protein structure prediction via a combination of threading, lattice folding, clustering, and structure refinement.

A combination of sequence comparison, threading, lattice, and off-lattice Monte Carlo (MC) simulations and clustering of MC trajectories was used to predict the structure of all (but one) targets of the CASP4 experiment on protein structure prediction. Although this method is automated and is operationally the same regardless of the level of uniqueness of the query proteins, here we focus on the more difficult targets at the border of the fold recognition and new fold categories. For a few targets (T0110 is probably the best example), the ab initio method produced more accurate models than models obtained by the fold recognition techniques. For the most difficult targets from the new fold categories, substantial fragments of structures have been correctly predicted. Possible improvements of the method are briefly discussed.

Cluster Analysis↗

Tandem clusters of membrane proteins in complete genome sequences.

The distribution of genes coding for membrane proteins was investigated in 16 complete genomes: 4 archaea, 11 bacteria, and 1 eukaryote. Membrane proteins were identified by our new method of predicting transmembrane segments () after the removal of amino-terminal signal peptides. Interestingly, about half of the membrane protein genes in each genome were found to be located next to another, forming tandem clusters. Roughly 10%-30% of the tandem clusters were conserved among organisms, and most of the conserved tandem clusters belonged to one of the three functional groups, namely, transporters, the electron transport system, and cell motility. A tandem cluster sometimes contained paralogous membrane proteins, in which case the cluster size and the number of transmembrane segments could be related to a functional category, especially to transporters. In addition to the clustering of membrane proteins, the clustering of membrane proteins and ATP-binding proteins in the complete genomes was also analyzed. Although this clustering was not statistically significant, it was useful to identify candidate membrane protein partners of isolated ATP-binding protein components in the ABC transporters. Possible implications of tandem cluster organization of membrane protein genes are discussed including the complex formation and other functional coupling of protein products and the mechanism of protein translocation to the cell membrane.

ATP-Binding Cassette Transporters↗

Prediction of membrane proteins based on classification of transmembrane segments.

The number of transmembrane segments often corresponds to a structural or functional class of membrane proteins such as to seven-transmembrane receptors and six-transmembrane ion channels. We have developed a new prediction method to detect the membrane protein class that is defined by the number of transmembrane segments, as well as to locate the transmembrane segments in the amino acid sequence. Each membrane protein class is represented by a model of ordering different types of transmembrane segments. Specifically, we have classified the transmembrane segments in known membrane proteins into five groups (types) using the Mahalanobis distance with the average hydrophobicity and the periodicity of hydrophobicity as a measure of similarity. The discriminant functions derived for these groups were then used to detect transmembrane segments and to match with the models for one- to fourteen-spanning membrane proteins and for globular proteins. Using the test data set of 89 membrane proteins whose transmembrane positions are known by experimental evidence, 61.8% of the proteins and 85.1% of the transmembrane segments were correctly predicted. Because of the new feature to predict membrane protein classes, the method should be useful in the functional assignment of genomic sequences.

Animals↗