PubMed Health⌕ Search

Biomedical subjects

Jeffrey Skolnick

Publications and source records attributed to Jeffrey Skolnick.

At least 19 recordsLinked to original sources

TM-align: a protein structure alignment algorithm based on the TM-score.

We have developed TM-align, a new algorithm to identify the best structural alignment between protein pairs that combines the TM-score rotation matrix and Dynamic Programming (DP). The algorithm is approximately 4 times faster than CE and 20 times faster than DALI and SAL. On average, the resulting structure alignments have higher accuracy and coverage than those provided by these most often-used methods. TM-align is applied to an all-against-all structure comparison of 10 515 representative protein chains from the Protein Data Bank (PDB) with a sequence identity cutoff <95%: 1996 distinct folds are found when a TM-score threshold of 0.5 is used. We also use TM-align to match the models predicted by TASSER for solved non-homologous proteins in PDB. For both folded and misfolded models, TM-align can almost always find close structural analogs, with an average root mean square deviation, RMSD, of 3 A and 87% alignment coverage. Nevertheless, there exists a significant correlation between the correctness of the predicted structure and the structural similarity of the model to the other proteins in the PDB. This correlation could be used to assist in model selection in blind protein structure predictions. The TM-align program is freely downloadable at http://bioinformatics.buffalo.edu/TM-align.

Algorithms↗

The protein structure prediction problem could be solved using the current PDB library.

For single-domain proteins, we examine the completeness of the structures in the current Protein Data Bank (PDB) library for use in full-length model construction of unknown sequences. To address this issue, we employ a comprehensive benchmark set of 1,489 medium-size proteins that cover the PDB at the level of 35% sequence identity and identify templates by structure alignment. With homologous proteins excluded, we can always find similar folds to native with an average rms deviation (RMSD) from native of 2.5 A with approximately 82% alignment coverage. These template structures often contain a significant number of insertions/deletions. The tasser algorithm was applied to build full-length models, where continuous fragments are excised from the top-scoring templates and reassembled under the guide of an optimized force field, which includes consensus restraints taken from the templates and knowledge-based statistical potentials. For almost all targets (except for 2/1,489), the resultant full-length models have an RMSD to native below 6 A (97% of them below 4 A). On average, the RMSD of full-length models is 2.25 A, with aligned regions improved from 2.5 A to 1.88 A, comparable with the accuracy of low-resolution experimental structures. Furthermore, starting from state-of-the-art structural alignments, we demonstrate a methodology that can consistently bring template-based alignments closer to native. These results are highly suggestive that the protein-folding problem can in principle be solved based on the current PDB library by developing efficient fold recognition algorithms that can recover such initial alignments.

Algorithms↗

A scoring function for docking ligands to low-resolution protein structures.

We present a docking method that uses a scoring function for protein-ligand docking that is designed to maximize the docking success rate for low-resolution protein structures. We find that the resulting scoring function parameters are very different depending on whether they were optimized for high- or low-resolution protein structures. We show that this docking method can be successfully applied to predict the ligand-binding site of low-resolution structures. For a set of 25 protein-ligand complexes, in 76% of the cases, more than 50% of ligand-contacting residues are correctly predicted (using receptor crystal structures where the binding site is unspecified). Using decoys of the receptor structures having a 4 A RMSD from the native structure, for the same set of complexes, in 72% of the cases, we obtain at least one correctly predicted ligand-contacting residue. Furthermore, using an 81-protein-ligand set described by Jain, in 76 (93.8%) cases, the algorithm correctly predicts more than 50% of the ligand-contacting residues when native protein structures are used. Using 3 A RMSD from native decoys, in all but two cases (97.5%), the algorithm predicts at least one ligand-binding residue correctly. Finally, compared to the previously published Dolores method, for 298 protein-ligand pairs, the number of cases in which at least half of the specific contacts are correctly predicted is more than four times greater.

Algorithms↗

TASSER: an automated method for the prediction of protein tertiary structures in CASP6.

The recently developed TASSER (Threading/ASSembly/Refinement) method is applied to predict the tertiary structures of all CASP6 targets. TASSER is a hierarchical approach that consists of template identification by the threading program PROSPECTOR_3, followed by tertiary structure assembly via rearranging continuous template fragments. Assembly occurs using parallel hyperbolic Monte Carlo sampling under the guide of an optimized, reduced force field that includes knowledge-based statistical potentials and spatial restraints extracted from threading alignments. Models are automatically selected from the Monte Carlo trajectories in the low-temperature replicas using the clustering program SPICKER. For all 90 CASP targets/domains, PROSPECTOR_3 generates initial alignments with an average root-mean-square deviation (RMSD) to native of 8.4 A with 79% coverage. After TASSER reassembly, the average RMSD decreases to 5.4 A over the same aligned residues; the overall cumulative TM-score increases from 39.44 to 52.53. Despite significant improvements over the PROSPECTOR_3 template alignment observed in all target categories, the overall quality of the final models is essentially dictated by the quality of threading templates: The average TM-scores of TASSER models in the three categories are, respectively, 0.79 [comparative modeling (CM), 43 targets/domains], 0.47 [fold recognition (FR), 37 targets/domains], and 0.30 [new fold (NF), 10 targets/domains]. This highlights the need to develop novel (or improved) approaches to identify very distant targets as well as better NF algorithms.

Algorithms↗

Prediction of physical protein-protein interactions.

Many essential cellular processes such as signal transduction, transport, cellular motion and most regulatory mechanisms are mediated by protein-protein interactions. In recent years, new experimental techniques have been developed to discover the protein-protein interaction networks of several organisms. However, the accuracy and coverage of these techniques have proven to be limited, and computational approaches remain essential both to assist in the design and validation of experimental studies and for the prediction of interaction partners and detailed structures of protein complexes. Here, we provide a critical overview of existing structure-independent and structure-based computational methods. Although these techniques have significantly advanced in the past few years, we find that most of them are still in their infancy. We also provide an overview of experimental techniques for the detection of protein-protein interactions. Although the developments are promising, false positive and false negative results are common, and reliable detection is possible only by taking a consensus of different experimental approaches. The shortcomings of experimental techniques affect both the further development and the fair evaluation of computational prediction methods. For an adequate comparative evaluation of prediction and high-throughput experimental methods, an appropriately large benchmark set of biophysically characterized protein complexes would be needed, but is sorely lacking.

Algorithms↗

Scoring function for automated assessment of protein structure template quality.

We have developed a new scoring function, the template modeling score (TM-score), to assess the quality of protein structure templates and predicted full-length models by extending the approaches used in Global Distance Test (GDT)1 and MaxSub.2 First, a protein size-dependent scale is exploited to eliminate the inherent protein size dependence of the previous scores and appropriately account for random protein structure pairs. Second, rather than setting specific distance cutoffs and calculating only the fractions with errors below the cutoff, all residue pairs in alignment/modeling are evaluated in the proposed score. For comparison of various scoring functions, we have constructed a large-scale benchmark set of structure templates for 1489 small to medium size proteins using the threading program PROSPECTOR_3 and built the full-length models using MODELLER and TASSER. The TM-score of the initial threading alignments, compared to the GDT and MaxSub scoring functions, shows a much stronger correlation to the quality of the final full-length models. The TM-score is further exploited as an assessment of all 'new fold' targets in the recent CASP5 experiment and shows a close coincidence with the results of human-expert visual assessment. These data suggest that the TM-score is a useful complement to the fully automated assessment of protein structure predictions. The executable program of TM-score is freely downloadable at http://bioinformatics.buffalo.edu/TM-score.

Automation↗

EFICAz: a comprehensive approach for accurate genome-scale enzyme function inference.

EFICAz (Enzyme Function Inference by Combined Approach) is an automatic engine for large-scale enzyme function inference that combines predictions from four different methods developed and optimized to achieve high prediction accuracy: (i) recognition of functionally discriminating residues (FDRs) in enzyme families obtained by a Conservation-controlled HMM Iterative procedure for Enzyme Family classification (CHIEFc), (ii) pairwise sequence comparison using a family specific Sequence Identity Threshold, (iii) recognition of FDRs in Multiple Pfam enzyme families, and (iv) recognition of multiple Prosite patterns of high specificity. For FDR (i.e. conserved positions in an enzyme family that discriminate between true and false members of the family) identification, we have developed an Evolutionary Footprinting method that uses evolutionary information from homofunctional and heterofunctional multiple sequence alignments associated with an enzyme family. The FDRs show a significant correlation with annotated active site residues. In a jackknife test, EFICAz shows high accuracy (92%) and sensitivity (82%) for predicting four EC digits in testing sequences that are <40% identical to any member of the corresponding training set. Applied to Escherichia coli genome, EFICAz assigns more detailed enzymatic function than KEGG, and generates numerous novel predictions.

Amino Acid Sequence↗

Local propensities and statistical potentials of backbone dihedral angles in proteins.

The following three issues concerning the backbone dihedral angles of protein structures are presented. (1) How do the dihedral angles of the 20 amino acids depend on the identity and conformation of their nearest residues? (2) To what extent are the native dihedral angles determined by local (dihedral) potentials? (3) How to build a knowledge-based potential for a residue's dihedral angles, considering the identity and conformation of its nearest residues? We find that the dihedral angle distribution for a residue can significantly depend on the identity and conformation of its adjacent residues. These correlations are in sharp contrast to the Flory isolated-pair hypothesis. Statistical potentials are built for all combinations of residue triplets and depend on the dihedral angles between consecutive residues. First, a low-resolution potential is obtained, which only differentiates between the main populated basins in the dihedral angle density plots. Minimization of the dihedral potential for 125 test proteins reveals that most native alpha-helical residues (89%) and a large fraction of native beta-sheet residues (47%) adopt conformations close to their native one. For native loop residues, the percentage is 48%. It is also found that this fraction is higher for residues away from the ends of alpha or beta secondary structure elements. In addition, a higher resolution potential is built as a function of dihedral angles by a smoothing procedure and continuous functions interpolations. Monte Carlo energy minimization with this potential results in a lower fraction for native beta-sheet residues. Nevertheless, because of the higher flexibility and entropy of beta structures, they could be preferred under the influence of non-local interactions. In general, most alpha-helices and many beta-sheets are strongly determined by the local potential, while the conformations in loops and near the end of beta-sheets are more influenced by non-local interactions.

Amino Acids↗

Development and large scale benchmark testing of the PROSPECTOR_3 threading algorithm.

This article describes the PROSPECTOR_3 threading algorithm, which combines various scoring functions designed to match structurally related target/template pairs. Each variant described was found to have a Z-score above which most identified templates have good structural (threading) alignments, Z(struct) (Z(good)). 'Easy' targets with accurate threading alignments are identified as single templates with Z > Z(good) or two templates, each with Z > Z(struct), having a good consensus structure in mutually aligned regions. 'Medium' targets have a pair of templates lacking a consensus structure, or a single template for which Z(struct) < Z < Z(good). PROSPECTOR_3 was applied to a comprehensive Protein Data Bank (PDB) benchmark composed of 1491 single domain proteins, 41-200 residues long and no more than 30% identical to any threading template. Of the proteins, 878 were found to be easy targets, with 761 having a root mean square deviation (RMSD) from native of less than 6.5 A. The average contact prediction accuracy was 46%, and on average 17.6 residue continuous fragments were predicted with RMSD values of 2.0 A. There were 606 medium targets identified, 87% (31%) of which had good structural (threading) alignments. On average, 9.1 residue, continuous fragments with RMSD of 2.5 A were predicted. Combining easy and medium sets, 63% (91%) of the targets had good threading (structural) alignments compared to native; the average target/template sequence identity was 22%. Only nine targets lacked matched templates. Moreover, PROSPECTOR_3 consistently outperforms PSIBLAST. Similar results were predicted for open reading frames (ORFS) < or =200 residues in the M. genitalium, E. coli and S. cerevisiae genomes. Thus, progress has been made in identification of weakly homologous/analogous proteins, with very high alignment coverage, both in a comprehensive PDB benchmark as well as in genomes.

Algorithms↗

Automated structure prediction of weakly homologous proteins on a genomic scale.

We have developed TASSER, a hierarchical approach to protein structure prediction that consists of template identification by threading, followed by tertiary structure assembly via the rearrangement of continuous template fragments guided by an optimized C(alpha) and side-chain-based potential driven by threading-based, predicted tertiary restraints. TASSER was applied to a comprehensive benchmark set of 1,489 medium-sized proteins in the Protein Data Bank. With homologues excluded, in 927 cases, the templates identified by our threading algorithm PROSPECTOR_3 have a rms deviation from native <6.5 A with approximately 80% alignment coverage. After template reassembly, this number increases to 1,172. This shows significant and systematic improvement of the final models with respect to the initial template alignments. Furthermore, significant improvements in loop modeling are demonstrated. We then apply TASSER to the 1,360 medium-sized ORFs in the Escherichia coli genome; approximately 920 can be predicted with high accuracy based on confidence criteria established in the Protein Data Bank benchmark. These results from our unprecedented comprehensive folding benchmark on all protein categories provide a reliable basis for the application of TASSER to structural genomics, especially to proteins of low sequence identity to solved protein structures.

Algorithms↗

Microbial genomes have over 72% structure assignment by the threading algorithm PROSPECTOR_Q.

The genome scale threading of five complete microbial genomes is revisited using our state-of-the-art threading algorithm, PROSPECTOR_Q. Considering that structure assignment to an ORF could be useful for predicting biochemical function as well as for analyzing pathways, it is important to assess the current status of genome scale threading. The fraction of ORFs to which we could assign protein structures with a reasonably good confidence level to each genome sequences is over 72%, which is significantly higher than earlier studies. Using the assigned structures, we have predicted the function of several ORFs through "single-function" template structures, obtained from an analysis of the relationship between protein fold and function. The fold distribution of the genomes and the effect of the number of homologous sequences on structure assignment are also discussed.

Algorithms↗

SPICKER: a clustering approach to identify near-native protein folds.

We have developed SPICKER, a simple and efficient strategy to identify near-native folds by clustering protein structures generated during computer simulations. In general, the most populated clusters tend to be closer to the native conformation than the lowest energy structures. To assess the generality of the approach, we applied SPICKER to 1489 representative benchmark proteins </=200 residues that cover the PDB at the level of 35% sequence identity; each contains up to 280,000 structure decoys generated using the recently developed TASSER (Threading ASSembly Refinement) algorithm. The best of the top five identified folds has a root-mean-square deviation from native (RMSD) in the top 1.4% of all decoys. For 78% of the proteins, the difference in RMSD from native to the identified models and RMSD from native to the absolutely best individual decoy is below 1 A; the majority belong to the targets with converged conformational distributions. Although native fold identification from divergent decoy structures remains a challenge, our overall results show significant improvement over our previous clustering algorithms.

Algorithms↗

Large-scale assessment of the utility of low-resolution protein structures for biochemical function assignment.

MOTIVATION: Several protein function prediction methods employ structural features captured in three-dimensional (3D) descriptors of biologically relevant sites. These methods are successful when applied to high-resolution structures, but their detection ability in lower resolution predicted structures has only been tested for a few cases. RESULTS: A method that automatically generates a library of 3D functional descriptors for the structure-based prediction of enzyme active sites (automated functional templates, 593 in total for 162 different enzymes), based on functional and structural information automatically extracted from public databases, has been developed and evaluated using decoy structures. The applicability to predicted structures was investigated by analyzing decoys of varying quality, derived from enzyme native structures. For 35% of decoy structures, our method identifies the active site in models having 3-4 A coordinate root mean square deviation from the native structure, a quality that is reachable using state of the art protein structure prediction algorithms. AVAILABILITY: See http://www.bioinformatics.buffalo.edu/resources/aft/

Algorithms↗

Application of sparse NMR restraints to large-scale protein structure prediction.

The protein structure prediction algorithm TOUCHSTONEX that uses sparse distance restraints derived from NMR nuclear Overhauser enhancement (NOE) data to predict protein structures at low-to-medium resolution was evaluated as follows: First, a representative benchmark set of the Protein Data Bank library consisting of 1365 proteins up to 200 residues was employed. Using N/8 simulated long-range restraints, where N is the number of residues, 1023 (75%) proteins were folded to a C(alpha) root-mean-square deviation (RMSD) from native <6.5 A in one of the top five models. The average RMSD of the models for all 1365 proteins is 5.0 A. Using N/4 simulated restraints, 1206 (88%) proteins were folded to a RMSD <6.5 A and the average RMSD improved to 4.1 A. Then, 69 proteins with experimental NMR data were used. Using long-range NOE-derived restraints, 47 proteins were folded to a RMSD <6.5 A with N/8 restraints and 61 proteins were folded to a RMSD <6.5 A with N/4 restraints. Thus, TOUCHSTONEX can be a tool for NMR-based rapid structure determination, as well as used in other experimental methods that can provide tertiary restraint information.

Algorithms↗

Tertiary structure predictions on a comprehensive benchmark of medium to large size proteins.

We evaluate tertiary structure predictions on medium to large size proteins by TASSER, a new algorithm that assembles protein structures through rearranging the rigid fragments from threading templates guided by a reduced Calpha and side-chain based potential consistent with threading based tertiary restraints. Predictions were generated for 745 proteins 201-300 residues in length that cover the Protein Data Bank (PDB) at the level of 35% sequence identity. With homologous proteins excluded, in 365 cases, the templates identified by our threading program, PROSPECTOR_3, have a root-mean-square deviation (RMSD) to native < 6.5 angstroms, with >70% alignment coverage. After TASSER assembly, in 408 cases the best of the top five full-length models has a RMSD < 6.5 angstroms. Among the 745 targets are 18 membrane proteins, with one-third having a predicted RMSD < 5.5 A. For all representative proteins less than or equal to 300 residues that have corresponding multiple NMR structures in the Protein Data Bank, approximately 20% of the models generated by TASSER are closer to the NMR structure centroid than the farthest individual NMR model. These results suggest that reasonable structure predictions for nonhomologous large size proteins can be automatically generated on a proteomic scale, and the application of this approach to structural as well as functional genomics represent promising applications of TASSER.

Algorithms↗

The PDB is a covering set of small protein structures.

Structure comparisons of all representative proteins have been done. Employing the relative root mean square deviation (RMSD) from native enables the assessment of the statistical significance of structure alignments of different lengths in terms of a Z-score. Two conclusions emerge: first, proteins with their native fold can be distinguished by their Z-score. Second and somewhat surprising, all small proteins up to 100 residues in length have significant structure alignments to other proteins in a different secondary structure and fold class; i.e. 24.0% of them have 60% coverage by a template protein with a RMSD below 3.5A and 6.0% have 70% coverage. If the restriction that we align proteins only having different secondary structure types is removed, then in a representative benchmark set of proteins of 200 residues or smaller, 93% can be aligned to a single template structure (with average sequence identity of 9.8%), with a RMSD less than 4A, and 79% average coverage. In this sense, the current Protein Data Bank (PDB) is almost a covering set of small protein structures. The length of the aligned region (relative to the whole protein length) does not differ among the top hit proteins, indicating that protein structure space is highly dense. For larger proteins, non-related proteins can cover a significant portion of the structure. Moreover, these top hit proteins are aligned to different parts of the target protein, so that almost the entire molecule can be covered when combined. The number of proteins required to cover a target protein is very small, e.g. the top ten hit proteins can give 90% coverage below a RMSD of 3.5A for proteins up to 320 residues long. These results give a new view of the nature of protein structure space, and its implications for protein structure prediction are discussed.

Algorithms↗

TOUCHSTONEX: protein structure prediction with sparse NMR data.

TOUCHSTONEX, a new method for folding proteins that uses a small number of long-range contact restraints derived from NMR experimental NOE (nuclear Overhauser enhancement) data, is described. The method employs a new lattice-based, reduced model of proteins that explicitly represents C(alpha), C(beta), and the sidechain centers of mass. The force field consists of knowledge-based terms to produce protein-like behavior, including various short-range interactions, hydrogen bonding, and one-body, pairwise, and multibody long-range interactions. Contact restraints were incorporated into the force field as an NOE-specific pairwise potential. We evaluated the algorithm using a set of 125 proteins of various secondary structure types and lengths up to 174 residues. Using N/8 simulated, long-range sidechain contact restraints, where N is the number of residues, 108 proteins were folded to a C(alpha)-root-mean-square deviation (RMSD) from native below 6.5 A. The average RMSD of the lowest RMSD structures for all 125 proteins (folded and unfolded) was 4.4 A. The algorithm was also applied to limited experimental NOE data generated for three proteins. Using very few experimental sidechain contact restraints, and a small number of sidechain-main chain and main chain-main chain contact restraints, we folded all three proteins to low-to-medium resolution structures. The algorithm can be applied to the NMR structure determination process or other experimental methods that can provide tertiary restraint information, especially in the early stage of structure determination, when only limited data are available.

Algorithms↗