PubMed Health⌕ Search

Biomedical subjects

Sandor Vajda

Publications and source records attributed to Sandor Vajda.

At least 19 recordsLinked to original sources

Classification of protein complexes based on docking difficulty.

Based on the results of several groups using different docking methods, the key properties that determine the expected success rate in protein-protein docking calculations are measures of conformational change, interface area, and hydrophobicity. A classification of protein complexes in terms of these measures provides a prediction of docking difficulty. This classification is used to study the targets of the CAPRI docking experiment. Results show that targets with a moderate expected difficulty were indeed predicted well by a number of groups, whereas the use of additional a priori information was necessary to obtain good results for some very difficult targets. The analysis indicates that CAPRI and other relatively large-scale docking studies represent very important steps toward understanding the capabilities and limitations of current protein-protein docking methods.

Algorithms↗

Performance of the first protein docking server ClusPro in CAPRI rounds 3-5.

To evaluate the current status of the protein-protein docking field, the CAPRI experiment came to life. Researchers are given the receptor and ligand 3-dimensional (3D) coordinates before the cocrystallized complex is published. Human predictions of the complex structure are supposed to be submitted within 3 weeks, whereas the server ClusPro has only 24 h and does not make use of any biochemical information. From the 10 targets analyzed in the second evaluation meeting of CAPRI, ClusPro was able to predict meaningful models for 5 targets using only empirical free energy estimates. For two of the targets, the server predictions were assessed to be among the best in the field. Namely, for Targets 8 and 12, ClusPro predicted the model with the most accurate binding-site interface and the model with the highest percentage of nativelike contacts, among 180 and 230 submissions, respectively. After CAPRI, the server has been further developed to predict oligomeric assemblies, and new tools now allow the user to restrict the search for the complex to specific regions on the protein surface, significantly enhancing the predictive capabilities of the server. The performance of ClusPro in CAPRI Rounds 3-5 suggests that clustering the low free energy (i.e., desolvation and electrostatic energy) conformations of a homogeneous conformational sampling of the binding interface is a fast and reliable procedure to detect protein-protein interactions and eliminate false positives. Not including targets that had a significant structural rearrangement upon binding, the success rate of ClusPro was found to be around 71%.

Algorithms↗

Optimal clustering for detecting near-native conformations in protein docking.

Clustering is one of the most powerful tools in computational biology. The conventional wisdom is that events that occur in clusters are probably not random. In protein docking, the underlying principle is that clustering occurs because long-range electrostatic and/or desolvation forces steer the proteins to a low free-energy attractor at the binding region. Something similar occurs in the docking of small molecules, although in this case shorter-range van der Waals forces play a more critical role. Based on the above, we have developed two different clustering strategies to predict docked conformations based on the clustering properties of a uniform sampling of low free-energy protein-protein and protein-small molecule complexes. We report on significant improvements in the automated prediction and discrimination of docked conformations by using the cluster size and consensus as a ranking criterion. We show that the success of clustering depends on identifying the appropriate clustering radius of the system. The clustering radius for protein-protein complexes is consistent with the range of the electrostatics and desolvation free energies (i.e., between 4 and 9 Angstroms); for protein-small molecule docking, the radius is set by van der Waals interactions (i.e., at approximately 2 Angstroms). Without any a priori information, a simple analysis of the histogram of distance separations between the set of docked conformations can evaluate the clustering properties of the data set. Clustering is observed when the histogram is bimodal. Data clustering is optimal if one chooses the clustering radius to be the minimum after the first peak of the bimodal distribution. We show that using this optimal radius further improves the discrimination of near-native complex structures.

Amino Acid Sequence↗

Exploring the binding site structure of the PPAR gamma ligand-binding domain by computational solvent mapping.

Solvent mapping moves molecular probes, small organic molecules containing various functional groups, around the protein surface, finds favorable positions, clusters the conformations, and ranks the clusters based on the average free energy. Using at least six different solvents as probes, the probes cluster in major pockets of the functional site, providing detailed and reliable information on the amino acid residues that are important for ligand binding. Solvent mapping was applied to 12 structures of the peroxisome proliferator activated receptor gamma (PPARgamma) ligand-binding domain (LBD), including 2 structures without a ligand, 2 structures with a partial agonist, and 8 structures with a PPAR agonist bound. The analysis revealed 10 binding "hot spots", 4 in the ligand-binding pocket, 2 in the coactivator-binding region, 1 in the dimerization domain, 2 around the ligand entrance site, and 1 minor site without a known function. Mapping is a major source of information on the role and cooperativity of these sites. It shows that large portions of the ligand-binding site are already formed in the PPARgamma apostructure, but an important pocket near the AF-2 transactivation domain becomes accessible only in structures that are cocrystallized with strong agonists. Conformational changes were seen in several other sites, including one involved in the stabilization of the LBD and two others at the region of the coactivator binding. The number of probe clusters retained by these sites depends on the properties of the bound agonist, providing information on the origin of correlations between ligand and coactivator binding.

Alkanesulfonates↗

PRECISE: a Database of Predicted and Consensus Interaction Sites in Enzymes.

PRECISE (Predicted and Consensus Interaction Sites in Enzymes) is a database of interactions between the amino acid residues of an enzyme and its ligands (substrate and transition state analogs, cofactors, inhibitors and products). It is available online at http://precise.bu.edu/. In the current version, all information on interactions is extracted from the enzyme-ligand complexes in the Protein Data Bank (PDB) by performing the following steps: (i) clustering homologous enzyme chains such that, in each cluster, the proteins have the same EC number and all sequences are similar; (ii) selecting a representative chain for each cluster; (iii) selecting ligand types; (iv) finding non-bonded interactions and hydrogen bonds; and (v) summing the interactions for all chains within the cluster. The output of the search is the color-coded sequence of the representative. The colors indicate the total number of interactions found at each amino acid position in all chains of the cluster. Clicking on a residue displays a detailed list of interactions for that residue. Optional filters allow restricting the output to selected chains in the cluster, to non-bonded or hydrogen bonding interactions, and to selected ligand types. The binding site information is essential for understanding and altering substrate specificity and for the design of enzyme inhibitors.

Amino Acid Sequence↗

Deletion of Ser-171 causes inactivation, proteasome-mediated degradation and complete deficiency of human transaldolase.

Homozygous deletion of three nucleotides coding for Ser-171 (S171) of TAL-H (human transaldolase) has been identified in a female patient with liver cirrhosis. Accumulation of sedoheptulose 7-phosphate raised the possibility of TAL (transaldolase) deficiency in this patient. In the present study, we show that the mutant TAL-H gene was effectively transcribed into mRNA, whereas no expression of the TALDeltaS171 protein or enzyme activity was detected in TALDeltaS171 fibroblasts or lymphoblasts. Unlike wild-type TAL-H-GST fusion protein (where GST stands for glutathione S-transferase), TALDeltaS171-GST was solubilized only in the presence of detergents, suggesting that deletion of Ser-171 caused conformational changes. Recombinant TALDeltaS171 had no enzymic activity. TALDeltaS171 was effectively translated in vitro using rabbit reticulocyte lysates, indicating that the absence of TAL-H protein in TALDeltaS171 fibroblasts and lymphoblasts may be attributed primarily to rapid degradation. Treatment with cell-permeable proteasome inhibitors led to the accumulation of TALDeltaS171 in whole cell lysates and cytosolic extracts of patient lymphoblasts, suggesting that deletion of Ser-171 led to rapid degradation by the proteasome. Although the TALDeltaS171 protein became readily detectable in proteasome inhibitor-treated cells, it displayed no appreciable enzymic activity. The results suggest that deletion of Ser-171 leads to inactivation and proteasome-mediated degradation of TAL-H. Since TAL-H is a regulator of apoptosis signal processing, complete deficiency of TAL-H may be relevant for the pathogenesis of liver cirrhosis.

Cells, Cultured↗

Anchor residues in protein-protein interactions.

We show that the mechanism for molecular recognition requires one of the interacting proteins, usually the smaller of the two, to anchor a specific side chain in a structurally constrained binding groove of the other protein, providing a steric constraint that helps to stabilize a native-like bound intermediate. We identify the anchor residues in 39 protein-protein complexes and verify that, even in the absence of their interacting partners, the anchor side chains are found in conformations similar to those observed in the bound complex. These ready-made recognition motifs correspond to surface side chains that bury the largest solvent-accessible surface area after forming the complex (> or =100 A2). The existence of such anchors implies that binding pathways can avoid kinetically costly structural rearrangements at the core of the binding interface, allowing for a relatively smooth recognition process. Once anchors are docked, an induced fit process further contributes to forming the final high-affinity complex. This later stage involves flexible (solvent-exposed) side chains that latch to the encounter complex in the periphery of the binding pocket. Our results suggest that the evolutionary conservation of anchor side chains applies to the actual structure that these residues assume before the encounter complex and not just to their loci. Implications for protein docking are also discussed.

Antigen-Antibody Complex↗

ClusPro: a fully automated algorithm for protein-protein docking.

ClusPro (http://nrc.bu.edu/cluster) represents the first fully automated, web-based program for the computational docking of protein structures. Users may upload the coordinate files of two protein structures through ClusPro's web interface, or enter the PDB codes of the respective structures, which ClusPro will then download from the PDB server (http://www.rcsb.org/pdb/). The docking algorithms evaluate billions of putative complexes, retaining a preset number with favorable surface complementarities. A filtering method is then applied to this set of structures, selecting those with good electrostatic and desolvation free energies for further clustering. The program output is a short list of putative complexes ranked according to their clustering properties, which is automatically sent back to the user via email.

Algorithms↗

Consensus alignment server for reliable comparative modeling with distant templates.

Consensus is a server developed to produce high-quality alignments for comparative modeling, and to identify the alignment regions reliable for copying from a given template. This is accomplished even when target-template sequence identity is as low as 5%. Combining the output from five different alignment methods, the server produces a consensus alignment, with a reliability measure indicated for each position and a prediction of the regions suitable for modeling. Models built using the server predictions are typically within 3 A rms deviations from the crystal structure. Users can upload a target protein sequence and specify a template (PDB code); if no template is given, the server will search for one. The method has been validated on a large set of homologous protein structure pairs. The Consensus server should prove useful for modelers for whom the structural reliability of the model is critical in their applications. It is currently available at http://structure.bu.edu/cgi-bin/consensus/consensus.cgi.

Algorithms↗

ClusPro: an automated docking and discrimination method for the prediction of protein complexes.

MOTIVATION: Predicting protein interactions is one of the most challenging problems in functional genomics. Given two proteins known to interact, current docking methods evaluate billions of docked conformations by simple scoring functions, and in addition to near-native structures yield many false positives, i.e. structures with good surface complementarity but far from the native. RESULTS: We have developed a fast algorithm for filtering docked conformations with good surface complementarity, and ranking them based on their clustering properties. The free energy filters select complexes with lowest desolvation and electrostatic energies. Clustering is then used to smooth the local minima and to select the ones with the broadest energy wells-a property associated with the free energy at the binding site. The robustness of the method was tested on sets of 2000 docked conformations generated for 48 pairs of interacting proteins. In 31 of these cases, the top 10 predictions include at least one near-native complex, with an average RMSD of 5 A from the native structure. The docking and discrimination method also provides good results for a number of complexes that were used as targets in the Critical Assessment of PRedictions of Interactions experiment. AVAILABILITY: The fully automated docking and discrimination server ClusPro can be found at http://structure.bu.edu

Algorithms↗

Protein-protein docking: is the glass half-full or half-empty?

Are current docking methods capable of building complexes from putative component protein structures? Results of recent computational studies, including those of the CAPRI (Critical Assessment of Protein Interactions) competition, were used to determine the key properties for successful docking and introduce a classification of protein complexes based on docking difficulty. Enzyme-inhibitor complexes could be determined with reasonable accuracy - possibly to within a few alternative structures. Results for antigen-antibody pairs are less predictable, and data for small signaling complexes are generally poor. However, moderate amounts of experimental data can remove uncertainty and the methodology is rapidly improving. Transient complexes with large interface areas undergo substantial conformational change and are beyond the reach of current docking methods. The docking of such complexes might therefore require fundamentally new approaches.

Algorithms↗

Combination of scoring functions improves discrimination in protein-protein docking.

Two structure-based potentials are used for both filtering (i.e., selecting a subset of conformations generated by rigid-body docking), and rescoring and ranking the selected conformations. ACP (atomic contact potential) is an atom-level extension of the Miyazawa-Jernigan potential parameterized on protein structures, whereas RPScore (residue pair potential score) is a residue-level potential, based on interactions in protein-protein complexes. These potentials are combined with other energy terms and applied to 13 sets of protein decoys, as well as to the results of docking 10 pairs of unbound proteins. For both potentials, the ability to discriminate between near-native and non-native docked structures is substantially improved by refining the structures and by adding a van der Waals energy term. It is observed that ACP and RPScore complement each other in a number of ways (e.g., although RPScore yields more hits than ACP, mainly as a result of its better performance for charged complexes, ACP usually ranks the near-native complexes better). As a general solution to the protein-docking problem, we have found that the best discrimination strategies combine either an RPScore filter with an ACP-based scoring function, or an ACP-based filter with an RPScore-based scoring function. Thus, ACP and RPScore capture complementary structural information, and combining them in a multistage postprocessing protocol provides substantially better discrimination than the use of the same potential for both filtering and ranking the docked conformations.

Algorithms↗

Identification of substrate binding sites in enzymes by computational solvent mapping.

Enzyme structures determined in organic solvents show that most organic molecules cluster in the active site, delineating the binding pocket. We have developed algorithms to perform solvent mapping computationally, rather than experimentally, by placing molecular probes (small molecules or functional groups) on a protein surface, and finding the regions with the most favorable binding free energy. The method then finds the consensus site that binds the highest number of different probes. The probe-protein interactions at this site are compared to the intermolecular interactions seen in the known complexes of the enzyme with various ligands (substrate analogs, products, and inhibitors). We have mapped thermolysin, for which experimental mapping results are also available, and six further enzymes that have no experimental mapping data, but whose binding sites are well characterized. With the exception of haloalkane dehalogenase, which binds very small substrates in a narrow channel, the consensus site found by the mapping is always a major subsite of the substrate-binding site. Furthermore, the probes at this location form hydrogen bonds and non-bonded interactions with the same residues that interact with the specific ligands of the enzyme. Thus, once the structure of an enzyme is known, computational solvent mapping can provide detailed and reliable information on its substrate-binding site. Calculations on ligand-bound and apo structures of enzymes show that the mapping results are not very sensitive to moderate variations in the protein coordinates.

Algorithms↗

CAPRI: a Critical Assessment of PRedicted Interactions.

CAPRI is a communitywide experiment to assess the capacity of protein-docking methods to predict protein-protein interactions. Nineteen groups participated in rounds 1 and 2 of CAPRI and submitted blind structure predictions for seven protein-protein complexes based on the known structure of the component proteins. The predictions were compared to the unpublished X-ray structures of the complexes. We describe here the motivations for launching CAPRI, the rules that we applied to select targets and run the experiment, and some conclusions that can already be drawn. The results stress the need for new scoring functions and for methods handling the conformation changes that were observed in some of the target systems. CAPRI has already been a powerful drive for the community of computational biologists who development docking algorithms. We hope that this issue of Proteins will also be of interest to the community of structural biologists, which we call upon to provide new targets for future rounds of CAPRI, and to all molecular biologists who view protein-protein recognition as an essential process.

Algorithms↗

Algorithms for computational solvent mapping of proteins.

Computational mapping methods place molecular probes (small molecules or functional groups) on a protein surface to identify the most favorable binding positions by calculating an interaction potential. We have developed a novel computational mapping program called CS-Map (computational solvent mapping of proteins), which differs from earlier mapping methods in three respects: (i) it initially moves the ligands on the protein surface toward regions with favorable electrostatics and desolvation, (ii) the final scoring potential accounts for desolvation, and (iii) the docked ligand positions are clustered, and the clusters are ranked on the basis of their average free energies. To understand the relative importance of these factors, we developed alternative algorithms that use the DOCK and GRAMM programs for the initial search. Because of the availability of experimental solvent mapping data, lysozyme and thermolysin are considered as test proteins. Both DOCK and GRAMM speed up the initial search, and the combined algorithms yield acceptable mapping results. However, the DOCK-based approaches place the consensus site farther from its experimentally determined position than CS-Map, primarily because of the lack of a solvation term in the initial search. The GRAMM-based program also finds the correct consensus site for thermolysin. We conclude that good sampling is the most important requirement for successful mapping, but accounting for desolvation and clustering of ligand positions also help to reduce the number of false positives.

2-Propanol↗

Computational mapping identifies the binding sites of organic solvents on proteins.

Computational mapping places molecular probes--small molecules or functional groups--on a protein surface to identify the most favorable binding positions. Although x-ray crystallography and NMR show that organic solvents bind to a limited number of sites on a protein, current mapping methods result in hundreds of energy minima and do not reveal why some sites bind molecules with different sizes and polarities. We describe a mapping algorithm that explains the origin of this phenomenon. The algorithm has been applied to hen egg-white lysozyme and to thermolysin, interacting with eight and four different ligands, respectively. In both cases the search finds the consensus site to which all molecules bind, whereas other positions that bind only certain ligands are not necessarily found. The consensus sites are pockets of the active site, lined with partially exposed hydrophobic residues and with a number of polar residues toward the edge. These sites can accommodate each ligand in a number of rotational states, some with a hydrogen bond to one of the nearby donor/acceptor groups. Specific substrates and/or inhibitors of hen egg-white lysozyme and thermolysin interact with the same side chains identified by the mapping, but form several hydrogen bonds and bind in unique orientations.

Algorithms↗

Protein-protein association kinetics and protein docking.

Rigid body protein docking methods frequently yield false positive structures that have good surface complementarity, but are far from the native complex. The main reason for this is the uncertainty of the protein structures to be docked, including the positions of solvent-exposed sidechains. Substantial efforts have been devoted to finding near-native structures by rescoring the docked conformations and employing various filters. An alternative approach emulates the process of protein-protein association, that is, first finding the region in which binding is likely to occur and then refining the complex while allowing for flexibility.

Binding Sites↗