PubMed HealthSearch

Biomedical subjects

H J Wolfson

Publications and source records attributed to H J Wolfson.

9 recordsLinked to original sources

A dataset of protein-protein interfaces generated with a sequence-order-independent comparison technique.

While there are a number of structurally non-redundant datasets of protein monomers, there is none of protein-protein interfaces. Yet, the availability of such a dataset is expected to provide an added insight into a number of investigations. First and foremost among these is analyzing the interfaces to obtain their prevailing architectures, the forces that account for the protein-protein associations and their packing considerations. Their comparisons with those of the monomers are likely to shed additional light on protein-protein recognition on the one hand and on the folding of the polypeptide chain on the other. Docking simulations are also expected to benefit from the existence of such a dataset. A major stumbling block to the generation of a dataset of interfaces has been that the interface is composed of at least two chains. Furthermore, in the interfaces, each of the chains might be represented by non-contiguous pieces. Their order in the interfaces being compared might be different as well. This discontinuity stems from the definition of an interface. An interface consists of interacting residues between the chains, and those that are in their vicinity in the supporting scaffold, within a certain distance threshold. This necessarily yields unordered fragments, as well as isolated residues. Our novel, efficient, sequence-order-independent structural comparison technique is ideally suited to handle the task of the generation of a library of structurally non-redundant protein-protein interfaces. As it is computer-vision based, it views atoms as collections of points in space, disregarding their chain connectivity. In this work, 351 interface-families are created. Comparisons of the interfaces, and separately, of the chains which contribute to them, yield some interesting cases. In one of the cases, while two interfaces are similar, the structure of only one of the two chains is similar between the two complexes. The structure of the second chain of the first complex differs from that of the second chain of the second complex. Here the structure of the cleft in the first chain dictates the specific binding interactions. In another case, while the interfaces in the two complexes are similar, both chains composing them differ between the complexes. Lastly, the chains composing the complexes are similar, but the interfaces are dissimilar, providing a set of data for investigations of the favorable orientations of protein-protein associations.

Algorithms

Amino acid pair interchanges at spatially conserved locations.

Here we study the pattern of amino acid interchanges at spatially, locally conserved regions in globally dissimilar and unrelated proteins. By using a method which completely separates the amino acid sequence from its respective structure, this work addresses the question of which properties of the amino acids are the most crucial for the stability of conserved structural motifs. The proteins are taken from a structurally non-redundant dataset. The spatially conserved substructural motifs are defined as consisting of a "large enough" number of Calpha atoms found to provide a geometric match between two proteins, regardless of the order of the Calpha atoms in the sequence, or of the sequence composition of the substructures. This approach can apply to proteins with little or no sequence similarity but with sufficient structural similarity, and is unique in its ability to handle local, non-topological matches between pairs of dissimilar proteins. The method uses a computer-version based algorithm, the Geometric Hashing. Since the Geometric Hashing ignores sequence information it lends itself to answer the question posed above. The interchanges at geometrically similar positions that have been obtained with our method demonstrate the expected behaviour. Yet, a closer inspection reveals some distant characteristics, as compared with interchanges based upon sequence-order based techniques, or from energy-contact-based considerations. First, a pronounced division of the amino acids into two classes is displayed: Lys, Glu, Arg, Gln, Asp, Asn, Pro, Gly, Thr, Ser and His on the one hand, and Ile, Val, Leu, Phe, Met, Tyr, Trp, Cys and Ala on the other. These groups further cluster into subgroups: Lys, Glu, Arg, Gln; Asp Asn; Pro, Gly; Ile, Val, Leu, Phe. The other amino acids stand alone. Analysis of the conservation among amino acids indicates proline to be consistently, by far, the most conserved. Next are Asp, Glu, Lys and Gly. Cys is also highly conserved. Interestingly, oppositely charged amino acids are interchanged roughly as frequently as those of the same charge. These observations can be explained in terms of the three-dimensional structures of the proteins. Most of all, there is a clear distinction between residues which prefer to be on the protein surfaces, compared to those frequently buried in the interiors. Analysis of the interchanges indicates their low information content. This, together with the separation into two groups, suggest that the predictive value of the spatial positions of the Calpha+ atoms is not much greater than the sequence alone, aside from their hydrophobicity/hydrophillicity classification.

Amino Acid Sequence

Protein-protein interfaces: architectures and interactions in protein-protein interfaces and in protein cores. Their similarities and differences.

Protein structures generally consist of favorable folding motifs formed by specific arrangements of secondary structure elements. Similar architectures can be adopted by different amino acids sequences, although the details of the structures vary. It has long been known that despite the sequence variability, there is a striking preferential conservation of the hydrophobic character of the amino acids at the buried positions of these folding motifs. Differences in the sizes of the side-chains are accommodated by movements of the secondary structure elements with respect to each other, leading to compact packing. Scanning protein-protein interfaces reveals that similar architectures are also observed at and around their interacting surfaces, with preservation of the hydrophobic character, although not to the same extent. The general forces that determine the origin of the native structures of proteins have been investigated intensively. The major non-bonded forces operating on a protein chain as it folds into a three-dimensional structure are likely to be packing, the hydrophobic effect, and electrostatic interactions. While the substantial hydrophobic forces lead to a compact conformation, they are also nonspecific and cannot serve as a guide to a conformationally unique structure. For the general folding problem, it thus appears that packing is a prime candidate for determining a particular fold. Specific hydrogen-bonding patterns and salt-bridges have also been proposed to play a role. Inspection of protein-protein interfaces reveals that the hallmarks governing single chain protein structures also determine their interactions, suggesting that similar principles underlie protein folding and protein-protein associations. This review focuses on some aspects of protein-protein interfaces, particularly on the architectures and their interactions. These are compared with those present in protein monomers. This task is facilitated by the recently compiled, non-redundant structural dataset of protein-protein interfaces derived from the crystallographic database. In particular, although current view holds that protein-protein interfaces and interactions are similar to those found in the conformations of single-chain proteins, this review brings forth the differences as well. Not only is it logical that such differences would exist, it is these differences that further illuminate protein folding on the one hand and protein-protein recognition on the other. These are also particularly important in considering inhibitor (ligand) design.

Amino Acids

Molecular surface complementarity at protein-protein interfaces: the critical role played by surface normals at well placed, sparse, points in docking.

Rigid-body docking of two molecules involves matching of their surfaces. A successful docking methodology considers two key issues: molecular surface representation, and matching. While approaches to the problem differ, they all employ certain surface geometric features. While surface normals are routinely created with molecular surfaces, their employment has surprisingly been almost completely overlooked. Here we show how the normals to the surface, at specific, well placed points, can play a critical role in molecular docking. If the points for which the normals are calculated represent faithfully and accurately the molecular surfaces, the normals can substantially ameliorate the efficiency of the docking in a number of ways. The normals can drastically reduce the combinatorial complexity of the receptor-ligand docking. Furthermore, they can serve as a powerful filter in screening for quality docked conformations. Below we show how deploying such a straight forward device, which is easy to calculate, large protein-protein molecules are docked with unparalleled short times and with a manageable number of potential solutions. Considering the facts that here we dock (1) two large protein molecules, including several large immunoglobulin-lysozyme complexes; (2) that we use the entire molecular surfaces, without a predefinition of the active sites, or of the epitopes, of neither the ligand nor the receptor; that (3) the docking is completely automated, without any labelling, or pre-specification, of the input structural database, and (4) with a single set of parameters, without any further tuning whatsoever, such results are highly desirable. This approach is specifically geared towards matching of the surfaces of large protein molecules and is not applicable to small molecule drugs.

Computer Graphics

An automated computer vision and robotics-based technique for 3-D flexible biomolecular docking and matching.

The generation of binding modes between two molecules, also known as molecular docking, is a key problem in rational drug design and biomolecular recognition. Docking a ligand, e.g., a drug molecule or a protein molecule, to a protein receptor, involves recognition of molecular surfaces as molecules interact at their surface. Recent studies report that the activity of many molecules induces conformational transitions by 'hinge-bending', which involves movements of relatively rigid parts with respect to each other. In ligand-receptor binding, relative rotational movements of molecular substructures about their common hinges have been observed. For automatically predicting flexible molecular interactions, we adapt a new technique developed in Computer Vision and Robotics for the efficient recognition of partially occluded articulated objects. These type of objects consist of rigid parts which are connected by rotary joints (hinges). Our approach is based on an extension and generalization of the Geometric Hashing and Generalized Hough Transform paradigm for rigid object recognition. Unlike other techniques which match each part individually, our approach exploits forcefully and efficiently enough the fact that the different rigid parts do belong to the same flexible molecule. We show experimental results obtained by an implementation of the algorithm for rigid and flexible docking. While the 'correct', crystal-bound complex is obtained with a small RMSD, additional, predictive 'high scoring' binding modes are generated as well. The diverse applications and implications of this general, powerful tool are discussed.

Algorithms

Shape complementarity at protein-protein interfaces.

A matching algorithm using surface complementarity between receptor and ligand protein molecules is outlined. The molecular surfaces are represented by "critical points," describing holes and knobs. Holes (maxima of a shape function) are matched with knobs (minima). This simple and appealing surface representation has been previously described by Connolly [(1986) Biopolymers, Vol. 25, pp. 1229-1247]. However, attempts to implement this description in a docking scheme have been unsuccessful (e.g., Connolly, ibid.). In order to decrease the combinatorial complexity, and to make the execution time affordable, four critical hole/knob point matches were sought. This approach failed since some bound interfaces are relatively flat and do not possess four critical point matches. On the otherhand, matchings of fewer critical points require a very time-consuming, full conformational (grid) space search [Wang, (1991) Journal of Computational Chemistry, Vol. 12, pp. 746-750]. Here we show that despite the initial failure of this approach, with a simple and straightforward modification in the matching algorithm, this surface representation works well. Out of the 16 protein-protein complexes we have tried, 15 were successfully docked, including two immunoglobulins. The entire molecular surfaces were considered, with absolutely no additional information regarding the binding sites. The whole process is completely automated, with no manual intervention, either in the input atomic coordinate data, or in the matching. We have been able to reach this level of performance with the hole/knob surface description by using pairs of critical points along with their surface normals in the calculation of the transformation matrix. The success of this approach suggests that future docking methods should use geometric docking as the first screening filter.(ABSTRACT TRUNCATED AT 250 WORDS)

Algorithms

Molecular surface representations by sparse critical points.

We have defined a molecular surface representation that describes precisely and concisely the complete molecular surface. The representation consists of a limited number of critical points disposed at key locations over the surface. These points adequately represent the shape and the important characteristics of the surface, despite the fact that they are modest in number. We expect the representation to be useful in areas such as molecular recognition and visualization. In particular, using this representation, we are able to achieve accurate and efficient protein-protein and protein-small molecule docking.

Chymotrypsin

Molecular surface recognition by a computer vision-based technique.

Correct docking of a ligand onto a receptor surface is a complex problem, involving geometry and chemistry. Geometrically acceptable solutions require close contact between corresponding patches of surfaces of the receptor and of the ligand and no overlap between the van der Waals spheres of the remainder of the receptor and ligand atoms. In the quest for favorable chemical interactions, the next step involves minimization of the energy between the docked molecules. This work addresses the geometrical aspect of the problem. It is assumed that we have the atomic coordinates of each of the molecules. In principle, since optimally matching surfaces are sought, the entire conformational space needs to be considered. As the number of atoms residing on molecular surfaces can be several hundred, sampling of all rotations and translations of every patch of a surface of one molecule with respect to the other can reach immense proportions. The problem we are faced with here is reminiscent of object recognition problems in computer vision. Here we borrow and adapt the geometric hashing paradigm developed in computer vision to a central problem in molecular biology. Using an indexing approach based on a transformation invariant representation, the algorithm efficiently scans groups of surface dots (or atoms) and detects optimally matched surfaces. Potential solutions displaying receptor--ligand atomic overlaps are discarded. Our technique has been applied successfully to seven cases involving docking of small molecules, where the structures of the receptor--ligand complexes are available in the crystallographic database and to three cases where the receptors and ligands have been crystallized separately. In two of these three latter tests, the correct transformations have been obtained.

Algorithms

Efficient detection of three-dimensional structural motifs in biological macromolecules by computer vision techniques.

Macromolecules carrying biological information often consist of independent modules containing recurring structural motifs. Detection of a specific structural motif within a protein (or DNA) aids in elucidating the role played by the protein (DNA element) and the mechanism of its operation. The number of crystallographically known structures at high resolution is increasing very rapidly. Yet, comparison of three-dimensional structures is a laborious time-consuming procedure that typically requires a manual phase. To date, there is no fast automated procedure for structural comparisons. We present an efficient O(n3) worst case time complexity algorithm for achieving such a goal (where n is the number of atoms in the examined structure). The method is truly three-dimensional, sequence-order-independent, and thus insensitive to gaps, insertions, or deletions. This algorithm is based on the geometric hashing paradigm, which was originally developed for object recognition problems in computer vision. It introduces an indexing approach based on transformation invariant representations and is especially geared toward efficient recognition of partial structures in rigid objects belonging to large data bases. This algorithm is suitable for quick scanning of structural data bases and will detect a recurring structural motif that is a priori unknown. The algorithm uses protein (or DNA) structures, atomic labels, and their three-dimensional coordinates. Additional information pertaining to the structure speeds the comparisons. The algorithm is straightforwardly parallelizable, and several versions of it for computer vision applications have been implemented on the massively parallel connection machine. A prototype version of the algorithm has been implemented and applied to the detection of substructures in proteins.

Algorithms