PubMed HealthSearch

Biomedical subjects

A C May

Publications and source records attributed to A C May.

7 recordsLinked to original sources

Finding local structural similarities among families of unrelated protein structures: a generic non-linear alignment algorithm.

We have developed a generic tool for the automatic identification of regions of local structural similarity in unrelated proteins having different folds, as well as for defining more global similarities that result from homologous protein structures. The computer program GENFIT has evolved from the genetic algorithm-based three-dimensional protein structure comparison program GA_FIT. GENFIT, however, can locate and superimpose regions of local structural homology regardless of their position in a pair of structures, the fold topology, or the chain direction. Furthermore, it is possible to restrict the search to a volume centered about a region of interest (e.g., catalytic site, ligand-binding site) in two protein structures. We present a number of examples to illustrate the function of the program, which is a parallel processing implementation designed for distribution to multiple machines over a local network or to run on a single multiprocessor computer.

Algorithms

Pairwise iterative superposition of distantly related proteins and assessment of the significance of 3-D structural similarity.

A challenge lies in identifying distant protein 3-D structural similarity by rigid-body superposition. The most common measure of structural similarity is r.m.s. distance (r.m.s.d.) between topologically equivalent residues, and most automated methods of protein modelling rely on the assembly of rigid fragments from known 3-D structures. A fast method of improving the definition of a common protein fold by superposition, especially for distant relationships, is described. The definition of topological equivalence by the standard dynamic programming sequence alignment algorithm is extended by refining the entire structure alignment (not just those equivalenced residues within a given cut-off distance) and determining whether the alignment can be continued at the termini. The most appropriate distance-based definition of topological equivalence for a given comparison is identified. Despite the fact that hitherto the distant similarity between the globin fold and colicin A has not been recognized directly by rigid-body superposition, this new approach defines more equivalent residues with a lower r.m.s.d. between them than that obtained by the superposition of equivalences identified by a more elaborate method. A previous distance metric of 3-D structural similarity derived from rigid-body superposition has been extended to the assessment of superpositions where topological equivalences have been determined by methods other than rigid-body ones.

Algorithms

Improved genetic algorithm-based protein structure comparisons: pairwise and multiple superpositions.

Three major improvements to a previously described method for automatic protein structure comparison are described. First, a limit to translations for the rigid-body superposition is now assigned according to the dimensions of the structures being compared. Second, examination of the effect of the gap penalty on the derivation of a sequence alignment corresponding to a given structure superposition has led to a method to evaluate alternative structure-based sequence alignments. Third, the pairwise procedure has been generalized to multiple structure alignment. This implementation of rigid-body superposition can recognize well documented distant relationships which hitherto have required consideration of additional features and properties as well as those relationships between proteins of different sizes. A much larger common scaffold or framework between six globins can be extracted than that obtained using a standard algorithm for multiple structure superposition.

Algorithms

The recognition of protein structure and function from sequence: adding value to genome data.

The explosion of DNA sequence data from genome projects presents many challenges. For instance, we must extend our current knowledge of protein structure and function so that it can be applied to these new sequences. The derivation of rules for the relationships between sequence and structure allow us to recognize a common fold by the use of tertiary templates. New techniques enable us to begin to meet the challenge of rule-based modelling of distantly related proteins. This paper describes an integrated and knowledge-based approach to the prediction of protein structure and function which can maximize the value of sequence information.

Amino Acid Sequence

Automated comparative modelling of protein structures.

Most automated methods for protein modelling rely on the assembly of rigid fragments from known three-dimensional structures. Although the modelling of side chains has received attention recently, loop regions continue to be neglected, and these conformations contribute most errors to models. Comparative modelling procedures using distance restraints are particularly useful for modelling distantly related proteins with a common fold. Progress has also recently been made in the automatic evaluation of model structures.

Computer Simulation

Protein structure comparisons using a combination of a genetic algorithm, dynamic programming and least-squares minimization.

We introduce a completely automatic and objective procedure for the comparison of protein structures. A genetic algorithm is used to search for a near optimal solution of the rigid-body superposition of two whole protein structures. The specification of an initial set of equivalences is not required. Topological equivalences in the final structural alignment are defined by a conventional dynamic programming routine, which is commonly used to compare protein sequences. A least-squares fitting algorithm is then used to optimize the fit between the final set of equivalences. We have applied our method to the comparison of ribonucleic acid structures, as well as protein structures. The structural alignments are generally consistent with those previously published. In fact, on most occasions our method defines at least the same number of topological equivalences as other procedures, but always with a lower r.m.s. distance between them.

Algorithms