PubMed HealthSearch

Biomedical subjects

P Willett

Publications and source records attributed to P Willett.

12 recordsLinked to original sources

Three-dimensional structural resemblance between leucine aminopeptidase and carboxypeptidase A revealed by graph-theoretical techniques.

Using 3-D searching techniques based on algorithms derived from graph theory we have established a striking structural similarity between the structure of bovine carboxypeptidase A and that of the C-terminal domain of bovine leucine aminopeptidase. There is no significant sequence homology between the aminopeptidases and the carboxypeptidases but the strong structural relationship detected in this complex fold suggests that there may be a very remote divergent evolutionary relationship between these two enzyme classes.

Animals

Pharmacophoric pattern matching in files of three-dimensional chemical structures: use of bounded distance matrices for the representation and searching of conformationally flexible molecules.

This paper discusses the use of bounded distance matrices for the representation of conformationally flexible three-dimensional (3D) molecules. It is shown that pharmacophoric pattern searches of databases of flexible 3D molecules represented in this way can be carried out using screen and geometric searching algorithms that are analogous to those used for searching databases of rigid 3D structures. Molecules matching a query pattern after the geometric search must then undergo a final conformational search to determine whether they can, in fact, adopt a conformation that matches the query. An analysis of this three-stage searching procedure shows that searching databases of flexible 3D molecules is extremely demanding of computational resources.

Algorithms

Techniques for the calculation of three-dimensional structural similarity using inter-atomic distances.

This paper reports a comparison of several methods for measuring the degree of similarity between pairs of 3-D chemical structures that are represented by inter-atomic distance matrices. The methods that have been tested use the distance information in very different ways and have very different computational requirements. Experiments with 10 small datasets, for which both structural and biological activity data are available, suggest that the most cost-effective technique is based on a mapping procedure that tries to match pairs of atoms, one from each of the molecules that are being compared, that have neighbouring atoms at approximately the same distances.

Computer Simulation

Pharmacophoric pattern matching in files of three-dimensional chemical structures: use of smoothed bounded distances for incompletely specified query patterns.

This paper describes a technique for increasing the screen-out of pharmacophoric pattern searches in the databases of three-dimensional chemical structures when only some of the interatomic distances in the query pattern are specified. The technique involves the application of a distance bounds-smoothing procedure to the query distances; this smoothing allows the calculation of upper and lower bounds for the unspecified distances. The bounded distances can then be used to set screens additional to those that are set to describe the distances that have been specified by the searcher. Evidence is presented to suggest that use of the technique can lead to increases in the efficiency of substructure searches for partially specified query patterns.

Databases, Factual

Pharmacophoric pattern matching in files of three-dimensional chemical structures: characterization and use of generalized valence angle screens.

This paper describes the use of generalized valence angles for the screening of pharmacophoric pattern searches in databases of three-dimensional chemical structures. A generalized valence angle is defined as the angle between two vectors, AB and BC, which have a common vertex B, and in which both vectors correspond to formal chemical bonds; one vector corresponds to a bond and the other to a non-bonded interaction; or both vectors correspond to non-bonded interactions. The screens are identified by a statistical analysis of the frequencies of occurrence of these angle-based features in the Cambridge Structural Database. The occurrence frequencies are discussed and shown to be explicable in terms of small, commonly occurring structural features. The effectiveness of the screens is demonstrated by an extensive series of searches for representative pharmacophoric patterns. The results are compared with those obtained from a similar series of searches using distance-based screens: The latter are found to give a better level of performance, and evidence is presented to suggest that this is due to a high degree of association between the assignments of the angle-based screens.

Computers

Use of techniques derived from graph theory to compare secondary structure motifs in proteins.

A substructure matching algorithm is described that can be used for the automatic identification of secondary structural motifs in three-dimensional protein structures from the Protein Data Bank. The proteins and motifs are stored for searching as labelled graphs, with the nodes of a graph corresponding to linear representations of helices and strands and the edges to the inter-line angles and distances. A modification of Ullman's subgraph isomorphism algorithm is described that can be used to search these graph representations. Tests with patterns from the protein structure literature demonstrate both the efficiency and the effectiveness of the search procedure, which has been implemented in FORTRAN 77 on a MicroVAX-II system, coupled to the molecular fitting program FRODO on an Evans and Sutherland PS300 graphics system.

Algorithms

Structural resemblance between the families of bacterial signal-transduction proteins and of G proteins revealed by graph theoretical techniques.

The first application of a novel technique for the identification of common folding motifs in proteins is presented. Using techniques derived from graph theory, developed in order to compare secondary structure motifs in proteins, we have established that there is a striking resemblance in the tertiary fold of the Salmonella typhimurium Che Y chemotaxis protein and that of the GDP-binding domain of Escherichia coli elongation factor Tu (EF Tu). These two protein structures are representatives of two major macromolecular classes: CheY is a signal-transduction protein with sequence homologies to a wide range of bacterial proteins involved in regulation of chemotaxis, membrane synthesis and sporulation; whilst EF Tu is one of a family of guanosine-nucleotide-binding proteins which include the ras oncogene proteins and signal-transducing G proteins. The similarity we have found extends far beyond the previously recognized resemblances of each protein's fold to that of a generic nucleotide-binding domain. The lack of significant sequence homology between the two classes of proteins may mean that the common fold of the two proteins constitutes a particularly stable folding motif. However, an alternative possibility is that the strong three-dimensional structural resemblance may be indicative of a remote shared common ancestry between the bacterial signal-transduction proteins and the GDP-binding proteins.

Algorithms

Upperbound procedures for the identification of similar three-dimensional chemical structures.

This paper describes techniques for calculating the degree of similarity between an input query molecule and each of the molecules in a database of 3-D chemical structures. The inter-molecular similarity measure used is the number of atoms in the 3-D common substructure (CS) between the two molecules which are being compared. The identification of 3-D CSs is very demanding of computational resources, even when an efficient clique detection algorithm is used for this purpose. Two types of upperbound calculation are described which allow reductions in the number of exact CS searches which need to be carried out to identify those molecules from a database which are similar to a 3-D target molecule.

Algorithms

Donor eyes. A comparison of characteristics and outcomes for Eye Bank and local tissue.

The Corneal Recipient Registry was begun in 1985 to collect information on all recipients of corneal grafts in the province of Ontario, Canada, and on the donors providing tissue. While most of the tissue is handled by the Eye Bank of Canada (Ontario Division), ophthalmologists in centers away from the Eye Bank often use local tissue when it is available. Comparison of the donor characteristics of local tissue with that obtained through the Eye Bank revealed that local donors were 9-10 years younger (p less than 0.01), their times to enucleation were an hour less (p less than 0.02), and they were much more likely to be the victims of trauma than the donors of Eye Bank eyes. Prognosis of the graft, assessed using life table methods, suggested that success of local eyes was 89% after 6 months, compared with 80% for Eye Bank eyes in the same period, but this was not a significant difference (p greater than 0.05). While the Eye Bank is a more common source of tissue, eyes obtained locally are more likely to represent the "ideal" tissue for many corneal surgeons.

Adult

Compression of nucleic acid and protein sequence data.

This paper describes the application of text compression methods to machine-readable files of nucleic acid and protein sequence data. Two main methods are used to reduce the storage requirements of such files, these being n-gram coding and run-length coding. A Pascal program combining both of these techniques resulted in a compression figure of 74.6% for the GenBank data-base and a program that used only n-gram coding gave a compression figure of 42.8% for the Protein Identification Resource database.

Algorithms

Similarity searching in databases of three-dimensional molecules and macromolecules.

This paper discusses algorithmic techniques for measuring the degree of similarity between pairs of three-dimensional (3-D) chemical molecules represented by interatomic distance matrices. A comparison of four methods for the calculation of 3-D structural similarity suggests that the most effective one is a procedure that identifies pairs of atoms, one from each of the molecules that are being compared, that lie at the center of geometrically-related volumes of 3-D space. This atom mapping method enables the calculation of a wide range of types of intermolecular similarity coefficient, including measures that are based on physicochemical data. Massively-parallel implementations of the method are discussed, using the AMT Distributed Array Processor, that achieve a substantial increase in performance when compared with a sequential implementation on a UNIX workstation. Current work involves the use of angular information and the extension of the method to field-based similarity searching. Similarity searching in 3-D macromolecules is effected by the use of a maximal common subgraph (MCS) isomorphism algorithm with a novel, graph-based representation of the tertiary structures of proteins. This algorithm is being used to identify similarities between the 3-D structures of proteins in the Brookhaven Protein Data Bank; its use is exemplified by searches involving the NAD-binding fold motif.

Algorithms