PubMed Health⌕ Search

Biomedical subjects

Peter Willett

Publications and source records attributed to Peter Willett.

15 recordsLinked to original sources

Comparison of chemical clustering methods using graph- and fingerprint-based similarity measures.

This paper compares several published methods for clustering chemical structures, using both graph- and fingerprint-based similarity measures. The clusterings from each method were compared to determine the degree of cluster overlap. Each method was also evaluated on how well it grouped structures into clusters possessing a non-trivial substructural commonality. The methods which employ adjustable parameters were tested to determine the stability of each parameter for datasets of varying size and composition. Our experiments suggest that both graph- and fingerprint-based similarity measures can be used effectively for generating chemical clusterings; it is also suggested that the CAST and Yin-Chen methods, suggested recently for the clustering of gene expression patterns, may also prove effective for the clustering of 2D chemical structures.

Algorithms↗

A sphere-based descriptor for matching protein structures.

This paper describes the use of a descriptor based on the number of alpha-carbon atoms within a sphere centered on the alpha-carbon in each amino acid residue in a protein. The descriptor can be used instead of the residue types in a dynamic programming algorithm, thus providing an efficient way of aligning protein structures. The method is applied to the alignment of protein families and to database searching. The results indicate that the method can quickly align protein sequences considering the 3D structure and can find proteins that are 3D structurally similar and dissimilar to the target protein.

Amino Acid Sequence↗

Designing focused libraries using MoSELECT.

When designing a combinatorial library it is usually desirable to optimise multiple properties of the library simultaneously and often the properties are in competition with one another. For example, a library that is designed to be focused around a given target molecule should ideally have minimum cost and also contain molecules that are bioavailable. In this paper, we describe the program MoSELECT for multiobjective library design that is based on a multiobjective genetic algorithm (MOGA). MoSELECT searches the product-space of a virtual combinatorial library to generate a family of equivalent solutions where each solution represents a combinatorial subset of the virtual library optimised over multiple objectives. The family of solutions allows the relationships between the objectives to be explored and thus enables the library designer to make an informed choice on an appropriate compromise solution. Experiments are reported where MoSELECT has been applied to the design of various focused libraries.

Algorithms↗

Effectiveness of graph-based and fingerprint-based similarity measures for virtual screening of 2D chemical structure databases.

This paper reports an evaluation of both graph-based and fingerprint-based measures of structural similarity, when used for virtual screening of sets of 2D molecules drawn from the MDDR and ID Alert databases. The graph-based measures employ a new maximum common edge subgraph isomorphism algorithm, called RASCAL, with several similarity coefficients described previously for quantifying the similarity between pairs of graphs. The effectiveness of these graph-based searches is compared with that resulting from similarity searches using BCI, Daylight and Unity 2D fingerprints. Our results suggest that graph-based approaches provide an effective complement to existing fingerprint-based approaches to virtual screening.

Algorithms↗

Maximum common subgraph isomorphism algorithms for the matching of chemical structures.

The maximum common subgraph (MCS) problem has become increasingly important in those aspects of chemoinformatics that involve the matching of 2D or 3D chemical structures. This paper provides a classification and a review of the many MCS algorithms, both exact and approximate, that have been described in the literature, and makes recommendations regarding their applicability to typical chemoinformatics tasks.

Algorithms↗

Combinatorial library design using a multiobjective genetic algorithm.

Early results from screening combinatorial libraries have been disappointing with libraries either failing to deliver the improved hit rates that were expected or resulting in hits with characteristics that make them undesirable as lead compounds. Consequently, the focus in library design has shifted toward designing libraries that are optimized on multiple properties simultaneously, for example, diversity and "druglike" physicochemical properties. Here we describe the program MoSELECT that is based on a multiobjective genetic algorithm and which is able to suggest a family of solutions to multiobjective library design where all the solutions are equally valid and each represents a different compromise between the objectives. MoSELECT also allows the relationships between the different objectives to be explored with competing objectives easily identified. The library designer can then make an informed choice on which solution(s) to explore. Various performance characteristics of MoSELECT are reported based on a number of different combinatorial libraries.

Algorithms↗

Heuristics for similarity searching of chemical graphs using a maximum common edge subgraph algorithm.

Recently a method (RASCAL) for determining graph similarity using a maximum common edge subgraph algorithm has been proposed which has proven to be very efficient when used to calculate the relative similarity of chemical structures represented as graphs. This paper describes heuristics which simplify a RASCAL similarity calculation by taking advantage of certain properties specific to chemical graph representations of molecular structure. These heuristics are shown experimentally to increase the efficiency of the algorithm, especially at more distant values of chemical graph similarity.

Journal Article↗

Generation and display of activity-weighted chemical hyperstructures.

A chemical hyperstructure is a single graph representation of a set of molecules that minimizes the degree of structural redundancy in the data set. This paper describes the use of a genetic algorithm to generate an activity-weighted chemical hyperstructure (AWCH) by sequentially mapping each molecule in the data set to the hyperstructure and then assigning activity and inactivity frequency weights to the nodes and edges of the hyperstructure. Experiments with several data sets demonstrate the level of activity clustering in an AWCH.

Journal Article↗

Comparison of ranking methods for virtual screening in lead-discovery programs.

This paper discusses the use of several rank-based virtual screening methods for prioritizing compounds in lead-discovery programs, given a training set for which both structural and bioactivity data are available. Structures from the NCI AIDS data set and from the Syngenta corporate database were represented by two types of fragment bit-string and by sets of high-level molecular features. These representations were processed using binary kernel discrimination, similarity searching, substructural analysis, support vector machine, and trend vector analysis, with the effectiveness of the methods being judged by the extent to which active test set molecules were clustered toward the top of the resultant rankings. The binary kernel discrimination approach yielded consistently superior rankings and would appear to have considerable potential for chemical screening applications.

Journal Article↗

Calculation of intersubstituent similarity using R-group descriptors.

This paper discusses the calculation of the similarities between pairs of substituents on ring systems. An R-group descriptor characterizes the distribution of some atom-based property, such as elemental type or partial atomic charge, at increasing numbers of bonds distant from the point of substitution on the parent ring. The similarity between a pair of descriptors is then calculated by a comparison of the corresponding property vectors. Experiments with the BIOSTER database demonstrate the ability of such similarity measures to discriminate between bioisosteric and nonbioisosteric functional groups.

Journal Article↗

Evaluation of similarity measures for searching the dictionary of natural products database.

Similarity searches using combinations of seven different similarity coefficients and six different representations have been carried out on the Dictionary of Natural Products database. The objective was to discover if any special methods of searching apply to this database, which is very different in nature from the many synthetic databases that have been the subject of previous studies of similarity searching. Search effectiveness was assessed by a recall analysis of the search outputs from sets of pharmacologically active target structures. The different target sets produce exceptional but contradictory results for the Russell-Rao and Forbes coefficients, which have been shown to be due to a dependence on molecular size; these are the coefficients of choice in the case of large and small structures, respectively. Rankings from these results have been combined using a data fusion scheme and some small gains in performance were normally obtained by using substructural fingerprints and molecular holograms in combination with the Squared Euclidean or Tanimoto coefficients.

Biological Products↗

Similarity searching using reduced graphs.

Reduced graphs provide summary representations of chemical structures. In this work, the effectiveness of reduced graphs for similarity searching is investigated. Different types of reduced graphs are introduced that aim to summarize features of structures that have the potential to form interactions with receptors while retaining the topology between the features. Similarity searches have been carried out across a variety of different activity classes. The effectiveness of the reduced graphs at retrieving compounds with the same activity as known target compounds is compared with searching using Daylight fingerprints. The reduced graphs are shown to be effective for similarity searching and to retrieve more diverse active compounds than those found using Daylight fingerprints; they thus represent a complementary similarity searching tool.

Journal Article↗

Combination of fingerprint-based similarity coefficients using data fusion.

Many different types of similarity coefficients have been described in the literature. Since different coefficients take into account different characteristics when assessing the degree of similarity between molecules, it is reasonable to combine them to further optimize the measures of similarity between molecules. This paper describes experiments in which data fusion is used to combine several binary similarity coefficients to get an overall estimate of similarity for searching databases of bioactive molecules. The results show that search performances can be improved by combining coefficients with little extra computational cost. However, there is no single combination which gives a consistently high performance for all search types.

Journal Article↗

Searching for patterns of amino acids in 3D protein structures.

This paper describes the program ASSAM, which has been developed to search for patterns of amino acid side-chains in the 3D structures in the Protein Data Bank. ASSAM represents an amino acid by a vector drawn from the main chain towards the functional part of the amino acid and then computes a graph representation of a protein in which the individual side-chain vectors are the nodes and the intervector distances are the edges. The presence of a query pattern in a Protein Data Bank structure can then be searched for by means of a subgraph isomorphism algorithm. Recent enhancements to ASSAM allow searches to include the following: the main-chain structure in addition to the side-chains; the secondary structure and solvent accessibility of side-chains; allowable distances from a known binding-site; disulfide bridges; and improved generic and wild-card queries. The effectiveness of these approaches is demonstrated by extensive searches of the Protein Data Bank for typical 3D query patterns.

Algorithms↗

CLIP: similarity searching of 3D databases using clique detection.

This paper describes a program for 3D similarity searching, called CLIP (for Candidate Ligand Identification Program), that uses the Bron-Kerbosch clique detection algorithm to find those structures in a file that have large structures in common with a target structure. Structures are characterized by the geometric arrangement of pharmacophore points and the similarity between two structures calculated using modifications of the Simpson and Tanimoto association coefficients. This modification takes into account the fact that a distance tolerance is required to ensure that pairs of interatomic distances can be regarded as equivalent during the clique-construction stage of the matching algorithm. Experiments with HIV assay data demonstrate the effectiveness and the efficiency of this approach to virtual screening.

Journal Article↗