PubMed Health⌕ Search

Biomedical subjects

J W Godden

Publications and source records attributed to J W Godden.

13 recordsLinked to original sources

Mini-fingerprints for virtual screening: design principles and generation of novel prototypes based on information theory.

Binary fingerprint representations of molecular structure and properties are convenient computational tools for similarity searching in compound databases and virtual screening (VS). We are investigating the design of relatively simple fingerprints for the identification of molecules having similar biological activity and recognition of remote similarity relationships. Since our designs are considerably shorter than other fingerprints used in VS, we have previously termed them "mini-fingerprints" (MFPs). A key aspect of the design strategy is the identification of suitable molecular descriptors. Whereas our initial fingerprint designs have relied on descriptor combinations that performed well in compound classification according to biological activity, second generation MFPs encode combinations of descriptors with high information content in large compound databases and high frequency of occurrence in drug-like molecules. Thus, the design of these new fingerprints does not depend on the analysis of specific classes of bioactive compounds, but rather on descriptor information content in large compound databases. Systematic evaluation of fingerprint performance in VS test calculations demonstrates that these new prototypes perform better than previously generated MFPs. The analysis described herein provides an example for the development of search tools for VS.

Environmental Pollutants↗

Searching for molecules with similar biological activity: analysis by fingerprint profiling.

We have recently developed a mini-fingerprint (MFP) representation for small molecules that performs well in database searches for compounds with similar biological activity. The MFP consists of only 54 bit positions that account for numerical ranges of three two-dimensional (2D) descriptors or the presence or absence of defined structural fragments. Here we present an analysis method, termed fingerprint profiling, to systematically compare bit patterns of compounds belonging to different biological activity classes. Some but not all bit positions were variably occupied in seven different activity classes and responsible for the detection of structure-activity differences. The analysis has made it possible to rank bit positions and encoded molecular descriptors according to their importance for our similarity search calculations. Fingerprint profiling can be applied to any keyed bit string representation and should be helpful, for example, to analyze descriptor distributions in large compound databases.

Computer-Aided Design↗

Molecular scaffold-based design and comparison of combinatorial libraries focused on the ATP-binding site of protein kinases.

Compound libraries were designed to target specifically the ATP cofactor-binding site in protein kinases by combining knowledge- and diversity-based design elements. A key aspect of the approach is the identification of molecular building blocks or scaffolds that are compatible with the binding site and therefore capture some aspects of target specificity. Scaffolds were selected on the basis of docking calculations and analysis of known inhibitors. We have generated 75 molecular scaffolds and applied different strategies to compute diverse compounds from scaffolds or, alternatively, to screen compound databases for molecules containing these scaffolds. The resulting libraries had a similar degree of molecular diversity, with at most 12% of the compounds being identical. However, their scaffold distributions differed significantly and a small number of scaffolds dominated the majority of compounds in each library.

Adenosine Triphosphate↗

Evaluation of docking strategies for virtual screening of compound databases: cAMP-dependent serine/threonine kinase as an example.

In an effort to establish efficient docking routines for computational screening of compound databases on protein structures, cAMP-dependent protein kinase has been selected as a test case and a variety of docking options and scoring functions were compared. These included rigid-body and flexible docking and scoring based on surface complementarity and/or force field energy. Inhibitors were removed from complex crystal structures and added to compound libraries in their binding conformations and, in addition, deliberately modified conformations. Rigid-body docking and contact scoring well reproduced two of three experimental enzyme-inhibitor complexes. Ligand docking with flexible torsional angles failed to do so but anchored search of some inhibitors converged near to experimental structures, however, only when energy scoring was applied.

1-(5-Isoquinolinesulfonyl)-2-Methylpiperazine↗

The structure of copper-nitrite reductase from Achromobacter cycloclastes at five pH values, with NO2- bound and with type II copper depleted.

High resolution x-ray crystallographic structures of nitrite reductase from Achromobacter cycloclastes, undertaken in order to understand the pH optimum of the reaction with nitrite, show that at pH 5.0, 5.4, 6.0, 6.2, and 6.8, no significant changes occur, other than in the occupancy of the type II copper at the active site. An extensive network of hydrogen bonds, both within and between subunits of the trimer, maintains the rigidity of the protein structure. A water occupies a site approximately 1.5 A from the site of the type II copper in the structure of the type II copper-depleted structure (at pH 5.4), again with no other significant changes in structure. In nitrite-soaked crystals, nitrite binds via its oxygens to the type II copper and replaces the water normally bound to the type II copper. The active-site cavity of the protein is distinctly hydrophobic on one side and hydrophilic on the other, providing a possible path for diffusion of the product NO. Asp-98 exhibits thermal parameter values higher than its surroundings, suggesting a role in shuttling the two protons necessary for the overall reaction. The strong structural homology with cupredoxins is described.

Alcaligenes↗

The 2.3 angstrom X-ray structure of nitrite reductase from Achromobacter cycloclastes.

The three-dimensional crystal structure of the copper-containing nitrite reductase (NIR) from Achromobacter cycloclastes has been determined to 2.3 angstrom (A) resolution by isomorphous replacement. The monomer has two Greek key beta-barrel domains similar to that of plastocyanin and contains two copper sites. The enzyme is a trimer both in the crystal and in solution. The two copper atoms in the monomer comprise one type I copper site (Cu-I; two His, one Cys, and one Met ligands) and one putative type II copper site (Cu-II; three His and one solvent ligands). Although ligated by adjacent amino acids Cu-I and Cu-II are approximately 12.5 A apart. Cu-II is bound with nearly perfect tetrahedral geometry by residues not within a single monomer, but from each of two monomers of the trimer. The Cu-II site is at the bottom of a 12 A deep solvent channel and is the site to which the substrate (NO2-) binds, as evidenced by difference density maps of substrate-soaked and native crystals.

Alcaligenes↗

Mini-fingerprints detect similar activity of receptor ligands previously recognized only by three-dimensional pharmacophore-based methods.

Mini-fingerprints (MFPs) are short binary bit string representations of molecular structure and properties, composed of few selected two-dimensional (2D) descriptors and a number of structural keys. MFPs were specifically designed to recognize compounds with similar activity. Here we report that MFPs are capable of detecting similar activities of some druglike molecules, including endothelin A antagonists and alpha(1)-adrenergic receptor ligands, the recognition of which was previously thought to depend on the use of multiple point three-dimensional (3D) pharmacophore methods. Thus, in these cases, MFPs and pharmacophore fingerprints produce similar results, although they define, in terms of their complexity, opposite ends of the spectrum of methods currently used to study molecular similarity or diversity. For each of the studied compound classes, comparison of MFP bit settings identified a consensus or signature pattern. Scaling factors can be applied to these bits in order to increase the probability of finding compounds with similar activity by virtual screening.

Angiotensin II↗

Fingerprint scaling increases the probability of identifying molecules with similar activity in virtual screening calculations.

Results of systematic virtual screening calculations using a structural key-type fingerprint are reported for compounds belonging to 14 activity classes added to randomly selected synthetic molecules. For each class, a fingerprint profile was calculated to monitor the relative occupancy of fingerprint bit positions. Consensus bit patterns were determined consisting of all bits that were always set on in compounds belonging to a specific activity class. In virtual screening calculations, scale factors were applied to each consensus bit position in fingerprints of query molecules. This technique, called "fingerprint scaling", effectively increases the weight of consensus bit positions in fingerprint comparisons. Although overall prediction accuracy was satisfactory using unscaled calculations, scaling significantly increased the number of correct predictions but only slightly increased the rate of false positives. These observations suggest that fingerprint scaling is an attractive approach to increase the probability of identifying molecules with similar activity by virtual screening. It requires the availability of a series of related compounds and can be easily applied to any keyed fingerprint representation that associates bit positions with specific molecular features.

Algorithms↗

Evaluation of descriptors and mini-fingerprints for the identification of molecules with similar activity.

Combinations of 65 preferred 1D/2D molecular descriptors and 143 single structural keys were evaluated for their performance in compound classification focused on biological activity. The analysis was based on principal component analysis of descriptor combinations and facilitated by use of a genetic algorithm and different scoring functions. In these calculations, several descriptor combinations with greater than 95% prediction accuracy were identified. A set of 40 preferred structural keys was incorporated into a small binary fingerprint designed to search databases for compounds with biological activity similar to query molecules. The performance of mini-fingerprints was tested by systematic similarity search calculations in a database consisting of compounds belonging to seven biological activity classes, which had not been used to select effective descriptors. In these blind test calculations, mini-fingerprints correctly identified approximately 54% of compounds sharing similar biological activity and with 1% false positives. Thus, although the design of mini-fingerprints is conceptually simple, they perform well in activity-oriented similarity searching.

Computing Methodologies↗

Distinguishing between natural products and synthetic molecules by descriptor Shannon entropy analysis and binary QSAR calculations.

Molecular descriptors were identified by Shannon entropy analysis that correctly distinguished, in binary QSAR calculations, between naturally occurring molecules and synthetic compounds. The Shannon entropy concept was first used in digital communication theory and has only very recently been applied to descriptor analysis. Binary QSAR methodology was originally developed to correlate structural features and properties of compounds with a binary formulation of biological activity (i.e., active or inactive) and has here been adapted to correlate molecular features with chemical source (i.e., natural or synthetic). We have identified a number of molecular descriptors with significantly different Shannon entropy and/or "entropic separation" in natural and synthetic compound databases. Different combinations of such descriptors and variably distributed structural keys were applied to learning sets consisting of natural and synthetic molecules and used to derive predictive binary QSAR models. These models were then applied to predict the source of compounds in different test sets consisting of randomly collected natural and synthetic molecules, or, alternatively, sets of natural and synthetic molecules with specific biological activities. On average, greater than 80% prediction accuracy was achieved with our best models. For the test case consisting of molecules with specific activities, greater than 90% accuracy was achieved. From our analysis, some chemical features were identified that systematically differ in many naturally occurring versus synthetic molecules.

Algorithms↗

Differential Shannon Entropy as a sensitive measure of differences in database variability of molecular descriptors.

A method termed Differential Shannon Entropy (DSE) is introduced to compare differences in information content and variance of molecular descriptors between compound databases. The analysis is based on histograms recording the individual and grouped distributions of molecular descriptors and calculation of Shannon entropy (SE), a formalism originally applied to digital communication. We have recently shown that SE values reflect the nonparametric variability of descriptor settings. Now the analysis has been advanced to assess differences in information content of 143 molecular descriptors in databases containing synthetic compounds, natural products, or drug-like molecules. The DSE metric captures the degree to which descriptor distributions complement or duplicate information contained in molecular databases. In our analysis, we observe significant differences for a number of descriptors and rank them according to their associated DSE values. Using DSE calculations, relative information content of different types of descriptors can be quantified, even if differences are subtle.

Journal Article↗

Database searching for compounds with similar biological activity using short binary bit string representations of molecules.

In an effort to identify biologically active molecules in compound databases, we have investigated similarity searching using short binary bit strings with a maximum of 54 bit positions. These "minifingerprints" (MFPs) were designed to account for the presence or absence of structural fragments and/or aromatic character, flexibility, and hydrogen-bonding capacity of molecules. MFP design was based on an analysis of distributions of molecular descriptors and structural fragments in two large compound collections. The performance of different MFPs and a reference fingerprint was tested by systematic "one-against-all" similarity searches of molecules in a database containing 364 compounds with different biological activities. For each fingerprint, the most effective similarity cutoff value was determined. An MFP accounting for only 32 structural fragments showed less than 2% false positive similarity matches and correctly assigned on average approximately 40% of the compounds with the same biological activity to a query molecule. Inclusion of three numerical two-dimensional (2D) molecular descriptors increased the performance by 15%. This MFP performed better than a complex 2D fingerprint. At a similarity cutoff value of 0.85, the 2D fingerprint totally eliminated false positives but recognized less than 10% of the compounds within the same activity class.

Cyclooxygenase Inhibitors↗