PubMed Health⌕ Search

Biomedical subjects

David Hartsough

Publications and source records attributed to David Hartsough.

2 recordsLinked to original sources

Surrogate docking: structure-based virtual screening at high throughput speed.

Structure-based screening using fully flexible docking is still too slow for large molecular libraries. High quality docking of a million molecule library can take days even on a cluster with hundreds of CPUs. This performance issue prohibits the use of fully flexible docking in the design of large combinatorial libraries. We have developed a fast structure-based screening method, which utilizes docking of a limited number of compounds to build a 2D QSAR model used to rapidly score the rest of the database. We compare here a model based on radial basis functions and a Bayesian categorization model. The number of compounds that need to be actually docked depends on the number of docking hits found. In our case studies reasonable quality models are built after docking of the number of molecules containing approximately 50 docking hits. The rest of the library is screened by the QSAR model. Optionally a fraction of the QSAR-prioritized library can be docked in order to find the true docking hits. The quality of the model only depends on the training set size - not on the size of the library to be screened. Therefore, for larger libraries the method yields higher gain in speed no change in performance. Prioritizing a large library with these models provides a significant enrichment with docking hits: it attains the values of approximately 13 and approximately 35 at the beginning of the score-sorted libraries in our two case studies: screening of the NCI collection and a combinatorial libraries on CDK2 kinase structure. With such enrichments, only a fraction of the database must actually be docked to find many of the true hits. The throughput of the method allows its use in screening of large compound collections and in the design of large combinatorial libraries. The strategy proposed has an important effect on efficiency but does not affect retrieval of actives, the latter being determined by the quality of the docking method itself.

Bayes Theorem↗

Computational identification of proteins for selectivity assays.

At the stage of optimization of a chemical series the compounds are normally assayed for binding or inhibition on the target protein as well as on several proteins from a selectivity panel. These proteins are normally identified on the basis of sequence homology to the target protein. Experimental selectivity data are also taken into account if available. Cases when a nonhomologous protein has a significant affinity to the compound series are going to be missed if the selectivity panel is identified by homology. Experimental data is usually either unavailable or limited to a small fraction of proteins that should be considered. We have developed a computational method of identification of selectivity panel proteins. It is based on the evaluation of binding site similarity to the target protein using docking scores of target-selected molecular probes. These probes are obtained by docking a large library of drug-like compounds to the target protein followed by selecting a diverse subset from the best virtual binders. Docking scores of these probes to other proteins measure binding site similarity to the target. Because the method does not require prior knowledge of either affinities or structures of inhibitors for the target, it can be applied to any protein with known 3D structure. Validation of the method includes rediscovery of nonhomologous proteins that bind common ligands: estradiol, tamoxifen, and riboflavin. Given 3D structures, the method can effectively discriminate proteins with similar binding sites from random proteins independent of sequence homology.

Algorithms↗