PubMed Health⌕ Search

PubMed · 11604043

How does consensus scoring work for virtual library screening? An idealized computer experiment.

Abstract

It has been reported recently that consensus scoring, which combines multiple scoring functions in binding affinity estimation, leads to higher hit-rates in virtual library screening studies. This method seems quite independent to the target receptor, the docking program, or even the scoring functions under investigation. Here we present an idealized computer experiment to explore how consensus scoring works. A hypothetical set of 5000 compounds is used to represent a chemical library under screening. The binding affinities of all its member compounds are assigned by mimicking a real situation. Based on the assumption that the error of a scoring function is a random number in a normal distribution, the predicted binding affinities were generated by adding such a random number to the "observed" binding affinities. The relationship between the hit-rates and the number of scoring functions employed in scoring was then investigated. The performance of several typical ranking strategies for a consensus scoring procedure was also explored. Our results demonstrate that consensus scoring outperforms any single scoring for a simple statistical reason: the mean value of repeated samplings tends to be closer to the true value. Our results also suggest that a moderate number of scoring functions, three or four, are sufficient for the purpose of consensus scoring. As for the ranking strategy, both the rank-by-number and the rank-by-rank strategy work more effectively than the rank-by-vote strategy.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

R Wang, S Wang. How does consensus scoring work for virtual library screening? An idealized computer experiment.. https://doi.org/10.1021/ci010025x

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation↗

Computational analysis of the chiral action of type II DNA topoisomerases.

It was found recently that bacterial type II DNA topoisomerase, topo IV, is much more efficient in relaxing (+) DNA supercoiling than (-) supercoiling. This means that the DNA-enzyme complex is chiral. This chirality can appear upon binding the first segment that participates in the strand passing reaction (G segment) or only after the second segment (T segment) joins the complex. The former possibility is analyzed here. We assume that upon binding the enzyme, the G segment forms a part of left-handed helical turn. This model is an extension of the hairpin model introduced earlier to explain simplification of DNA topology by these enzymes. Using statistical-mechanical simulation of DNA properties, we estimated different consequences of the model: (1) relative rates of relaxation of (+) and (-) supercoiling by the enzyme; (2) the distribution of positions of the G segment in supercoiled molecules; (3) steady-state distribution of knots in circular molecules created by the topoisomerase; (4) the variance of topoisomer distribution created by the enzyme; (5) the effect of (+) and (-) supercoiling on the binding topo II with G segment. The simulation results are capable of explaining nearly all available experimental data, at least semiquantitatively. A few predictions obtained in the model analysis can be tested experimentally.

Computer Simulation↗