PubMed Health⌕ Search

Biomedical subjects

Bruce S Williams

Publications and source records attributed to Bruce S Williams.

3 recordsLinked to original sources

Evaluating real-life high-throughput screening data.

High-throughput screening (HTS) is the result of a concerted effort of chemistry, biology, information technology, and engineering. Many factors beyond the biology of the assay influence the quality and outcome of the screening process, yet data analysis and quality control are often focused on the analysis of a limited set of control wells and the calculated values derived from these wells. Taking into account the large number of variables and the amount of data generated, multiple views of the screening data are necessary to guarantee quality and validity of HTS results. This article does not aim to give an exhaustive outlook on HTS data analysis but tries to illustrate the shortfalls of a reductionist approach focused on control wells and give examples for further analysis.

Biological Assay↗

Nonlinear prediction of quantitative structure-activity relationships.

Predicting the log of the partition coefficient P is a long-standing benchmark problem in Quantitative Structure-Activity Relationships (QSAR). In this paper we show that a relatively simple molecular representation (using 14 variables) can be combined with leading edge machine learning algorithms to predict logP on new compounds more accurately than existing benchmark algorithms which use complex molecular representations.

Journal Article↗

Data visualization during the early stages of drug discovery.

Multidimensional compound optimization is a new paradigm in the drug discovery process, yielding efficiencies during early stages and reducing attrition in the later stages of drug development. The success of this strategy relies heavily on understanding this multidimensional data and extracting useful information from it. This paper demonstrates how principled visualization algorithms can be used to understand and explore a large data set created in the early stages of drug discovery. The experiments presented are performed on a real-world data set comprising biological activity data and some whole-molecular physicochemical properties. Data visualization is a popular way of presenting complex data in a simpler form. We have applied powerful principled visualization methods, such as generative topographic mapping (GTM) and hierarchical GTM (HGTM), to help the domain experts (screening scientists, chemists, biologists, etc.) understand and draw meaningful decisions. We also benchmark these principled methods against relatively better known visualization approaches, principal component analysis (PCA), Sammon's mapping, and self-organizing maps (SOMs), to demonstrate their enhanced power to help the user visualize the large multidimensional data sets one has to deal with during the early stages of the drug discovery process. The results reported clearly show that the GTM and HGTM algorithms allow the user to cluster active compounds for different targets and understand them better than the benchmarks. An interactive software tool supporting these visualization algorithms was provided to the domain experts. The tool facilitates the domain experts by exploration of the projection obtained from the visualization algorithms providing facilities such as parallel coordinate plots, magnification factors, directional curvatures, and integration with industry standard software.

Drug Design↗