PubMed Health⌕ Search

Biomedical subjects

D K Agrafiotis

Publications and source records attributed to D K Agrafiotis.

8 recordsLinked to original sources

Combinatorial networks.

A novel approach for the analysis and virtual screening of large combinatorial libraries is presented. The method attempts to relieve the computational burden by computing the properties of the products in a way that does not require their explicit enumeration. In particular, a small subset of compounds from the virtual library is identified and their descriptors are calculated in a conventional manner. The resulting data is used as input to a multilayer perceptron, which is trained to predict the descriptors of the products from the descriptors of their respective building blocks. Once trained, the neural network is able to estimate the descriptors of the remaining members of the virtual library with remarkable accuracy, without ever, generating their connection tables. This method eliminates the two most time-consuming steps in virtual screening and allows the processing of very large combinatorial libraries that are intractable with conventional techniques.

Chemistry Techniques, Analytical↗

A new method for analyzing protein sequence relationships based on Sammon maps.

Recent advances in gene sequencing and rational drug design have re-emphasized the need for new methods for protein analysis, classification, and structure and function prediction. In this article, we introduce a new method for analyzing protein sequences based on Sammon's non-linear mapping algorithm. When applied to a family of homologous sequences, the method is able to capture the essential features of the similarity matrix, and provides a faithful representation of chemical or evolutionary distance in a simple and intuitive way. The merits of the new algorithm are demonstrated using examples from the protein kinase family.

Algorithms↗

Kolmogorov-Smirnov statistic and its application in library design.

After several years of frantic development, the dream of an "ideal" library remains elusive. Traditionally, combinatorial chemistry has been used primarily for lead generation, and molecular diversity has been the method of choice for designing and prioritizing experiments. One aspect that often has been overlooked is the drug likeness of the resulting collections. Recently, there have been several attempts to quantify this concept and incorporate it directly into the design process. This article demonstrates the limitations of some conventional methodologies and proposes a new paradigm for experimental design based on the principles of multiobjective optimization. This method allows traditional design objectives such as diversity or similarity to be combined with secondary selection criteria in order to bias the selection toward more pharmacologically relevant regions of chemical space. The method is robust, general, and easily extensible, and it allows the medicinal chemist to create designs that represent the best compromise between several, often conflicting, objectives. Two types of designs are discussed (singles, arrays), and a novel criterion based on the Kolmogorov-Smirnov statistic is proposed as a means to enforce a particular distribution on key molecular properties that are related to drug likeness. The potential of this approach is illustrated in the design of an exploratory library based on the simultaneous optimization of five different parameters. These parameters are combined in an intuitive manner to produce a design that is sufficiently diverse, exhibits a molecular weight and logP profile that is consistent with the respective distributions of known drugs, requires a small number of reagents, and can be synthesized easily in array format using robotic hardware.

Databases, Factual↗

Nonlinear mapping networks.

Among the many dimensionality reduction techniques that have appeared in the statistical literature, multidimensional scaling and nonlinear mapping are unique for their conceptual simplicity and ability to reproduce the topology and structure of the data space in a faithful and unbiased manner. However, a major shortcoming of these methods is their quadratic dependence on the number of objects scaled, which imposes severe limitations on the size of data sets that can be effectively manipulated. Here we describe a novel approach that combines conventional nonlinear mapping techniques with feed-forward neural networks, and allows the processing of data sets orders of magnitude larger than those accessible with conventional methodologies. Rooted on the principle of probability sampling, the method employs a classical algorithm to project a small random sample, and then "learns" the underlying nonlinear transform using a multilayer neural network trained with the back-propagation algorithm. Once trained, the neural network can be used in a feed-forward manner to project the remaining members of the population as well as new, unseen samples with minimal distortion. Using examples from the fields of image processing and combinatorial chemistry, we demonstrate that this method can generate projections that are virtually indistinguishable from those derived by conventional approaches. The ability to encode the nonlinear transform in the form of a neural network makes nonlinear mapping applicable to a wide variety of data mining applications involving very large data sets that are otherwise computationally intractable.

Journal Article↗

A constant time algorithm for estimating the diversity of large chemical libraries.

We describe a novel diversity metric for use in the design of combinatorial chemistry and high-throughput screening experiments. The method estimates the cumulative probability distribution of intermolecular dissimilarities in the collection of interest and then measures the deviation of that distribution from the respective distribution of a uniform sample using the Kolmogorov-Smirnov statistic. The distinct advantage of this approach is that the cumulative distribution can be easily estimated using probability sampling and does not require exhaustive enumeration of all pairwise distances in the data set. The function is intuitive, very fast to compute, does not depend on the size of the collection, and can be used to perform diversity estimates on both global and local scale. More importantly, it allows meaningful comparison of data sets of different cardinality and is not affected by the curse of dimensionality, which plagues many other diversity indices. The advantages of this approach are demonstrated using examples from the combinatorial chemistry literature.

Journal Article↗

Design and prioritization of plates for high-throughput screening.

A general algorithm for the prioritization and selection of plates for high-throughput screening is presented. The method uses a simulated annealing algorithm to search through the space of plate combinations for the one that maximizes some user-defined objective function. The algorithm is robust and convergent, and permits the simultaneous optimization of multiple design objectives, including molecular diversity, similarity to known actives, predicted activity or binding affinity, and many others. It is shown that the arrangement of compounds among the plates may have important consequences on the ability to design a well-targeted and cost-effective experiment. To that end, two simple and effective schemes for the construction of homogeneous and heterogeneous plates are outlined, using a novel similarity sorting algorithm based on one-dimensional nonlinear mapping.

Algorithms↗

Advances in diversity profiling and combinatorial series design.

Rapid advances in synthetic and screening technology have recently enabled the simultaneous synthesis and biological evaluation of large chemical libraries containing hundreds to tens of thousands of compounds, using molecular diversity as a means to design and prioritize experiments. This paper reviews some of the most important computational work in the field of diversity profiling and combinatorial library design, with particular emphasis on methodology and applications. It is divided into four sections that address issues related to molecular representation, dimensionality reduction, compound selection, and visualization.

Chemistry, Pharmaceutical↗