PubMed Health⌕ Search

PubMed · 10614024

Analysis of a large structure/biological activity data set using recursive partitioning.

Abstract

Combinatorial chemistry and high-throughput screening are revolutionizing the process of lead discovery in the pharmaceutical industry. Large numbers of structures and vast quantities of biological assay data are quickly being accumulated, overwhelming traditional structure/activity relationship (SAR) analysis technologies. Recursive partitioning is a method for statistically determining rules that classify objects into similar categories or, in this case, structures into groups of molecules with similar potencies. SCAM is a computer program implemented to make extremely efficient use of this methodology. Depending on the size of the data set, rules explaining biological data can be determined interactively. An example data set of 1650 monoamine oxidase inhibitors exemplifies the method, yielding substructural rules and leading to general classifications of these inhibitors. The method scales linearly with the number of descriptors, so hundreds of thousands of structures can be analyzed utilizing thousands to millions of molecular descriptors. There are currently no methods to deal with statistical analysis problems of this size. An important aspect of this analysis is the ability to deal with mixtures, i.e., identify SAR rules for classes of compounds in the same data set that might be binding in different ways. Most current quantitative structure/activity relationship methods require that the compounds follow a single mechanism. Advantages and limitations of this methodology are presented.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

A Rusinko, M W Farmen, C G Lambert, P L Brown, S S Young. Analysis of a large structure/biological activity data set using recursive partitioning.. https://doi.org/10.1021/ci9903049

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

New solid support for the synthesis of 3'-oligonucleotide conjugates through glyoxylic oxime bond formation.

A novel solid support 1 was synthesized to incorporate glyoxylic aldehyde functionality at the oligonucleotide 3'-terminus. 6-mer and 11-mer oligonucleotide sequences containing 3'-glyoxylic aldehyde functionality were prepared by using this support. These modified oligonucleotides were coupled to reporters containing an aminooxy group to prepare oligonucleotide 3'-conjugates through glyoxylic oxime bond formation. The hydrolytic stability of a glyoxylic oxime linkage was also investigated. [reaction: see text].

Combinatorial Chemistry Techniques↗

Chemical genetics: an evolving toolbox for target identification and lead optimization.

Chemical genetics combines chemistry with biology as a means of exploring the function of unknown proteins or identifying the proteins responsible for a particular phenotype. Chemical genetics is thus a valuable tool in the identification of novel drug targets. This chapter describes the application of chemical genetics in traditional and systems-based approaches to drug target discovery and the tools/approaches that appear most promising for guiding future pharmaceutical development.

Combinatorial Chemistry Techniques↗

Protein library design and screening: working out the probabilities.

In designing protein libraries for selection, we must coordinate our capacity to create a large diversity of protein variants with the physical limitations of what we can actually screen. This chapter aims to bring the language of probabilities into the protein engineer's laboratory to answer some of our common questions: How can we most efficiently design a library? What fraction of the theoretical library diversity have we actually sampled at the end of the day? What is the probability of missing an individual of the library? Are the mutations present in the variants we have selected statistically meaningful or the product of random variation? The computation of these criteria throughout the process of experimental protein engineering will enable us to better design and evaluate the products of our libraries of protein variants.

Combinatorial Chemistry Techniques↗