PubMed Health⌕ Search

Biomedical subjects

Ansgar Schuffenhauer

Publications and source records attributed to Ansgar Schuffenhauer.

14 recordsLinked to original sources

Quest for the rings. In silico exploration of ring universe to identify novel bioactive heteroaromatic scaffolds.

Bioactive molecules only contain a relatively limited number of unique ring types. To identify those ring properties and structural characteristics that are necessary for biological activity, a large virtual library of nearly 600 000 heteroaromatic scaffolds was created and characterized by calculated properties, including structural features, bioavailability descriptors, and quantum chemical parameters. A self-organizing neural network was used to cluster these scaffolds and to identify properties that best characterize bioactive ring systems. The analysis shows that bioactivity is very sparsely distributed within the scaffold property and structural space, forming only several relatively small, well-defined "bioactivity islands". Various possible applications of a large database of rings with calculated properties and bioactivity scores in the drug design and discovery process are discussed, including virtual screening, support for the design of combinatorial libraries, bioisosteric design, and scaffold hopping.

Biological Availability↗

A chemoinformatics analysis of hit lists obtained from high-throughput affinity-selection screening.

The high-throughput affinity-selection screening platform SpeedScreen was recently reported by the Novartis Institutes for BioMedical Research as a homogeneous, label-free screening technology with mass-spectrometry readout. SpeedScreen relies on the screening of compound mixtures with various target proteins and uses fast size-exclusion chromatography to separate target-bound from unbound substances. After disintegration of the target-binder complex, the binder molecules are identified by their molecular masses using liquid chromatography/mass spectrometry. The authors report an analysis of the molecular properties of hits obtained with SpeedScreen on 26 targets screened within the past few years at Novartis using this technology. Affinity-based SpeedScreen is a robust high-throughput screening technology that does not accumulate frequent hitters or potential covalent binders. The hits are representative of the most commonly identified scaffold classes observed for known drugs. Validated SpeedScreen hits tend to be enriched on more lipophilic and larger-molecular-weight compounds compared to the whole library. The potential for a reduced SpeedScreen screening set to be used in case only limited protein quantities are available is evaluated. Such a reduced compound set should also maximize the coverage of the high-performing regions of the chemical property and class spaces; chemoinformatics methods including genetic algorithms and divisive K-means clustering are used for this aim.

Chromatography, High Pressure Liquid↗

Charting biologically relevant chemical space: a structural classification of natural products (SCONP).

The identification of small molecules that fall within the biologically relevant subfraction of vast chemical space is of utmost importance to chemical biology and medicinal chemistry research. The prerequirement of biological relevance to be met by such molecules is fulfilled by natural product-derived compound collections. We report a structural classification of natural products (SCONP) as organizing principle for charting the known chemical space explored by nature. SCONP arranges the scaffolds of the natural products in a tree-like fashion and provides a viable analysis- and hypothesis-generating tool for the design of natural product-derived compound collections. The validity of the approach is demonstrated in the development of a previously undescribed class of selective and potent inhibitors of 11beta-hydroxysteroid dehydrogenase type 1 with activity in cells guided by SCONP and protein structure similarity clustering. 11beta-hydroxysteroid dehydrogenase type 1 is a target in the development of new therapies for the treatment of diabetes, the metabolic syndrome, and obesity.

11-beta-Hydroxysteroid Dehydrogenase Type 1↗

Enhancing the effectiveness of similarity-based virtual screening using nearest-neighbor information.

We test the hypothesis that fusing the outputs of similarity searches based on a single bioactive reference structure and on its nearest neighbors (of unknown activity) is more effective (in terms of numbers of high-ranked active structures) than a similarity search involving just the reference structure. This turbo similarity searching approach provides a simple way to enhance the effectiveness of simulated virtual screening searches of the MDL Drug Data Report database.

Computing Methodologies↗

Complex molecules: do they add value?

The concept of complexity in chemistry has received much interest in the research community. Various measures to assess molecular complexity have been published, ranging from abstract complexity definitions to very specific application-oriented definitions. In this article we focus on molecular complexity in relation to biological activity. Connectivity and feature-based structural descriptors have been evaluated with reference to their potential as complexity measures. Our goal was to discuss the potential of the complexity concept to support the drug discovery process, helping to design suitable lead candidates. The studies have shown that highly active compounds, on average, are more complex than inactive compounds. However, complexity must be balanced with other molecular properties because more complex molecules have a higher probability to exhibit pharmacokinetic problems.

Combinatorial Chemistry Techniques↗

Key aspects of the Novartis compound collection enhancement project for the compilation of a comprehensive chemogenomics drug discovery screening collection.

The NIBR (Novartis Institutes for BioMedical Research) compound collection enrichment and enhancement project integrates corporate internal combinatorial compound synthesis and external compound acquisition activities in order to build up a comprehensive screening collection for a modern drug discovery organization. The main purpose of the screening collection is to supply the Novartis drug discovery pipeline with hit-to-lead compounds for today's and the future's portfolio of drug discovery programs, and to provide tool compounds for the chemogenomics investigation of novel biological pathways and circuits. As such, it integrates designed focused and diversity-based compound sets from the synthetic and natural paradigms able to cope with druggable and currently deemed undruggable targets and molecular interaction modes. Herein, we will summarize together with new trends published in the literature, scientific challenges faced and key approaches taken at NIBR to match the chemical and biological spaces.

Animals↗

Library design for fragment based screening.

According to Hann's model of molecular complexity an increased probability of detection binding to a target protein can be expected when small, low complex molecular fragments are screened with high sensitivity instead of full-sized ligands with lower sensitivity. Analysis of the HTS summary data of Novartis and comparison with NMR screening results obtained on generic fragment libraries indicate this expectation to be true with hitrates of 0.001% - 0.151% observed in the identification of ligands with an IC(50) threshold in the micromolar range in an HTS setup and hitrates above or equal to 3% observed in NMR screening of fragments with an affinity threshold in the millimolar range. It is however necessary to keep in mind that the sets of target studied were not identical for both method and the experience in NMR screening is too limited for a final conclusion. The term hitrate as used here reflects only the success rate in the observation of ligand binding event. It must not be confused with the overall success rate of fragment and high throughput screening in the lead finding process, which can be entirely different, since the steps required to follow-up a ligand binding event to a lead are different for both methods. A survey of fragment-based lead discovery case studies given in the literature shows that in approximately half of the cases the initial hit fragment was discovered by screening a generic library, whereas in the other cases some knowledge about an initial ligands or the protein binding site has been used, whereas systematic virtual screening of fragment databases has been only rarely reported. As comparatively high hitrates were obtained, further consideration to optimize the generic fragment screening library were directed to the chemical tractability of the fragment. As several functional groups preferred by chemists for modification and linking of the fragments are also preferentially involved in interactions between the fragments and the target protein, a set of screening fragments was derived from chemical building blocks by masking its linker group by a chemical transformation which can be later on used in the chemical follow-up of the fragment hit. For example primary amines can be masked as acetamides. If the screening fragment is active the related building block can then be used for synthesis of a follow-up library.

Combinatorial Chemistry Techniques↗

Comparison of topological descriptors for similarity-based virtual screening using multiple bioactive reference structures.

This paper reports a detailed comparison of a range of different types of 2D fingerprints when used for similarity-based virtual screening with multiple reference structures. Experiments with the MDL Drug Data Report database demonstrate the effectiveness of fingerprints that encode circular substructure descriptors generated using the Morgan algorithm. These fingerprints are notably more effective than fingerprints based on a fragment dictionary, on hashing and on topological pharmacophores. The combination of these fingerprints with data fusion based on similarity scores provides both an effective and an efficient approach to virtual screening in lead-discovery programmes.

Algorithms↗

Chemogenomics knowledge-based strategies in drug discovery.

In the postgenomic age of drug discovery, targets can no longer be viewed as singular objects having no relationship to one another. All targets are now visible and the systematic exploration of selected target families appears to be a promising way to speed up and further industrialize target-based drug discovery. Chemogenomics refers to such systematic exploration of target families and aims to identify all possible ligands of all target families. Because biology works by applying prior knowledge to an unknown entity, chemogenomics approaches are expected to be especially effective within the previously well-explored target families, for which, in addition to the protein sequence and structure information, considerable knowledge of pharmacologically active structural classes and structure-activity relationships exists. For the new target families, chemical knowledge will have to be generated and beyond biological target validation, the emphasis is on chemistry to provide the molecules with which their novel biology and pharmacology can be studied. Using examples from the previously most successfully explored target families, the GPCR family in particular, we summarize herein our current chemogenomics knowledge-based strategies for drug discovery, which are founded on the high integration of chem and bioinformatics, thereby providing a molecular informatics frame for the exploration of the new target families.

Computational Biology↗

An ontology for pharmaceutical ligands and its application for in silico screening and library design.

Annotation efforts in biosciences have focused in past years mainly on the annotation of genomic sequences. Only very limited effort has been put into annotation schemes for pharmaceutical ligands. Here we propose annotation schemes for the ligands of four major target classes, enzymes, G protein-coupled receptors (GPCRs), nuclear receptors (NRs), and ligand-gated ion channels (LGICs), and outline their usage for in silico screening and combinatorial library design. The proposed schemes cover ligand functionality and hierarchical levels of target classification. The classification schemes are based on those established by the EC, GPCRDB, NuclearDB, and LGICDB. The ligands of the MDL Drug Data Report (MDDR) database serve as a reference data set of known pharmacologically active compounds. All ligands were annotated according to the schemes when attribution was possible based on the activity classification provided by the reference database. The purpose of the ligand-target classification schemes is to allow annotation-based searching of the ligand database. In addition, the biological sequence information of the target is directly linkable to the ligand, hereby allowing sequence similarity-based identification of ligands of next homologous receptors. Ligands of specified levels can easily be retrieved to serve as comprehensive reference sets for cheminformatics-based similarity searches and for design of target class focused compound libraries. Retrospective in silico screening experiments within the MDDR01.1 database, searching for structures binding to dopamine D2, all dopamine receptors and all amine-binding class A GPCRs using known dopamine D2 binding compounds as a reference set, have shown that such reference sets are in particular useful for the identification of ligands binding to receptors closely related to the reference system. The potential for ligand identification drops with increasing phylogenetic distance. The analysis of the focus of a tertiary amine based combinatorial library compared to known amine binding class A GPCRs, peptide binding class A GPCRs, and LGIC ligands constitutes a second application scenario which illustrates how the focus of a combinatorial library can be treated quantitatively. The provided annotation schemes, which bridge chem- and bioinformatics by linking ligands to sequences, are expected to be of key utility for further systematic chemogenomics exploration of previously well explored target families.

Combinatorial Chemistry Techniques↗

Similarity metrics for ligands reflecting the similarity of the target proteins.

In this study we evaluate how far the scope of similarity searching can be extended to identify not only ligands binding to the same target as the reference ligand(s) but also ligands of other homologous targets without initially known ligands. This "homology-based similarity searching" requires molecular representations reflecting the ability of a molecule to interact with target proteins. The Similog keys, which are introduced here as a new molecular representation, were designed to fulfill such requirements. They are based only on the molecular constitution and are counts of atom triplets. Each triplet is characterized by the graph distances and the types of its atoms. The atom-typing scheme classifies each atom by its function as H-bond donor or acceptor and by its electronegativity and bulkiness. In this study the Similog keys are investigated in retrospective in silico screening experiments and compared with other conformation independent molecular representations. Studied were molecules of the MDDR database for which the activity data was augmented by standardized target classification information from public protein classification databases. The MDDR molecule set was split randomly into two halves. The first half formed the candidate set. Ligands of four targets (dopamine D2 receptor, opioid delta-receptor, factor Xa serine protease, and progesterone receptor) were taken from the second half to form the respective reference sets. Different similarity calculation methods are used to rank the molecules of the candidate set by their similarity to each of the four reference sets. The accumulated counts of molecules binding to the reference target and groups of targets with decreasing homology to it were examined as a function of the similarity rank for each reference set and similarity method. In summary, similarity searching based on Unity 2D-fingerprints or Similog keys are found to be equally effective in the identification of molecules binding to the same target as the reference set. However, the application of the Similog keys is more effective in comparison with the other investigated methods in the identification of ligands binding to any target belonging to the same family as the reference target. We attribute this superiority to the fact that the Similog keys provide a generalization of the chemical elements and that the keys are counted instead of merely noting their presence or absence in a binary form. The second most effective molecular representation are the occurrence counts of the public ISIS key fragments, which like the Similog method, incorporates key counting as well as a generalization of the chemical elements. The results obtained suggest that ligands for a new target can be identified by the following three-step procedure: 1. Select at least one target with known ligands which is homologous to the new target. 2. Combine the known ligands of the selected target(s) to a reference set. 3. Search candidate ligands for the new targets by their similarity to the reference set using the Similog method. This clearly enlarges the scope of similarity searching from the classical application for a single target to the identification of candidate ligands for whole target families and is expected to be of key utility for further systematic chemogenomics exploration of previously well explored target families.

Algorithms↗

Comparison of fingerprint-based methods for virtual screening using multiple bioactive reference structures.

Fingerprint-based similarity searching is widely used for virtual screening when only a single bioactive reference structure is available. This paper reviews three distinct ways of carrying out such searches when multiple bioactive reference structures are available: merging the individual fingerprints into a single combined fingerprint; applying data fusion to the similarity rankings resulting from individual similarity searches; and approximations to substructural analysis. Extended searches on the MDL Drug Data Report database suggest that fusing similarity scores is the most effective general approach, with the best individual results coming from the binary kernel discrimination technique.

Molecular Structure↗

New methods for ligand-based virtual screening: use of data fusion and machine learning to enhance the effectiveness of similarity searching.

Similarity searching using a single bioactive reference structure is a well-established technique for accessing chemical structure databases. This paper describes two extensions of the basic approach. First, we discuss the use of group fusion to combine the results of similarity searches when multiple reference structures are available. We demonstrate that this technique is notably more effective than conventional similarity searching in scaffold-hopping searches for structurally diverse sets of active molecules; conversely, the technique will do little to improve the search performance if the actives are structurally homogeneous. Second, we make the assumption that the nearest neighbors resulting from a similarity search, using a single bioactive reference structure, are also active and use this assumption to implement approximate forms of group fusion, substructural analysis, and binary kernel discrimination. This approach, called turbo similarity searching, is notably more effective than conventional similarity searching.

Artificial Intelligence↗

Relationships between Molecular Complexity, Biological Activity, and Structural Diversity.

Following the theoretical model by Hann et al. moderately complex structures are preferable lead compounds since they lead to specific binding events involving the complete ligand molecule. To make this concept usable in practice for library design, we studied several complexity measures on the biological activity of ligand molecules. We applied the historical IC50/EC50 summary data of 160 assays run at Novartis covering a diverse range of targets, among them kinases, proteases, GPCRs, and protein-protein interactions, and compared this to the background of "inactive" compounds which have been screened for 2 years but have never shown any activity in any primary screen. As complexity measures we used the number of structural features present in various molecular fingerprints and descriptors. We found generally that with increasing activity of the ligands, their average complexity also increased, and we could therefore establish a minimum number of structural features in each descriptor needed for biological activity. Especially well suited in this context were the Similog keys and circular substructure fingerprints. These are those descriptors, which also perform especially well in the identification of bioactive compounds by similarity search, suggesting that structural features encoded in these descriptors have a high relevance for bioactivity. Since the number of features correlates with the number of atoms present in the molecule, also the number of atoms serves as a reasonable complexity measure and larger molecules have, in general, higher activities. Due to the relationship between feature counts and densities on one hand and biological activity on the other, the size bias present in almost all similarity coefficients becomes especially important. Diversity selections using these coefficients can influence the overall complexity of the resulting set of molecules, which has an impact on the biological activity that they exhibit. Using sphere-exclusion based diversity selection methods, such as OptiSim together with the Tanimoto dissimilarity, the average feature count distribution of the resulting selections is shifted toward lower complexity than that of the original set, particularly when applying tight diversity constraints. This size bias reduces the fraction of molecules in the subsets having the complexity required for a high, submicromolar activity. None of the diversity selection methods studied, namely OptiSim, divisive K-means clustering, and self-organizing maps, yielded subsets covering the activity space of the IC50 summary data set better than subsets selected randomly.

Drug Design↗