PubMed Health⌕ Search

Biomedical subjects

Sukjoon Yoon

Publications and source records attributed to Sukjoon Yoon.

6 recordsLinked to original sources

Surrogate docking: structure-based virtual screening at high throughput speed.

Structure-based screening using fully flexible docking is still too slow for large molecular libraries. High quality docking of a million molecule library can take days even on a cluster with hundreds of CPUs. This performance issue prohibits the use of fully flexible docking in the design of large combinatorial libraries. We have developed a fast structure-based screening method, which utilizes docking of a limited number of compounds to build a 2D QSAR model used to rapidly score the rest of the database. We compare here a model based on radial basis functions and a Bayesian categorization model. The number of compounds that need to be actually docked depends on the number of docking hits found. In our case studies reasonable quality models are built after docking of the number of molecules containing approximately 50 docking hits. The rest of the library is screened by the QSAR model. Optionally a fraction of the QSAR-prioritized library can be docked in order to find the true docking hits. The quality of the model only depends on the training set size - not on the size of the library to be screened. Therefore, for larger libraries the method yields higher gain in speed no change in performance. Prioritizing a large library with these models provides a significant enrichment with docking hits: it attains the values of approximately 13 and approximately 35 at the beginning of the score-sorted libraries in our two case studies: screening of the NCI collection and a combinatorial libraries on CDK2 kinase structure. With such enrichments, only a fraction of the database must actually be docked to find many of the true hits. The throughput of the method allows its use in screening of large compound collections and in the design of large combinatorial libraries. The strategy proposed has an important effect on efficiency but does not affect retrieval of actives, the latter being determined by the quality of the docking method itself.

Bayes Theorem↗

Rapid assessment of contact-dependent secondary structure propensity: relevance to amyloidogenic sequences.

We have previously demonstrated that calculation of contact-dependent secondary structure propensity (CSSP) is highly sensitive in detecting non-native beta-strand propensities in the core sequences of known amyloidogenic proteins. Here we describe a CSSP method based on an artificial neural network that rapidly and accurately quantifies the influence of tertiary contacts (TCs) on secondary structure propensity in local regions of protein sequences. The present method exhibited 72% accuracy in predicting the alternate secondary structure adopted by chameleon sequences located in highly disparate TC regions. Analysis of 1930 nonhomologous protein domains reveals that the alpha-helix and the beta-strand largely share the same sequence context, and that tertiary context is a major determinant of the native conformation. Conversely, it appears that the propensity of random coils for either the alpha-helix or the beta-strand is largely invariant to tertiary effects. The present CSSP method successfully reproduced the amyloidogenic character observed in local regions of the human islet amyloid polypeptide (hIAPP). Furthermore, CSSP profiles were strongly correlated (r = 0.76) with the observed mutational effects on the aggregation rate of acylphosphatase. Taken together, these results provide compelling evidence in support of the present CSSP approach as a sensitive probe useful for analysis of full-length proteins and for detection of core sequences that may trigger amyloid fibril formation. The combined speed and simplicity of the CSSP method lends itself to proteome-wide analysis of the amyloidogenic nature of common proteins.

Algorithms↗

Computational identification of proteins for selectivity assays.

At the stage of optimization of a chemical series the compounds are normally assayed for binding or inhibition on the target protein as well as on several proteins from a selectivity panel. These proteins are normally identified on the basis of sequence homology to the target protein. Experimental selectivity data are also taken into account if available. Cases when a nonhomologous protein has a significant affinity to the compound series are going to be missed if the selectivity panel is identified by homology. Experimental data is usually either unavailable or limited to a small fraction of proteins that should be considered. We have developed a computational method of identification of selectivity panel proteins. It is based on the evaluation of binding site similarity to the target protein using docking scores of target-selected molecular probes. These probes are obtained by docking a large library of drug-like compounds to the target protein followed by selecting a diverse subset from the best virtual binders. Docking scores of these probes to other proteins measure binding site similarity to the target. Because the method does not require prior knowledge of either affinities or structures of inhibitors for the target, it can be applied to any protein with known 3D structure. Validation of the method includes rediscovery of nonhomologous proteins that bind common ligands: estradiol, tamoxifen, and riboflavin. Given 3D structures, the method can effectively discriminate proteins with similar binding sites from random proteins independent of sequence homology.

Algorithms↗

Improved method for predicting beta-turn using support vector machine.

MOTIVATION: Numerous methods for predicting beta-turns in proteins have been developed based on various computational schemes. Here, we introduce a new method of beta-turn prediction that uses the support vector machine (SVM) algorithm together with predicted secondary structure information. Various parameters from the SVM have been adjusted to achieve optimal prediction performance. RESULTS: The SVM method achieved excellent performance as measured by the Matthews correlation coefficient (MCC = 0.45) using a 7-fold cross validation on a database of 426 non-homologous protein chains. To our best knowledge, this MCC value is the highest achieved so far for predicting beta-turn. The overall prediction accuracy Qtotal was 77.3%, which is the best among the existing prediction methods. Among its unique attractive features, the present SVM method avoids overtraining and compresses information and provides a predicted reliability index.

Algorithms↗

Detecting hidden sequence propensity for amyloid fibril formation.

The preponderance of evidence implicates protein misfolding in many unrelated human diseases. In all cases, normal correctly folded proteins transform from their proper native structure into an abnormal beta-rich structure known as amyloid fibril. Here we introduce a computational algorithm to detect nonnative (hidden) sequence propensity for amyloid fibril formation. Analyzing sequence-structure relationships in terms of tertiary contact (TC), we find that the hidden beta-strand propensity of a query local sequence can be quantitatively estimated from the secondary structure preferences of template sequences of known secondary structure found in regions of high TC. The present method correctly pinpoints the minimal peptide fragment shown experimentally as the likely local mediator of amyloid fibril formation in beta-amyloid peptide, islet amyloid polypeptide (hIAPP), alpha-synuclein, and human acetylcholinesterase (AChE). It also found previously unrecognized beta-strand propensities in the prototypical helical protein myoglobin that has been reported as amyloidogenic. Analysis of 2358 nonhomologous protein domains provides compelling evidence that most proteins contain sequences with significant hidden beta-strand propensity. The present method may find utility in many medically relevant applications, such as the engineering of protein sequences and the discovery of therapeutic agents that specifically target these sequences for the prevention and treatment of amyloid diseases.

Acetylcholinesterase↗

Identification of a minimal subset of receptor conformations for improved multiple conformation docking and two-step scoring.

Docking and scoring are critical issues in virtual drug screening methods. Fast and reliable methods are required for the prediction of binding affinity especially when applied to a large library of compounds. The implementation of receptor flexibility and refinement of scoring functions for this purpose are extremely challenging in terms of computational speed. Here we propose a knowledge-based multiple-conformation docking method that efficiently accommodates receptor flexibility thus permitting reliable virtual screening of large compound libraries. Starting with a small number of active compounds, a preliminary docking operation is conducted on a large ensemble of receptor conformations to select the minimal subset of receptor conformations that provides a strong correlation between the experimental binding affinity (e.g., Ki, IC50) and the docking score. Only this subset is used for subsequent multiple-conformation docking of the entire data set of library (test) compounds. In conjunction with the multiple-conformation docking procedure, a two-step scoring scheme is employed by which the optimal scoring geometries obtained from the multiple-conformation docking are re-scored by a molecular mechanics energy function including desolvation terms. To demonstrate the feasibility of this approach, we applied this integrated approach to the estrogen receptor alpha (ERalpha) system for which published binding affinity data were available for a series of structurally diverse chemicals. The statistical correlation between docking scores and experimental values was significantly improved from those of single-conformation dockings. This approach led to substantial enrichment of the virtual screening conducted on mixtures of active and inactive ERalpha compounds.

Protein Binding↗