PubMed Health⌕ Search

PubMed · 11373721

Incorporating knowledge-based biases into an energy-based side-chain modeling method: application to comparative modeling of protein structure.

Abstract

The performance of the self-consistent mean field theory (SCMFT) method for side-chain modeling, employing rotamer energies calculated with the flexible rotamer model (FRM), is evaluated in the context of comparative modeling of protein structure. Predictions were carried out on a test set of 56 model backbones of varying accuracy, to allow side-chain prediction accuracy to be analyzed as a function of backbone accuracy. A progressive decrease in the accuracy of prediction was observed as backbone accuracy decreased. However, even for very low backbone accuracy, prediction was substantially higher than random, indicating that the FRM can, in part, compensate for the errors in the modeled tertiary environment. It was also investigated whether the introduction in the FRM-SCMFT method of knowledge-based biases, derived from a backbone-dependent rotamer library, could enhance its performance. A bias derived from the backbone-dependent rotamer conformations alone did not improve prediction accuracy. However, a bias derived from the backbone-dependent rotamer probabilities improved prediction accuracy considerably. This bias was incorporated through two different strategies. In one (the indirect strategy), rotamer probabilities were used to reject unlikely rotamers a priori, thus restricting prediction by FRM-SCMFT to a subset containing only the most probable rotamers in the library. In the other (the direct strategy), rotamer energies were transformed into pseudo-energies that were added to the average potential energies of the respective rotamers, thereby creating hybrid energy-based/knowledge-based average rotamer energies, which were used by the FRM-SCMFT method for prediction. For all degrees of backbone accuracy, an optimal strength of the knowledge-based bias existed for both strategies for which predictions were more accurate than pure energy-based predictions, and also than pure knowledge-based predictions. Hybrid knowledge-based/energy-based methods were obtained from both strategies and compared with the SCWRL method, a hybrid method based on the same backbone-dependent rotamer library. The accuracy of the indirect method was approximately the same as that of the SCWRL method, but that of the direct method was significantly higher.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

J Mendes, H A Nagarajaram, C M Soares, T L Blundell, M A Carrondo. 2001. Incorporating knowledge-based biases into an energy-based side-chain modeling method: application to comparative modeling of protein structure.. https://doi.org/10.1002/1097-0282(200108)59%3A2%3C72%3A%3Aaid-bip1007%3E3.0.co%3B2-s

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation↗

Theoretical distribution of truncation lengths in incremental truncation libraries.

Incremental truncation is a method for constructing libraries of every one base pair truncation of a segment of DNA. Incremental truncation libraries can be created using a time-dependent nuclease method or through the incorporation of alpha-phosphothioate dNTPs by PCR or by primer extension (THIO(pcr) truncation and THIO(extension) truncation, respectively). Libraries created by the fusion of two truncation libraries, known as ITCHY libraries, can be created using the above methods or by the incremental truncation-like method SHIPREC. Knowing and being able to tailor the distribution of truncations in incremental truncation, ITCHY and SHIPREC libraries would be beneficial for their use in protein engineering and other applications. However, the experimental determination of the distributions would require extensive, cost-prohibitive, DNA sequencing to obtain statistically relevant data. Instead, a theoretical prediction of the distributions was developed. Time-dependent incremental truncation libraries had the most uniform distribution of truncation lengths, but were biased against longer truncations. Essentially uniform distribution over the desired truncation range (from zero to N(max) base pairs) required that truncations be prepared up to at least 1.2-1.5 N(max). THIO(pcr) and THIO(extension) truncation libraries had a very nonuniform distribution of truncation lengths with a bias against longer truncations. Such nonuniformity could be significantly diminished by decreasing the incorporation rate of alphaS-dNTPs but at the expense of having a large fraction of the DNA truncated beyond the desired range or completely degraded. ITCHY libraries created using time-dependent truncation had the most uniform distribution of possible fusions and had the highest fraction of the library being parental-length fusions. However, the distribution of parental-length fusions was biased against fusions near the beginning/ends of genes unless the truncation libraries are prepared with a uniform distribution up to N(max). In contrast, SHIPREC libraries and THIO(pcr) ITCHY libraries, by the very nature of the nonuniform distributions of the truncated DNA, are ensured of having a uniform distribution of fusion points in parental-length fusions. This comes at the expense of having a smaller fraction of the library being parental-length fusions; however, this limitation can be overcome by performing size selection on the library.

Computer Simulation↗