PubMed Health⌕ Search

PubMed · 16363875

Protein structure evaluation using an all-atom energy based empirical scoring function.

Abstract

Arriving at the native conformation of a polypeptide chain characterized by minimum most free energy is a problem of long standing interest in protein structure prediction endeavors. Owing to the computational requirements in developing free energy estimates, scoring functions--energy based or statistical--have received considerable renewed attention in recent years for distinguishing native structures of proteins from non-native like structures. Several cleverly designed decoy sets, CASP (Critical Assessment of Techniques for Protein Structure Prediction) structures and homology based internet accessible three dimensional model builders are now available for validating the scoring functions. We describe here an all-atom energy based empirical scoring function and examine its performance on a wide series of publicly available decoys. Barring two protein sequences where native structure is ranked second and seventh, native is identified as the lowest energy structure in 67 protein sequences from among 61,659 decoys belonging to 12 different decoy sets. We further illustrate a potential application of the scoring function in bracketing native-like structures of two small mixed alpha/beta globular proteins starting from sequence and secondary structural information. The scoring function has been web enabled at www.scfbio-iitd.res.in/utility/proteomics/energy.jsp.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Pooja Narang, Kukum Bhushan, Surojit Bose, B Jayaram. 2006. Protein structure evaluation using an all-atom energy based empirical scoring function.. https://doi.org/10.1080/07391102.2006.10531234

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation↗

Twin concordances test for ascertained trichotomous traits data.

In human genetics, twin studies are widely underwent for investigation of the genetic influence on diseases. There are several measures that had been proposed to evaluate the similarity between two twins for dichotomous traits when the twins are sampled at random. These measures include the correlations, odds ratios, and casewise and pairwise concordances. When data are sampled through an ascertainment procedure, truncated data are formed. Under this circumstance, odds ratio cannot be defined and correlations cannot be correctly estimated. However, concordance measures for dichotomous traits can still be estimated using the likelihood method. On the other hand, though theoretically concordance measures can be extended to trichotomous traits, how to define them and to derive their estimators for ascertained trichotomous traits data have not been thoroughly discussed. In this study, we aim to address several relevant issues for ascertained trichotomous traits data. We define two new (casewise and pairwise) concordance measures for trichotomous traits and demonstrate how to apply a so-called 'self-contained subsets method (SCSM)' to estimation of twin concordances for ascertained data. We show that this method can obtain the same estimates as the likelihood method in an easier way and derive the asymptotic variances of the SCSM estimates under ascertainment. We establish the testing procedure for test of the equality of concordance measures between monozygotic twin pairs and dizygotic twin pairs, and illustrate the methods with a real data set and conduct Monte Carlo simulation to investigate its power performance.

Computer Simulation↗