PubMed Health⌕ Search

PubMed · 14768902

Validation subset selections for extrapolation oriented QSPAR models.

Abstract

One of the most important features of QSPAR models is their predictive ability. The predictive ability of QSPAR models should be checked by external validation. In this work we examined three different types of external validation set selection methods for their usefulness in in-silico screening. The usefulness of the selection methods was studied in such a way that: 1) We generated thousands of QSPR models and stored them in 'model banks'. 2) We selected a final top model from the model banks based on three different validation set selection methods. 3) We predicted large data sets, which we called 'chemical universe sets', and calculated the corresponding SEPs. The models were generated from small fractions of the available water solubility data during a GA Variable Subset Selection procedure. The external validation sets were constructed by random selections, uniformly distributed selections or by perimeter-oriented selections. We found that the best performing models on the perimeter-oriented external validation sets usually gave the best validation results when the remaining part of the available data was overwhelmingly large, i.e., when the model had to make a lot of extrapolations. We also compared the top final models obtained from external validation set selection methods in three independent and different sizes of 'chemical universe sets'.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Csaba Szántai-Kis, István Kövesdi, György Kéri, László Orfi. 2003. Validation subset selections for extrapolation oriented QSPAR models.. https://doi.org/10.1023/b%3Amodi.0000006538.99122.00

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation↗

Conformational sampling and dynamics of membrane proteins from 10-nanosecond computer simulations.

In the current report, we provide a quantitative analysis of the convergence of the sampling of conformational space accomplished in molecular dynamics simulations of membrane proteins of duration in the order of 10 nanoseconds. A set of proteins of diverse size and topology is considered, ranging from helical pores such as gramicidin and small beta-barrels such as OmpT, to larger and more complex structures such as rhodopsin and FepA. Principal component analysis of the C(alpha)-atom trajectories was employed to assess the convergence of the conformational sampling in both the transmembrane domains and the whole proteins, while the time-dependence of the average structure was analyzed to obtain single-domain information. The membrane-embedded regions, particularly those of small or structurally simple proteins, were found to achieve reasonable convergence. By contrast, extra-membranous domains lacking secondary structure are often markedly under-sampled, exhibiting a continuous structural drift. This drift results in a significant imprecision in the calculated B-factors, which detracts from any quantitative comparison to experimental data. In view of such limitations, we suggest that similar analyses may be valuable in simulation studies of membrane protein dynamics, in order to attach a level of confidence to any biologically relevant observations.

Computer Simulation↗