PubMed Health⌕ Search

PubMed · 11343364

Logistic regression when binary predictor variables are highly correlated.

Abstract

Standard logistic regression can produce estimates having large mean square error when predictor variables are multicollinear. Ridge regression and principal components regression can reduce the impact of multicollinearity in ordinary least squares regression. Generalizations of these, applicable in the logistic regression framework, are alternatives to standard logistic regression. It is shown that estimates obtained via ridge and principal components logistic regression can have smaller mean square error than estimates obtained through standard logistic regression. Recommendations for choosing among standard, ridge and principal components logistic regression are developed. Published in 2001 by John Wiley & Sons, Ltd.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

L Barker, C Brown. Logistic regression when binary predictor variables are highly correlated.. https://doi.org/10.1002/sim.680

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Generating correlated data for omics simulation.

Simulation of realistic omics data is a key input for benchmarking studies that help users obtain optimal computational pipelines. Omics data involves large numbers of measured features on each sample and these measures are generally correlated with each other. However, simulation too often ignores these correlations, perhaps due to computational and statistical hurdles of doing so. To alleviate this, we describe three approaches for generating omics-scale data with correlated measures which mimic real datasets. These approaches are all based on a Gaussian copula approach with a covariance matrix that decomposes into a diagonal part and a low-rank part. This decomposition allows for extremely efficient simulation, overcoming a hurdle for adoption of past methods. We use these approaches to demonstrate the importance of including correlation in two benchmarking applications. First, we show that variance of results from the popular DESeq2 method increases when dependence is included. Second, we demonstrate that CYCLOPS, a method for inferring circadian time of collection from transcriptomics, improves in performance when given gene-gene dependencies in some circumstances. We provide an R package, dependentsimr, that has efficient implementations of these methods and can generate dependent data with arbitrary marginal distributions, including discrete (binary, ordered categorical, Poisson, negative binomial), continuous (normal), or with an empirical distribution.

Computer Simulation↗

Addressing current challenges in cancer immunotherapy with mathematical and computational modelling.

The goal of cancer immunotherapy is to boost a patient's immune response to a tumour. Yet, the design of an effective immunotherapy is complicated by various factors, including a potentially immunosuppressive tumour microenvironment, immune-modulating effects of conventional treatments and therapy-related toxicities. These complexities can be incorporated into mathematical and computational models of cancer immunotherapy that can then be used to aid in rational therapy design. In this review, we survey modelling approaches under the umbrella of the major challenges facing immunotherapy development, which encompass tumour classification, optimal treatment scheduling and combination therapy design. Although overlapping, each challenge has presented unique opportunities for modellers to make contributions using analytical and numerical analysis of model outcomes, as well as optimization algorithms. We discuss several examples of models that have grown in complexity as more biological information has become available, showcasing how model development is a dynamic process interlinked with the rapid advances in tumour-immune biology. We conclude the review with recommendations for modellers both with respect to methodology and biological direction that might help keep modellers at the forefront of cancer immunotherapy development.

Computer Simulation↗

Effect of rate of chemical or thermal renaturation on refolding and aggregation of a simple lattice protein.

We used dynamic Monte Carlo simulation to investigate how changing the rate of chemical or thermal renaturation affects the folding and aggregation behavior of a system of simple, two-dimensional lattice protein molecules. Four renaturation methods were simulated: infinitely slow cooling; slow but finite cooling; quenching; and pulse renaturation. The infinitely slow cooling method, which is equivalent to dialysis or diafiltration, provides refolding yields that are relatively high and aggregates that are relatively small (mostly dimers or trimers). The slow but finite cooling method, which is equivalent to multiple-step dilution, provides refolding yields that are almost as high as those observed in the infinitely slow cooling case, but in a relatively short period of time. Quenching, which is equivalent to one-step dilution or quick quenching, is extremely slow and has low re- folding yields. A maximum appears in the refolding yield as a function of denaturant concentration in the simulation but disappears after a very long duration. Finally, the pulse renaturation method provides refolding yields that are substantially higher than those observed in the other three methods, even at high packing fractions. As in the early stages of quenching, there is a maximum in the refolding yield as a function of denaturant concentration when relatively large numbers of denatured chains are added to the refolding solution at each step.

Computer Simulation↗