PubMed Health⌕ Search

PubMed · 14709437

Advanced statistics: linear regression, part II: multiple linear regression.

Abstract

The applications of simple linear regression in medical research are limited, because in most situations, there are multiple relevant predictor variables. Univariate statistical techniques such as simple linear regression use a single predictor variable, and they often may be mathematically correct but clinically misleading. Multiple linear regression is a mathematical technique used to model the relationship between multiple independent predictor variables and a single dependent outcome variable. It is used in medical research to model observational data, as well as in diagnostic and therapeutic studies in which the outcome is dependent on more than one factor. Although the technique generally is limited to data that can be expressed with a linear function, it benefits from a well-developed mathematical framework that yields unique solutions and exact confidence intervals for regression coefficients. Building on Part I of this series, this article acquaints the reader with some of the important concepts in multiple regression analysis. These include multicollinearity, interaction effects, and an expansion of the discussion of inference testing, leverage, and variable transformations to multivariate models. Examples from the first article in this series are expanded on using a primarily graphic, rather than mathematical, approach. The importance of the relationships among the predictor variables and the dependence of the multivariate model coefficients on the choice of these variables are stressed. Finally, concepts in regression model building are discussed.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Keith A Marill. 2004. Advanced statistics: linear regression, part II: multiple linear regression.. https://doi.org/10.1197/j.aem.2003.09.006

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Assessment of blinding in pharmacotherapy and noninvasive neuromodulation randomized controlled trials for neuropathic pain in adults.

In randomized controlled trials (RCTs), study participants and research personnel are often blinded to minimize biases related to knowing treatment allocation. To determine if blinding was effective, participants may be asked which treatment they believe they received ("treatment guess"). This descriptive review characterized blinding assessment (BA) reporting in pharmacotherapy and neuromodulation neuropathic pain RCTs. Of 288 papers, 36 (12.5%) reported a BA. One paper reported the results of 2 studies, so in total 37 studies with a BA were assessed. Of these, 19 were crossover, 17 parallel, and 1 partial crossover in design. All 37 studies assessed participant blinding, and 10 also assessed investigator blinding. Approximately 27% included an "unsure" answer option for treatment guess, and 38% asked the reason for the guess. There were no clear patterns in BA reporting across time nor based on treatment type. Seventeen trials provided sufficient data to calculate Bang Blinding Index (BI) to determine blinding success. Participants remained blinded (BI = 0 &#xb1; 0.2) in 10/17 placebo and 10/17 treatment arms, 6 placebo and 5 treatment arms had a BI > 0.2 suggesting possible unblinding, whereas 1 placebo and 2 treatment arms had a BI < -0.2 suggesting misinformed guessing. Overall, we found that BAs are done in a minority of published neuropathic pain trials and with variable methodology. Given the importance of minimizing risk of bias because of treatment unblinding, future studies should consider including BAs, and further consensus building is necessary to determine if and how BAs should be conducted and interpreted in analgesic clinical trials.

Bias↗

Bias due to aggregation of individual covariates in the Cox regression model.

The impact of covariate aggregation, well studied in relation to linear regression, is less clear in the Cox model. In this paper, the authors use real-life epidemiologic data to illustrate how aggregating individual covariate values may lead to important underestimation of the exposure effect. The issue is then systematically assessed through simulations, with six alternative covariate representations. It is shown that aggregation of important predictors results in a systematic bias toward the null in the Cox model estimate of the exposure effect, even if exposure and predictors are not correlated. The underestimation bias increases with increasing strength of the covariate effect and decreasing censoring and, for a strong predictor and moderate censoring, may exceed 20%, with less than 80% coverage of the 95% confidence interval. However, covariate aggregation always induces smaller bias than covariate omission does, even if the two phenomena are shown to be related. The impact of covariate aggregation, but not omission, is independent of the covariate-exposure correlation. Simulations involving time-dependent aggregates demonstrate that bias results from failure of the baseline covariate mean to account for nonrandom changes over time in the risk sets and suggest a simple approach that may reduce the bias if individual data are available but have to be aggregated.

Bias↗