PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Linear Regression”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Advanced statistics: linear regression, part I: simple linear regression.

Simple linear regression is a mathematical technique used to model the relationship between a single independent predictor variable and a single dependent outcome variable. In this, the first of a two-part series exploring concepts in linear regression analysis, the four fundamental assumptions and the mechanics of simple linear regression are reviewed. The most common technique used to derive the regression line, the method of least squares, is described. The reader will be acquainted with other important concepts in simple linear regression, including: variable transformations, dummy variables, relationship to inference testing, and leverage. Simplified clinical examples with small datasets and graphic models are used to illustrate the points. This will provide a foundation for the second article in this series: a discussion of multiple linear regression, in which there are multiple predictor variables.

Biometry↗

Advanced statistics: linear regression, part II: multiple linear regression.

The applications of simple linear regression in medical research are limited, because in most situations, there are multiple relevant predictor variables. Univariate statistical techniques such as simple linear regression use a single predictor variable, and they often may be mathematically correct but clinically misleading. Multiple linear regression is a mathematical technique used to model the relationship between multiple independent predictor variables and a single dependent outcome variable. It is used in medical research to model observational data, as well as in diagnostic and therapeutic studies in which the outcome is dependent on more than one factor. Although the technique generally is limited to data that can be expressed with a linear function, it benefits from a well-developed mathematical framework that yields unique solutions and exact confidence intervals for regression coefficients. Building on Part I of this series, this article acquaints the reader with some of the important concepts in multiple regression analysis. These include multicollinearity, interaction effects, and an expansion of the discussion of inference testing, leverage, and variable transformations to multivariate models. Examples from the first article in this series are expanded on using a primarily graphic, rather than mathematical, approach. The importance of the relationships among the predictor variables and the dependence of the multivariate model coefficients on the choice of these variables are stressed. Finally, concepts in regression model building are discussed.

Bias↗

Comparison of logistic regression and linear regression in modeling percentage data.

Percentage is widely used to describe different results in food microbiology, e.g., probability of microbial growth, percent inactivated, and percent of positive samples. Four sets of percentage data, percent-growth-positive, germination extent, probability for one cell to grow, and maximum fraction of positive tubes, were obtained from our own experiments and the literature. These data were modeled using linear and logistic regression. Five methods were used to compare the goodness of fit of the two models: percentage of predictions closer to observations, range of the differences (predicted value minus observed value), deviation of the model, linear regression between the observed and predicted values, and bias and accuracy factors. Logistic regression was a better predictor of at least 78% of the observations in all four data sets. In all cases, the deviation of logistic models was much smaller. The linear correlation between observations and logistic predictions was always stronger. Validation (accomplished using part of one data set) also demonstrated that the logistic model was more accurate in predicting new data points. Bias and accuracy factors were found to be less informative when evaluating models developed for percentage data, since neither of these indices can compare predictions at zero. Model simplification for the logistic model was demonstrated with one data set. The simplified model was as powerful in making predictions as the full linear model, and it also gave clearer insight in determining the key experimental factors.

Clostridium botulinum↗

BMDP program for piecewise linear regression.

Piecewise linear regression has potentially broad applications in medical data analysis as well as other types of regression. Various kinds of algorithms have been proposed for finding optimum piecewise linear regressions. This paper presents a BMDP program for obtaining near optimum piecewise linear regression equations. An idea intrinsic to the method is that restricting parameter space to a discrete set makes the difficult problems become standard problems. Any software having the variable selection feature in the multiple linear regression can be used to apply the method.

Computers↗

A case study found that a regression tree outperformed multiple linear regression in predicting the relationship between impairments and Social and Productive Activities scores.

OBJECTIVE: Many important physiologic and clinical predictors are continuous. Clinical investigators and epidemiologists' interest in these predictors lies, in part, in the risk they pose for adverse outcomes, which may be continuous as well. The relationship between continuous predictors and a continuous outcome may be complex and difficult to interpret. Therefore, methods to detect levels of a predictor variable that predict the outcome and determine the threshold for clinical intervention would provide a beneficial tool for clinical investigators and epidemiologists. STUDY DESIGN AND SETTING: We present a case study using regression tree methodology to predict Social and Productive Activities score at 3 years using five modifiable impairments. The predictive ability of regression tree methodology was compared with multiple linear regression using two independent data sets, one for development and one for validation. RESULTS: The regression tree approach and the multiple linear regression model provided similar fit (model deviances) on the development cohort. In the validation cohort, the deviance of the multiple linear regression model was 31% greater than the regression tree approach. CONCLUSION: Regression tree analysis developed a better model of impairments predicting Social and Productive Activities score that may be more easily applied in research settings than multiple linear regression alone.

Aged↗

Ordinal regression model and the linear regression model were superior to the logistic regression models.

OBJECTIVE: Ordinal scales often generate scores with skewed data distributions. The optimal method of analyzing such data is not entirely clear. The objective was to compare four statistical multivariable strategies for analyzing skewed health-related quality of life (HRQOL) outcome data. HRQOL data were collected at 1 year following catheterization using the Seattle Angina Questionnaire (SAQ), a disease-specific quality of life and symptom rating scale. STUDY DESIGN AND SETTING: In this methodological study, four regression models were constructed. The first model used linear regression. The second and third models used logistic regression with two different cutpoints and the fourth model used ordinal regression. To compare the results of these four models, odds ratios, 95% confidence intervals, and 95% confidence interval widths (i.e., ratios of upper to lower confidence interval endpoints) were assessed. RESULTS: Relative to the two logistic regression analysis, the linear regression model and the ordinal regression model produced more stable parameter estimates with smaller confidence interval widths. CONCLUSION: A combination of analysis results from both of these models (adjusted SAQ scores and odds ratios) provides the most comprehensive interpretation of the data.

Adolescent↗

Determination of antibody affinity by ELISA with a non-linear regression program. Evaluation of linearized approximations.

ELISA experiments based on competition between immobilized and soluble antibody for soluble antigen, and on the formation of a ternary complex of immobilized antibody, antigen and soluble antibody were used by Hoylaerts et al. (J. Immunol. Methods 126 (1990) 253-261) for the determination of dissociation constants. The dissociation constant was taken from linearized plots according to a theory that required several approximations. The effect of these approximations on the resulting dissociation constants has been investigated using two computer programs, CBEIA-C and CBEIA-S. Since most approximations that have to be made for linearization can be avoided by non-linear regression, the programs provide a more reliable basis for calculation of the dissociation constant. Experiments measuring formation of a ternary complex were found to be unsuitable for affinity determination.

Animals↗

Prediction and evolutionary information analysis of protein solvent accessibility using multiple linear regression.

A multiple linear regression method was applied to predict real values of solvent accessibility from the sequence and evolutionary information. This method allowed us to obtain coefficients of regression and correlation between the occurrence of an amino-acid residue at a specific target and its sequence neighbor positions on the one hand, and the solvent accessibility of that residue on the other. Our linear regression model based on sequence information and evolutionary models was found to predict residue accessibility with 18.9% and 16.2% mean absolute error respectively, which is better than or comparable to the best available methods. A correlation matrix for several neighbor positions to examine the role of evolutionary information at these positions has been developed and analyzed. As expected, the effective frequency of hydrophobic residues at target positions shows a strong negative correlation with solvent accessibility, whereas the reverse is true for charged and polar residues. The correlation of solvent accessibility with effective frequencies at neighboring positions falls abruptly with distance from target residues. Longer protein chains have been found to be more accurately predicted than their smaller counterparts.

Amino Acids↗

A simplified statistical method for local INR using linear regression. European Concerted Action on Anticoagulation.

A simplified method of International Normalized Ratio (INR) derivation using linear regression of certified INR plotted against local prothrombin time (PT) results has been compared with INR from conventional orthogonal regression. Linear regression assumes error only with the local PT results whereas orthogonal regression assumes error with both reference and local results. The reliability of local INR derivation using lyophilized plasmas has been assessed in a collaborative study. INR from conventional fresh plasma International Sensitivity Index (ISI) calibrations have been compared with INR from calibrations with two types of lyophilized plasma, artificially depleted and coumarin. Although calibration slopes differed with the two types of analysis and the different lyophilized plasmas, both gave reasonable approximations to fresh plasma ISI calibrations. With orthogonal regression the overall percentage INR deviation was 5.25% with the artificially depleted plasmas and 6.85% for the results with lyophilized coumarins. With the linear regression, deviation was 8.40% for the artificially depleted plasmas and 5.05% for coumarin-treated patients' lyophilised-plasma. The simpler regression method appears to be worthy of further study as the present report has demonstrated that if the calibrant plasmas are accurately certified with the thromboplastin International Reference Plasma (IRP) results approximate to the conventionally determined INR using the manual PT technique. Coagulometers require further assessment.

Blood Coagulation Tests↗

Accurate prediction of heat of formation by combining Hartree-Fock/density functional theory calculation with linear regression correction approach.

A linear regression correction approach has been developed successfully to account for the electron correlation energy missing in Hartree-Fock calculation and to reduce the calculation errors of density functional theory. The numbers of lone-pair electrons, bonding electrons and inner layer electrons in molecules, and the number of unpaired electrons in the composing atoms in their ground states are chosen to be the most important physical descriptors to determine the correlation energy unaccounted by Hartree-Fock method or to improve the results calculated by B3LYP density functional theory method. As a demonstration, this proposed linear regression correction approach has been applied to evaluate the standard heats of formation DeltaH(f) (Theta) of 180 small-sized to medium-sized organic molecules at 298.15 K. Upon correction, the mean absolute deviation for the 150 molecules in the training set decreases from 351.0 to 4.6 kcal/mol and 360.9 to 4.6 kcal/mol for HF/6-31G(d) and HF/6-311+G(d,p) methods, respectively. For B3LYP method, the mean absolute deviations are reduced from 9.2 and 18.2 kcal/mol to 2.7 and 2.4 kcal/mol for 6-31G(d) and 6-311+G(d,p) basis sets, respectively.

Journal Article↗

[Locally adjusted linear regression and its possibilities for application].

A new numerical method is presented for a model free representation of the mean course in a series of measured data of an unknown functional connection. The usefulness of the method is demonstrated by some examples of different time series (chemical reaction kinetics, growth, damped oscillations of a physiological system, pharmacokinetics). The advantages of the new method compared with nonlinear regression or segmented linear regression are noted. For any argument within the interval of measurements, the local adjusted linear regression value may be calculated. In this way one get a continuous curve with a continuous 1st derivative, too. This local adjustment will be obtained by using a weight function which is from Gaussianlike type in our procedure. The calculated continuous approximation curve represents the mean course in the measured values and may serve for a further quantitative evaluation of the measured functional connection, for interpolative and smoothing purposes, or even for model construction and model proving procedures. The nonparametric estimation of a continuous (1 dimensional) distribution density function from realizations of a continuous random variable is another field of application of this method.

Animals↗

A weighted estimating equation for linear regression with missing covariate data.

Linear regression is one of the most popular statistical techniques. In linear regression analysis, missing covariate data occur often. A recent approach to analyse such data is a weighted estimating equation. With weighted estimating equations, the contribution to the estimating equation from a complete observation is weighted by the inverse 'probability of being observed'. In this paper, we propose a weighted estimating equation in which we wrongly assume that the missing covariates are multivariate normal, but still produces consistent estimates as long as the probability of being observed is correctly modelled. In simulations, these weighted estimating equations appear to be highly efficient when compared to the most efficient weighted estimating equation as proposed by Robins et al. and Lipsitz et al. However, these weighted estimating equations, in which we wrongly assume that the missing covariates are multivariate normal, are much less computationally intensive than the weighted estimating equations given by Lipsitz et al. We compare the weighted estimating equations proposed in this paper to the efficient weighted estimating equations via an example and a simulation study. We only consider missing data which are missing at random; non-ignorably missing data are not addressed in this paper.

Adolescent↗

Residual plots for the censored data linear regression model.

To be consistent, censored data linear regression estimators typically require a correctly specified linear regression function and independent and identically distributed errors. For uncensored data one can assess these model assumptions informally by examining plots of the residuals against the independent variables or fitted values. In this paper I propose plots for censored data analogous to these uncensored data residual plots. One can use such plots in the same way as their uncensored data counterparts for checking model assumptions; if the model assumptions are correct, then the plots should exhibit a random scatter. I show that the proposed plots are useful in selecting a linear regression model for the Stanford heart transplant data.

Heart Transplantation↗

Neural network and linear regression models in residency selection.

For many years, multiple linear regression models have been used at a residency program to generate preliminary rank lists of residency applicants. These lists are then used by the admissions committee as an aid in developing a final ranking to submit to the National Residency Match Program (NRMP). A study was undertaken to compare predictions made using linear regression with those generated by a newer technique, an artificial neural network. A prospective cohort design was used. Seventy-four applicants to an emergency medicine program were evaluated by faculty and resident interviewers with regard to medical school grades, autobiography, interviews, letters of recommendation, and National Board scores. Normalization of these scores (by linear transformation of interviewer means) was used to correct for differences among interviewers. Multivariate linear regression and neural network models were developed using data from the previous 5 years' applicants. These models were used to forecast provisional rank orderings of the candidates. These rankings were combined into a single hybrid list that was used by the admissions committee as the starting point for development of the final rank list by consensus. Each model's predictions were tested for goodness of fit against the final NRMP rank using Wilks' test. Using the final submitted NRMP rank order as the dependent variable, the neural network yielded a correlation coefficient of 0.77 and an R2 of 59.4%. The linear regression model exhibited a correlation coefficient of 0.74 and an R2 of 54.0%. No significant difference was found (chi 2 = 1.08, P = .7). A neural network performs as well as a linear regression model when used for forecasting the rank order of residency applicants.

Cohort Studies↗

Application of linear regression to videokeratoscope data for tilted surfaces.

PURPOSE: To investigate the use of linear regression analysis performed on the tabular data display of the EyeSys videokeratoscope (VK). When a radius squared vs distance squared scatterplot is produced from aspheric surface data the equivalent conic section can be deduced from the intercept and slope of the linear regression line. Non-linear plots are often produced. Linear regression may then be applied in a number of ways. METHOD: Topographical data derived from both the EyeSys VK and a computer model of the instrument were analysed by three methods of linear regression. The resultant apical radii, p-values and predicted surface tilts were compared with known values. RESULTS: The three methods predict different surface characteristics whose errors were found to vary depending upon the asphericity of the surface and its tilt. CONCLUSIONS: Apical radius is most accurately predicted by linear regression method (1). Both p-value and tilt are best predicted by averaging the radius and position data for corresponding points in each semi-meridian before squaring the resultant points and performing linear regression (method 3).

Cornea↗

Quantification of dark adaptation dynamics in retinitis pigmentosa using non-linear regression analysis.

PURPOSE: Non-linear regression analysis was used to determine dark adaptation indices in people with retinitis pigmentosa and in control subjects. METHODS: Dark adaptation data were collected for 13 people with retinitis pigmentosa and 21 controls using the Goldmann-Weekers Dark Adaptometer. Data were analysed using an exponential non-linear regression model and dark adaptation indices derived. The results were compared to age-related values. RESULTS: The mean cone threshold of the group with RP (4.73 +/- 0.19 log units) was significantly greater than that found in the control group (3.69 +/- 0.12 log units). The rate of cone dark adaptation in the RP group was not significantly different from that of the control group. The a break in the RP group (6.46 +/- 0.70 minutes) was delayed when compared to the control group (4.29 +/- 0.21 minutes) and the rate of rod dark adaptation in the RP group was slower (10 +/- 2 per cent per minute) than that of the control group (15 +/- 1 per cent per minute). CONCLUSIONS: This study has shown that a relatively simple data analysis can provide a more quantitative and intuitive description of dark adaptation rates in people with retinal disease. This technique will enable more effective use of dark adaptometry as a supplement to objective electrophysiology, when monitoring people with retinitis pigmentosa.

Adult↗

[Quantitative analysis of compound injection of aminopyrine by dual-wavelength linear regression spectrophotometry].

A dual-wavelength linear regression spectrophotometry has been introduced and evaluated. Depending on a group of standard mixture solutions the optimal wavelengths and calibration curve can be determined simultaneously by linear regression method. The deviation of absorption resulting from molecular interaction can be calibrated by this method and the results are more accurate. The validity of this method has been confirmed through its use in the analysis of compound injection of antipyrine with satisfactory recoveries. Results obtained by Kalman filtering (KF), partial least squares (PLS) and target factor analysis (TFA) are also given.

Aminopyrine↗

Validity of linear regression in method comparison studies: is it limited by the statistical model or the quality of the analytical input data?

We compared the application of ordinary linear regression, Deming regression, standardized principal component analysis, and Passing-Bablok regression to real-life method comparison studies to investigate whether the statistical model of regression or the analytical input data have more influence on the validity of the regression estimates. We took measurements of serum potassium as an example for comparisons that cover a narrow data range and measurements of serum estradiol-17beta as an example for comparisons that cover a wide data range. We demonstrate that, in practice, it is not the statistical model but the quality of the analytical input data that is crucial for interpretation of method comparison studies. We show the usefulness of ordinary linear regression, in particular, because it gives a better estimate of the standard deviation of the residuals than the other procedures. The latter is important for distinguishing whether the observed spread across the regression line is caused by the analytical imprecision alone or whether sample-related effects also contribute. We further demonstrate the usefulness of linear correlation analysis as a first screening test for the validity of linear regression data. When ordinary linear regression (in combination with correlation analysis) gives poor estimates, we recommend investigating the analytical reason for the poor performance instead of assuming that other linear regression procedures add substantial value to the interpretation of the study. This investigation should address whether (a) the x and y data are linearly related; (b) the total analytical imprecision (s(a,tot)) is responsible for the poor correlation; (c) sample-related effects are present (standard deviation of the residuals >> s(a,tot)); (d) the samples are adequately distributed over the investigated range; and (e) the number of samples used for the comparison is adequate.

Chromatography, Ion Exchange↗