PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Linear Models”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

Interpretation of linear regression models that include transformations or interaction terms.

In linear regression analyses, we must often transform the dependent variable to meet the statistical assumptions of normality, variance stability, or linearity. Transformations, however, can complicate the interpretation of results because they change the scale on which the dependent variable is measured. In this setting, the inclusion of product terms or the transformation of some independent (or predictor) variables may further complicate interpretation. In this article, we present some interpretations of linear models that include transformations or product terms. We illustrate these interpretations using regression analyses designed to study determinants of serum testosterone levels. These examples show how one can present results using simple measures, such as medians, and interpret regression parameters.

Epidemiologic Methods↗

Evaluation of logistic versus linear regression models for predicting pulmonary hypertension syndrome (ascites) using cold exposure or pulmonary artery clamp models in broilers.

Syndromes such as ascites (pulmonary hypertension syndrome) present difficulties both in the interpretation of associated physiological observations and in their analyses. The ability to predict which physiological variables have the greatest influence on survival or, more importantly, which individuals are most susceptible or resistant to ascites would be very useful selection tools. When addressed in this manner, ascites data become binary data sets (healthy or affected). Binary data can be problematic in that they do not meet all of the assumptions necessary for more traditional analyses such as ANOVA and linear regression. Binary data are discrete and do not have normally distributed errors, which violates a fundamental assumption of linear models. The predictive abilities of linear and logistic regression were evaluated in two replicated experiments using two methods to induce ascites, cold exposure (COLD) and surgical clamping of one pulmonary artery (PAC). The logistic and linear predictive models were derived using the same data and variables. The first data set from PAC and COLD were used to develop the predictive models and the replicate data sets of PAC and COLD were used as "test data sets" for the prediction of ascites. The linear models developed were complex, using four or five variables and requiring up to seven different measurements. On average, the linear models predicted ascites correctly 87.6% of the time. The logistic models were simple (single variable) models that predicted ascites correctly 92.0% of the time. The variables used in the logistic models were derivations of the ratio of right ventricular weight to total ventricular weight, either corrected for age or the body weight of the bird. Although linear regression predicted the incidence of ascites almost as well as logistic regression did, logistic regression is the more appropriate test statistic to use.

Analysis of Variance↗

Biomass growth rate during the prokaryote cell cycle.

The rate of biomass growth throughout the cell cycle of prokaryotes is important in the study of global regulation. Two limiting cases have generally been considered: the exponential model and the linear model. The exponential model is a logical expectation because protein, the main component of biomass of a bacterial cell, increases continuously during the cell cycle and therefore the means for synthesis of other cell components and metabolites also increases. In addition, during the cell cycle, ribosomes, the means of production of proteins, increase monotonically. As a consequence, the increase of all should be autocatalytic and the content of cell substance should be an exponential function of time. Two cellular components would not be expected to increase exponentially: the DNA and the cell envelope. The former because of the intermittent synthesis of the chromosome, and the latter because of changes in the surface-to-volume ratio with growth and division. In contrast to the exponential model, the linear model of Kubitschek postulates that the cell only increases its membrane transport capability over a brief period during the cell cycle, and, thus limited by transport, all cell components can increase only at a constant linear rate during most of the cell cycle. Other proposed models are intermediate and assume that the growth rate of the cell depends on some cell cycle event, such as the initiation of chromosome replication. The models have relevance to prokaryotes undergoing balanced growth; they may not be relevant to eukaryotic microbes or to eukaryotic cells in tissue culture that have endogenous rhythms or are controlled by protein growth factors. Logically, the models could possibly apply to a free-living cell that does not respond to environmental cues. Even under rigidly constant conditions, however, cells may try to respond to a stimulus that was periodic or regulatory under natural conditions, but is present at a constant level under the experimental culture condition. There are four classes of experiments that have been used to measure the accumulation of dry biomass or its components during the cell cycle of a bacterium, as typified by Escherichia coli. For the first class of experiments, the dimensions of living cells are measured under the microscope. So far, the experiments have been limited by the resolving power of the phase microscope, but adequate resolution should be possible with the confocal scanning light microscope or various video computer systems. Such experiments are called integral because augmentation of cell constituents is followed. The second class involves pulse-chase labeling of cells and then their separation into different phases of the cycle or age groups and measurement of the radioactivity per cell in the fractions. Such experiments are called differential in that the rate is measured directly instead of being deduced by comparing the total size at different times.(ABSTRACT TRUNCATED AT 400 WORDS)

Cell Cycle↗

[An application of a linear statistical model in clinical anesthesia practice].

One of the important topics in clinical practice of anesthesia, is "prediction". We have to predict what will happen next after obtaining some imperfect data about the patient. The standard and most popularized method is the one utilizing a linear statistical model. This linear model in prediction is based, in part, on the information theory, which says that the volume of information is the linear combination of its components. One of the characteristic models is linear regression analysis. In this paper, basic structure of linear statistical model and its application are discussed. In particular, problems and pitfalls in its application are demonstrated, focusing on variable selection, multi-colinearity and model checking, which will, I believe, lead us to have an access to this clinically important subject "prediction".

Anesthesia↗

Linear invariants under Jukes' and Cantor's one-parameter model.

Linear invariants are random variables with zero expectations under certain assumptions. In this paper, linear invariants under Jukes' and Cantor's one-parameter model, both with and without the assumption that nucleotide frequencies are at equilibrium, are studied using the method developed in a previous paper. Phylogenetic linear invariants (random variables that are linear invariants of some but not all trees of the same number of species) for trees with up to seven species are derived and bases of phylogenetic linear invariant spaces for unrooted trees with four, five and six species are presented. All these bases consist of invariants of simple form. The constraints that specify non-phylogenetic linear invariants (invariants shared by all trees of the same number of species) are determined. Under the assumption that nucleotide frequencies are at equilibrium, it is found that (i) each five-species tree has 17 independent phylogenetic linear invariants, and for two different trees with five species, there are at least three phylogenetic linear invariants of one tree that are not invariants of the other tree; (ii) each six-species tree has 98 independent phylogenetic linear invariants, and for two different trees of six species there are at least nine independent phylogenetic linear invariants that are not invariants of the other tree; and (iii) each seven-species tree has 482 independent phylogenetic linear invariants. It is also found that the number of independent phylogenetic linear invariants is much larger without the assumption of equilibrium than with it, but the reverse is true for the number of non-phylogenetic linear invariants. A class of random variables that are phylogenetic linear invariants with or without the equilibrium assumption is also identified.

Animals↗

Bivariate linear mixed models using SAS proc MIXED.

Bivariate linear mixed models are useful when analyzing longitudinal data of two associated markers. In this paper, we present a bivariate linear mixed model including random effects or first-order auto-regressive process and independent measurement error for both markers. Codes and tricks to fit these models using SAS Proc MIXED are provided. Limitations of this program are discussed and an example in the field of HIV infection is shown. Despite some limitations, SAS Proc MIXED is a useful tool that may be easily extendable to multivariate response in longitudinal studies.

Antiretroviral Therapy, Highly Active↗

Estimation of protein secondary structure and error analysis from circular dichroism spectra.

The estimation of protein secondary structure from circular dichroism spectra is described by a multivariate linear model with noise (Gauss-Markoff model). With this formalism the adequacy of the linear model is investigated, paying special attention to the estimation of the error in the secondary structure estimates. It is shown that the linear model is only adequate for the alpha-helix class. Since the failure of the linear model is most likely due to nonlinear effects, a locally linearized model is introduced. This model is combined with the selection of the estimate whose fractions of secondary structure summate to approximately one. Comparing the estimation from the CD spectra with the X-ray data (by using the data set of W.C. Johnson Jr., 1988, Annu. Rev. Biophys. Chem. 17, 145-166) the root mean square residuals are 0.09 (alpha-helix), 0.12 (anti-parallel beta-sheet), 0.08 (parallel beta-sheet), 0.07 (beta-turn), and 0.09 (other). These residuals are somewhat larger than the errors estimated from the locally linearized model. In addition to alpha-helix, in this model the beta-turn and "other" class are estimated adequately. But the estimation of the antiparallel and parallel beta-sheet class remains unsatisfactory. We compared the linear model and the locally linearized model with two other methods (S. W. Provencher and J. Glöckner, 1981, Biochemistry 20, 1085-1094; P. Manavalan and W. C. Johnson Jr., 1988, Anal. Biochem. 167, 76-85). The locally linearized model and the Provencher and Glöckner method provided the smallest residuals. However, an advantage of the locally linearized model is the estimation of the error in the secondary structure estimates.

Circular Dichroism↗

A comparison of some statistical techniques for road accident analysis.

At the TRRL/SWOV Workshop on Accident Analysis Methodology, held in Amsterdam in 1988, the need to establish a methodology for the analysis of road accidents was firmly stated by all participants. Data from different countries cannot be compared because there is no agreement on research methodology, data collection, and analysis. Linear and log-linear models are regularly used for the analysis of such data. This paper discusses the background of these models and the model assumptions. It is stated that two relevant and fundamentally different classes of models exist. However, in most cases the difference between these classes remains implicit. These classes of models are called design oriented and object oriented models. Within the context of linear models, the basic assumptions deal with this distinction and the generalization of linear to log-linear models. It is argued that there are advantages and disadvantages for each choice and that a check on the tenability of a particular model is to be recommended. Two different types of generalized linear models, both object-oriented, are used by SWOV and TRRL for the analysis of accidents of road users or on particular road locations. The qualitative data analysis (QDA) models used at SWOV are primarily descriptive and useful for model exploration. The generalized linear interactive modeling (GLIM) technique used at TRRL has stricter assumptions, but is better suited for model testing. In order to compare these models, a number of analyses has been carried out on roundabout data from TRRL. The outcomes are consistent and highly comparable. The results show that for this kind of data these two types of models are to be preferred over classical multiple linear regression analysis (MLR) and log-linear analysis of contingency tables (LLA). QDA turned out to be the most powerful in data exploration. A QDA analysis showed that the dominant accident type of entering and circulating accidents had a different relation with the road and traffic characteristics of the roundabouts than the other accident types and should be analyzed separately. Furthermore, it was shown by a QDA analysis that the choice of a log-link function was preferable over an identity relation as used in MLR. This log-link function in combination with the choice of a Poisson error distribution results in a GLIM model with parameter estimates and errorbounds for these estimates. The errorbounds can be used as an indication of the reliability of the parameters. It is stated that the combination of QDA and GLIM is most powerful for the analysis of this type of problems.

Accidents, Traffic↗

A componential model of human interaction with graphs: 1. Linear regression modeling.

Task analyses served as the basis for developing the Mixed Arithmetic-Perceptual (MA-P) model, which proposes (1) that people interacting with common graphs to answer common questions apply a set of component processes--searching for indicators, encoding the value of indicators, performing arithmetic operations on the values, making spatial comparisons among the indicators, and responding; and (2) that the type of graph and user's task determine the combination and order of the components applied (i.e., the processing steps). Two experiments investigated the prediction that response time will be linearly related to the number of processing steps according to the MA-P model. Subjects used line graphs, scatter plots, and stacked bar graphs to answer comparison questions and questions requiring arithmetic calculations. A one-parameter version of the model (with equal weights for all components) and a two-parameter version (with different weights for arithmetic and nonarithmetic processes) accounted for 76%-85% of individual subjects' variance in response time and 61%-68% of the variance taken across all subjects. The discussion addresses possible modifications in the MA-P model, alternative models, and design implications from the MA-P model.

Computer Graphics↗

The mechanism of involuntary visual spatial attention revealed by a new linear computation model.

EEG reveals brain electrical activities with high temporal resolution. Yet, multiple implicit variables may be involved in limited event related potential (ERP) measures. Special computation techniques are needed to recover these parameters. In the study of involuntary visual spatial attention, we may obtain the ERP in valid cued (V), invalid cued (I) and neutral cued (N) conditions. Usually, the effect of involuntary attention is computed by the subtraction model with the assumption that V/I and N are independent. Yet, they should be related. Treating V/I as a function of N, a linear model V(I) = W + GN is assumed, where W and G are implicit in the ERP measures. G is the gain control on the neutral function. Provided G and W are constant over a local brain region, we may use the Total Least Square (TLS) algorithm to compute their values. The values of W and G computed from an involuntary attention experiment data show that multiple implicit variables are involved in obtained ERPs. Here G acts as a "top-down" sensory modulator on the neutral ERPs and W is related to possible newly involved neural activities. The parameters derived from the new linear model also suggest that there are different mechanisms involved in involuntary attention and voluntary allocation of attention.

Adult↗

Linear discriminant models for unbalanced longitudinal data.

This paper discusses statistical methods for the classification of observations into one of two or more groups based on longitudinal observations. Measurements on subjects in longitudinal medical studies are often collected at different times and on a different number of occasions. Classical multivariate methods for linear discriminant analysis are difficult to apply to repeated measurements due to the highly unbalanced structure observed in these data. Linear models for the analysis of longitudinal data proposed by Laird and Ware and non-linear models proposed by Lindstrom and Bates can be used to estimate population parameters for a discriminant model that classifies individuals into distinct predefined groups or populations. An example is presented using data from a study in 150 pregnant women in Santiago, Chile, in order to predict normal versus abnormal pregnancy outcomes.

Abortion, Spontaneous↗

Analysis of litter size and days to lambing in the Ripollesa ewe. I. Comparison of models with linear and threshold approaches.

The analysis focused on model fitting of 2 ewe reproductive traits, litter size, and days to lambing (interval between the introduction of the ram into the flock and the subsequent parturition of the ewes). The experimental data set of the Universitat Autònoma of Barcelona flock was used, including 1,598 records of litter size and 1,699 records of days to lambing from 376 Ripollesa ewes between 1986 and 2005. Univariate and bivariate models were considered as beginning points with linear or threshold approximation for litter size. Model fitting was evaluated in terms of goodness-of-fit and predictive ability, using the mean square error and the correlation between phenotypic and predicted records (rho(y,ŷ)) as reference parameters. The bivariate model was preferable for both variables, minimizing mean square error and maximizing rho(y,ŷ). A threshold approximation for litter size was preferable over a linear approximation. Models were also compared with a simulation study, comparing the correlation coefficient between simulated and predicted breeding values (rho(a,â)). The bivariate threshold model was favored, with a rho(y,ŷ) of 0.677 and 0.834 for litter size and days to lambing, respectively. Correlation coefficients between simulated and predicted breeding values in the bivariate linear model were reduced slightly to 0.651 and 0.831, respectively, and they were lowest with linear univariate models (0.642 and 0.802). Although the bivariate models for ewe litter size and days to lambing were more accurate than the univariate models, the threshold approaches showed a greater advantage under the bivariate model. For the purpose of genetic evaluation of litter size in sheep, use of the threshold-linear model seems justified. In the Ripollesa breed, the evaluation of litter size can benefit from recording birth weight.

Animals↗

Comparison of logistic regression and linear regression in modeling percentage data.

Percentage is widely used to describe different results in food microbiology, e.g., probability of microbial growth, percent inactivated, and percent of positive samples. Four sets of percentage data, percent-growth-positive, germination extent, probability for one cell to grow, and maximum fraction of positive tubes, were obtained from our own experiments and the literature. These data were modeled using linear and logistic regression. Five methods were used to compare the goodness of fit of the two models: percentage of predictions closer to observations, range of the differences (predicted value minus observed value), deviation of the model, linear regression between the observed and predicted values, and bias and accuracy factors. Logistic regression was a better predictor of at least 78% of the observations in all four data sets. In all cases, the deviation of logistic models was much smaller. The linear correlation between observations and logistic predictions was always stronger. Validation (accomplished using part of one data set) also demonstrated that the logistic model was more accurate in predicting new data points. Bias and accuracy factors were found to be less informative when evaluating models developed for percentage data, since neither of these indices can compare predictions at zero. Model simplification for the logistic model was demonstrated with one data set. The simplified model was as powerful in making predictions as the full linear model, and it also gave clearer insight in determining the key experimental factors.

Clostridium botulinum↗

A Macintosh BASIC program for fitting linear additive models to data by weighted least squares methods, with automatic elimination of redundant parameters from the model.

A BASIC program for fitting linear additive models to data with weighted or unweighted least squares methods and matrix procedures is described. Written and compiled to run on Macintosh II type machines (68020/68030/68040) with coprocessor (68881/68882), the program fits the model supplied by the user as an X or design matrix and provides the option of making the model's parameters orthogonal. Predicted values plus regression and residual terms are calculated. For the weighted fit, the significance of the residual term is evaluated and redundant parameters sequentially and automatically removed until a minimum adequate model is achieved. An appropriate design matrix allows weighted fitting of models for multiway analyses of variance (e.g. factorial, nested and split plot), simple or multiple linear or curvilinear regression, combinations of these in covariance analyses, and genetic analyses (including diallel cross analyses). Sums of squares for individual parameters and orthogonal comparisons can be calculated by the program.

Least-Squares Analysis↗

Linear multivariate models for physiological signal analysis: theory.

The general linear parametric multivariate modelling concept is presented. This model combines a variety of different kinds of multivariate linear models. The concept of partial spectral analysis is derived from the general model. Some emphasis is laid on the causality demands of the model, and it is shown that the classic strictly-causal structure must be abandoned in order to utilise the modelling in many practical situations. Two special sub-class models are described in detail: the multivariate autoregressive model and the multivariate dynamic adjustment model. Furthermore, time-varying modelling is considered. The modelling of the real system is presented on a general level as a system identification cycle. The application of the methods to real physiological data is presented in the companion paper.

Data Collection↗

Pressure-volume analysis of the lung with an exponential and linear-exponential model in asthma and COPD. Dutch CNSLD Study Group.

The prevalence of abnormalities in lung elasticity in patients with asthma or chronic obstructive pulmonary disease (COPD) is still unclear. This might be due to uncertainties concerning the method of analysis of quasistatic deflation lung pressure-volume curves. Pressure-volume curves were obtained in 99 patients with moderately severe asthma or COPD. These patients were a subgroup from a Dutch multicentre trial; the entire group was selected on the basis of a moderately lowered % predicted forced expiratory volume in one second (FEV1), and a provocative concentration of histamine producing a 20% decrease in FEV1 (PC20) < 8 mg.mL-1 obtained with the 2 min tidal breathing technique. The curves were fitted with an exponential (E) model and an exponential model which took the linear appearance in the mid vital capacity range into account (linear-exponential (LE)). The linear-exponential model showed a markedly better fit ability, yielding additional parameters, such as the compliance at functional residual capacity (FRC) level as slope of the linear part (b), and the volume at which the linear part changed into the exponential part of the curve (transition volume (Vtr)). Vtr (mean value Vtr/total lung capacity (TLC) = 0.79 (SD 0.07)) showed a close positive linear correlation with obstruction and hyperinflation variables, which might be due to airway closure, already starting at elevated lung volumes. The exponential shape factor K was closely correlated with b and mean values (K = 1.32 (SD 0.05) kPa-1; b = 2.96 (SD 1.16) L,kPa-1) and the relationship with age was comparable with data reported in healthy individuals. The shape factor of the linear-exponential fit showed no correlation with any elasticity related variable. Neither the elastic recoil at 90% TLC, as obtained from the linear-exponential fit, nor its relationship with age were significantly different from healthy individuals. We conclude that, for a more accurate description of the lung pressure-volume curve, a linear-exponential fit is preferable to an exponential model. However, the physiological relevance of the shape parameter (KLE) is still unclear. These results indicate that patients with moderately severe asthma or COPD had, on average, no appreciable loss of elastic lung recoil as compared with healthy individuals.

Adolescent↗

A linear regression model for the analysis of life times.

A linear model is suggested for the influence of covariates on the intensity function. This approach is less vulnerable than the Cox model to problems of inconsistency when covariates are deleted or the precision of covariate measurements is changed. A method of non-parametric estimation of regression functions is presented. This results in plots that may give information on the change over time in the influence of covariates. A test method and two goodness of fit plots are also given. The approach is illustrated by simulation as well as by data from a clinical trial of treatment of carcinoma of the oropharynx.

Clinical Trials as Topic↗

Properties of random regression models using linear splines.

Properties of random regression models using linear splines (RRMS) were evaluated with respect to scale of parameters, numerical properties, changes in variances and strategies to select the number and positions of knots. Parameters in RRMS are similar to those in multiple trait models with traits corresponding to points at knots. RRMS have good numerical properties because of generally superior numerical properties of splines compared with polynomials and sparser system of equations. These models also contain artefacts in terms of depression of variances and predictions in the middle of intervals between the knots, and inflation of predictions close to knots; the artefacts become smaller as correlations corresponding to adjacent knots increase. The artefacts can be greatly reduced by a simple modification to covariables. With the modification, the accuracy of RRMS increases only marginally if the correlations between the adjacent knots are > or =0.6. In practical analyses the knots for each effect in RRMS can be selected so that: (i) they cover the entire trajectory; (ii) changes in variances in intervals between the knots are approximately linear; and (iii) the correlations between the adjacent knots are at least 0.6. RRMS allow for simple and numerically stable implementations of genetic evaluations with artefacts present but transparent and easily controlled.

Animals↗