Further statistics in dentistry. Part 6: Multiple linear regression.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Many variables, including rat body weight (BWt), are sequentially measured in individuals yet the information obtained is often poorly utilised. Many advantages result if a linear relationship is established, on the original or transformed scales, between the variable and time. For individual rats log BWt, which is normally distributed, is linearly related to the reciprocal of age. Slopes (rates of BWt gain) and constants (ultimate BWts) then serve as data. Linearity extended from 4 weeks onwards in ovariectomised (OvX) and intact rats, implying that puberty does not affect BWt and that BWt increases before puberty following pre-pubertal OvX; the ovary inhibiting growth before puberty. For individual rats weaning BWt was correlated with rate of BWt gain but not with ultimate BWt, indicating postweaning compensatory growth which continues to maturity. Pre-treatment and post-treatment measurements were made on individuals, allowing within-animal comparisons and regression. Thus the BWt response to OvX or estrogen treatment depended on the pre-treatment rate of BWt gain and ultimate BWt, implying that the BWt response entails a form of compensatory growth. OvX at three ages (day 2, week 4, and week 7) produced different rates of BWt gain but similar ultimate BWts. These results support the hypothesis that ultimate BWt is predetermined, but can be modified by treatment.
A prospective investigation was undertaken in adults to assess the specificity and sensitivity of fever (greater than 38 degrees C) and leucocytosis (greater than 10 000/microliters) for the diagnosis of infection after operations with cardiopulmonary bypass. A log-linear model analysis of a multiway frequency table was used for statistical evaluation. The model parameters were separately evaluated for 2 periods: the early one until the 6th day, the late period from the 7th postoperative day until discharge. Seven out of 115 patients suffered infections during their hospital stay: Bacteremia occurred in 3, pneumonia in 2, and deep sternal wound infection in 2 patients, and a superficial wound infection in one. No significant interactions between fever, leucocytosis and/or infection were found in the first period, except an inverse relation between fever and elevated WBC (p = 0.0197). After the 6th postoperative day the model parameters did show significant interactions, fever and leucocytosis being more frequent in infected patients. However, the specificity was low: only 15% of the patients with fever or elevated WBC had an infection. The risk of in-hospital infection was significantly higher after a long duration of cardiopulmonary bypass (p = 0.009), and after transfusion of more than 2500 ml of blood on the day of operation (p = 0.001).
Computer fitting of binding data is discussed and it is concluded that the main problem is the choice of starting estimates and internal scaling parameters, not the optimization software. Solving linear overdetermined systems of equations for starting estimates is investigated. A function, Q, is introduced to study model discrimination with binding isotherms and the behaviour of Q as a function of model parameters is calculated for the case of 2 and 3 sites. The power function of the F test is estimated for models with 2 to 5 binding sites and necessary constraints on parameters for correct model discrimination are given. The sampling distribution of F test statistics is compared to an exact F distribution using the Chi-squared and Kolmogorov-Smirnov tests. For low order modes (n less than 3) the F test statistics are approximately F distributed but for higher order models the test statistics are skewed to the left of the F distribution. The parameter covariance matrix obtained by inverting the Hessian matrix of the objective function is shown to be a good approximation to the estimate obtained by Monte Carlo sampling for low order models (n less than 3). It is concluded that analysis of up to 2 or 3 binding sites presents few problems and linear, normal statistical results are valid. To identify correctly 4 sites is much more difficult, requiring very precise data and extreme parameter values. Discrimination of 5 from 4 sites is an upper limit to the usefulness of the F test.
Explore the source record for details and available documents.
Task analyses served as the basis for developing the Mixed Arithmetic-Perceptual (MA-P) model, which proposes (1) that people interacting with common graphs to answer common questions apply a set of component processes--searching for indicators, encoding the value of indicators, performing arithmetic operations on the values, making spatial comparisons among the indicators, and responding; and (2) that the type of graph and user's task determine the combination and order of the components applied (i.e., the processing steps). Two experiments investigated the prediction that response time will be linearly related to the number of processing steps according to the MA-P model. Subjects used line graphs, scatter plots, and stacked bar graphs to answer comparison questions and questions requiring arithmetic calculations. A one-parameter version of the model (with equal weights for all components) and a two-parameter version (with different weights for arithmetic and nonarithmetic processes) accounted for 76%-85% of individual subjects' variance in response time and 61%-68% of the variance taken across all subjects. The discussion addresses possible modifications in the MA-P model, alternative models, and design implications from the MA-P model.
Explore the source record for details and available documents.
The current study examined the impact of a censored independent variable, after adjusting for a second independent variable, when estimating regression coefficients using "naïve" ordinary least squares (OLS), "partial" OLS and full-likelihood models. We used Monte Carlo simulations to determine the bias associated with all three regression methods. We demonstrated that substantial bias was introduced in the estimation of the regression coefficient associated with the variable subject to a ceiling effect when naïve OLS regression was used. Furthermore, minor bias was transmitted to the estimation of the regression coefficient associated with the second independent variable. High correlation between the two independent variables improved estimation of the censored variable's coefficient at the expense of estimation of the other coefficient. The use of "partial" OLS and maximum-likelihood estimation were shown to result in, at most, negligible bias in estimation. Furthermore, we demonstrated that the full-likelihood method was robust under mis-specification of the joint distribution of the independent random variables. Lastly, we provided an empirical example using National Population Health Survey (NPHS) data to demonstrate the practical implications of our main findings and the simple methods available to circumvent the bias identified in the Monte Carlo simulations. Our results suggest that researchers need to be aware of the bias associated with the use of naïve ordinary least-squares estimation when estimating regression models in which at least one independent variable is subject to a ceiling effect.
Free radicals produced in chicken bone tissue by 137Cs gamma-rays were measured using electron paramagnetic resonance spectroscopy. The yield of radicals was found to be proportional to the absorbed dose. Additive re-irradiation of previously irradiated bones is the basis of a method to estimate the absorbed dose in radiation-processed foods. The ability of the method to provide accurate dose assessments for a range of doses (0.5-7.4 kGy) is tested here. A linear fit to the data yields reasonable dose estimates for bone irradiated less than 2 kGy, but fails at the higher doses using a linear approximation to the dose response. These data and their implications are discussed.
A MATHEMATICA package, 'CONDU.M', has been developed to find the polynomial in concentration and temperature which best fits conductimetric data of the type (kappa, c, T) or (kappa, c1, c2, T) of electrolyte solutions (kappa: specific conductivity; ci: concentration of component i; T: temperature). In addition, an interface, 'TKONDU', has been written in the TCL/Tk language to facilitate the use of CONDU.M by an operator not familiarised with MATHEMATICA. All this software is available on line (UPV/EHU, 2001). 'CONDU.M' has been programmed to: (i) select the optimum grade in c1 and/or c2; (ii) compare models with linear or quadratic terms in temperature; (iii) calculate the set of adjustable parameters which best fits data; (iv) simplify the model by elimination of 'a priori' included adjustable parameters which after the regression analysis result in low statistical significance; (v) facilitate the location of outlier data by graphical analysis of the residuals; and (vi) provide quantitative statistical information on the quality of the fit, allowing a critical comparison among different models. Due to the multiple options offered the software allows testing different conductivity models in a short time, even if a large set of conductivity data is being considered simultaneously. Then, the user can choose the best model making use of the graphical and statistical information provided in the output file. Although the program has been initially designed to treat conductimetric data, it can be also applied for processing data with similar structure, e.g. (P, c, T) or (P, c1, c2, T), being P any appropriate transport, physical or thermodynamic property.
OBJECTIVE: The field of Anaesthesia has recently witnessed numerous advances both in the drug administration and monitoring of anaesthetic state. This development has further boosted the efforts and interest of researchers in the automation of clinical Anaesthesia. The success in this direction is possible only when assessment of the depth of hypnotic component of anaesthesia is achieved accurately. This paper describes a technique to arrive at a reliable Depth of Hypnosis (DoH) index using electroencephalographic (EEG) parameters. METHODS: EEG data from nine patients was recorded and processed to obtain a total of 21 EEG parameters. They were reduced to a set of best five parameters after applying graphical variance analysis which evaluates their power to discriminate between awake and unresponsive states. These five parameters were normalized with respect to awake state and used in a first order equation to give DoH index. RESULTS: The value of computed DoH index varied from 0.37 to 0.58 for different patients during anesthetized state (awake value 1). For a single patient, the maximum variation in the index was observed as +/- 5% for different epochs at constant dose. CONCLUSIONS: A combination of irregularity of EEG waveform in time-domain and band powers in frequency domain best describes the difference between awake and anesthetized states. To characterize these states, a set of optimum EEG parameters exists. These parameters must be normalized to reduce interpatient variability. The calculated graded index may be used to assist the anaesthetist in the operating theatre.
UNLABELLED: The toxicity of red bone marrow is widely considered to be a key factor in restricting the activity administered in molecular radiotherapy to suboptimal levels. The assessment of marrow toxicity requires an assessment of the dose absorbed by red bone marrow which, in many cases, requires knowledge of the total red bone marrow mass in a given patient. Previous studies demonstrated, however, that a close surrogate-spongiosa volume (combined tissues of trabecular bone and marrow)-can be used to accurately scale reference patient red marrow dose estimates and that these dose estimates are predictive of marrow toxicity. Consequently, a predictive model of the total skeletal spongiosa volume (TSSV) would be a clinically useful tool for improving patient specificity in skeletal dosimetry. METHODS: In this study, 10 male and 10 female cadavers were subjected to whole-body CT scans. Manual image segmentation was used to estimate the TSSV in all 13 active marrow-containing skeletal sites within the adult skeleton. The age, total body height, and 14 CT-based skeletal measurements were obtained for each cadaver. Multiple regression was used with the dependent variables to develop a model to predict the TSSV. RESULTS: Os coxae height and width were the 2 skeletal measurements that proved to be the most important parameters for prediction of the TSSV. The multiple R(2) value for the statistical model with these 2 parameters was 0.87. The analysis revealed that these 2 parameters predicted the estimated the TSSV to within approximately +/-10% for 15 of the 20 cadavers and to within approximately +/-20% for all 20 cadavers in this study. CONCLUSION: Although the utility of spongiosa volume in estimating patient-specific active marrow mass has been shown, estimation of the TSSV in active marrow-containing skeletal sites via patient-specific image segmentation is not a simple endeavor. However, the alternate approach demonstrated in this study is fairly simple to implement in a clinical setting, as the 2 input measurements (os coxae height and width) can be made with either pelvic CT scanning or skeletal radiography.
This work demonstrates a novel computational approach combining flux balance modeling with statistical methods to identify correlations among fluxes in a metabolic network, providing insight as to how the fluxes should be redirected to achieve maximum product yield. The procedure is demonstrated using the example of amino acid production from an industrial Escherichia coli production strain and a hypothetical engineered strain overexpressing two heterologous genes. Regression analysis based on a random sampling of 5,000 points within the feasible solution space of the E. coli stoichiometric network suggested that increased activity of the glyoxylate cycle or PEP carboxylase and elimination of malic enzyme will improve lysine and arginine synthesis.
A lever-like EEG feature-extraction method based on the Hurst exponent and regression-fitting errors is proposed for identifying beta rhythms. The proposed method is superior to most methods using the time- and frequency-domain feature extraction parameters for identifying beta rhythms.
The aim of this article is to encourage good practice in the statistical analysis of dental research data. Our objective is to highlight the statistical problems of collinearity and multicollinearity. These are among the most common statistical pitfalls in oral health research when exploring the relationship between clinical variables using multiple regression analysis. We hope that this article will show why these problems arise and how they can be avoided and overcome. Examples from the periodontal literature will be used to illustrate how collinearity and multicollinearity can seriously distort the model development process as a result of the phenomenon of mathematical coupling. Knowledge of these problems can help to eliminate misleading results and improve any subsequent interpretations. Regression analyses are useful tools in oral health research when their limitations are recognized. However, care is required in planning and it is worthwhile seeking statistical advice when formulating the study's research questions.
Grains of 26 Turkish wheat cultivars and advanced breeding lines were used in this study. Simple correlations between a number of quality parameters to predict bulgur yield and bulgur cooking quality were determined. Highly significant correlations between bulgur yield and each of the thousand-kernel weight and the sum of the grain over 2.8 + 2.5 mm sieves were obtained for both durum and bread wheat samples (p < 0.01). The regression equations showed that the models involving two variables (the thousand-kernel weight and the thickness of the grain for durum wheat samples; the thousand-kernel weight and the length of the grain for bread wheat samples) resulted in the highest R2 values. For an assessment of the influence of all factors on bulgur cooking properties (total organic matter: TOM and colorimetric test values), simple and multiple regression analyses were used to find equations that predict best the relationship between various quality parameters and bulgur cooking properties. The models involving two variables; the vitreousness and the dry gluten contents for the durum wheat samples and SDS sedimentation test value and wheat protein content for the bread wheat samples resulted in the highest R2 for the TOM value.
The manuscript discusses the application of chemometrics to the handling of TLC response time data. Derivative treatment of chromatographic response data followed by convolution of the resulting derivative curves using 8-points sinx(i) polynomials (discrete Fourier functions) was found to be beneficial in eliminating the interference due to background noise in TLC-densitometric measurements. It also compares the application of Theil's method, a non-parametric regression method, in handling the response data, with the least squares parametric regression method, which is considered the de facto standard method used for regression. Theil's method was found to be superior to the method of least squares as it assumes that errors could occur in both x- and y-directions and they might not be normally distributed. In addition, it could effectively circumvent any outlier data points.
In the present work, a novel method was proposed for prediction of secondary structure. Over a database of 396 proteins (CB396) with a three-state-defining secondary structure, this method with jackknife procedure achieved an accuracy of 68.8% and SOV score of 71.4% using single sequence and an accuracy of 73.7% and SOV score of 77.3% using multiple sequence alignments. Combination of this method with DSC, PHD, PREDATOR, and NNSSP gives Q3 = 76.2% and SOV = 79.8%.