PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Models, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Charge sequence coding in statistical modeling of unfolded proteins.

Unfolded proteins recently attracted attention due to accumulation of experimental evidences for their significant role in different life processes. Modeling of electrostatic interactions (EI) in unfolded state of proteins is becoming increasingly important as well. In this paper, we stress on the importance of how the sequence of charged residues of a given protein is incorporated into the models for calculation of EI in the unfolded state. On the basis of the distributions of distances between titratable sites of charged residues calculated for polypeptide chains of various compositions, it was found that the distance distribution for a pair of residues, located close to each other along the sequence of a protein, depends on what residues constitute the pair in question. It was concluded that the consideration of these residue-specific distributions is essential for a statistical model to be accurate from the physical point of view. It was suggested that use of distance intervals in the spherical model of unfolded proteins accounts better for the charge sequence than the set of single distance values. This was illustrated by comparison of the pK values of the titratable groups of the unfolded N-terminal SH3 domain of the Drosophila protein drk to the available experimental data.

Amino Acids↗

A statistical model for the "N-of-1" study.

The controlled clinical trial has largely replaced case-reports as the authoritative source of information concerning the efficacy of treatment. However, many situations arise in clinical practice where treatment decisions cannot be made on the basis of such studies. The definitive clinical trial may not have been performed, or the results from a particular study may not be applicable to a particular patient. Recently, "N-of-1" studies have been proposed for the experimental evaluation of therapy in a single patient. Multiple courses of active and placebo treatments are administered, and efficacy is determined by following the response measure over a period of time. The purpose of this paper is to present a statistical model appropriate for data arising from this design. The model provides for serial correlation among the response measures captured from the subject, and for heteroskedasticity across the treatment periods. ML estimation procedures are considered, and their properties are investigated. A scoring algorithm is described to iterate to the solution of the ML equations, and considerations for hypothesis testing are presented. The techniques are illustrated through an example.

Aged↗

A statistical model of the VA/Q distribution.

A "blocks model" is proposed to model the distribution of the ventilation-perfusion ratio (VA/Q distribution). This model is developed from statistical principles and enables the estimated VA/Q distribution to be interpreted in a straightforward and intuitive manner. Estimation of parameters of the blocks model uses a constrained weighted least-squares procedure. Although developed initially to estimate VA/Q distributions from data generated by the multiple inert gas elimination technique (P. D. Wagner, H. A. Saltzman, and J. B. West, J. Appl. Physiol. 36: 588-599, 1974), the blocks method is applicable to any problem in which the unknown distribution is related to the data through an ill-posed integral equation and is particularly suited for problems in which the data are scarce. The method is illustrated with several examples--hypothetical data representing a wide range of VA/Q distributions as well as some real data.

Computer Simulation↗

Statistical limitations in functional neuroimaging. I. Non-inferential methods and statistical models.

Functional neuroimaging (FNI) provides experimental access to the intact living brain making it possible to study higher cognitive functions in humans. In this review and in a companion paper in this issue, we discuss some common methods used to analyse FNI data. The emphasis in both papers is on assumptions and limitations of the methods reviewed. There are several methods available to analyse FNI data indicating that none is optimal for all purposes. In order to make optimal use of the methods available it is important to know the limits of applicability. For the interpretation of FNI results it is also important to take into account the assumptions, approximations and inherent limitations of the methods used. This paper gives a brief overview over some non-inferential descriptive methods and common statistical models used in FNI. Issues relating to the complex problem of model selection are discussed. In general, proper model selection is a necessary prerequisite for the validity of the subsequent statistical inference. The non-inferential section describes methods that, combined with inspection of parameter estimates and other simple measures, can aid in the process of model selection and verification of assumptions. The section on statistical models covers approaches to global normalization and some aspects of univariate, multivariate, and Bayesian models. Finally, approaches to functional connectivity and effective connectivity are discussed. In the companion paper we review issues related to signal detection and statistical inference.

Bayes Theorem↗

Improving the design and analysis of high-throughput screening technology comparison experiments using statistical modeling.

Contemporary small-molecule drug discovery frequently involves the screening of large compound files as a core activity. Subsequently cost, speed, and safety become critical issues. In order to meet this need, numerous technologies have been developed to allow mix and measure approaches, facilitate miniaturization, and to increase speed and to minimize the use of potentially hazardous reagents such as radioactive materials. However, despite the on-paper advantages of these new technologies, risks can remain undefined. For example, the question of whether the novel method will facilitate identification of active chemical series in a way that is comparable with conventional methods arises. In order to address this question, we have taken the approach of carrying out experiments to directly compare the output of high-throughput screens using a given novel approach and a traditional method. The concordance between the screening methods can then be determined via comparison of the numbers and structures of the active molecules identified. This article describes the approach taken in our laboratory to minimize variability in such experiments and shows data that exemplifies the general result of lower than expected concordance. Statistical modeling was subsequently used to facilitate this interpretation. The model used beta-distribution function to generate a real-activity frequency relationship with added normal random error and occasional outliers to represent assay variability. Hence, the effect of assay parameters such as the threshold, the number of real actives, and the number of outliers and the standard deviation could readily be explored. The model was found to describe the data reasonably and moreover was found to be of great utility when it came to planning further optimal experiments. A key conclusion from the model was that concordance between screening methods could appear poor even when one approach is compared with itself. This occurs simply because the result is a function of assay threshold, standard deviation and the true compound % activity. In response to this finding we have adopted alternative experimental designs that more reliably measure the concordance between screening methods.

Biological Assay↗

Statistical modeling and projections of lung cancer mortality in 4 industrialized countries.

The purpose of this work was to model lung cancer mortality as a function of past exposure to tobacco and to forecast age-sex-specific lung cancer mortality rates. A 3-factor age-period-cohort (APC) model, in which the period variable is replaced by the product of average tar content and adult tobacco consumption per capita, was estimated for the US, UK, Canada and Australia by the maximum likelihood method. Age- and sex-specific tobacco consumption was estimated from historical data on smoking prevalence and total tobacco consumption. Lung cancer mortality was derived from vital registration records. Future tobacco consumption, tar content and the cohort parameter were projected by autoregressive moving average (ARIMA) estimation. The optimal exposure variable was found to be the product of average tar content and adult cigarette consumption per capita, lagged for 25-30 years for both males and females in all 4 countries. The coefficient of the product of average tar content and tobacco consumption per capita differs by age and sex. In all models, there was a statistically significant difference in the coefficient of the period variable by sex. In all countries, male age-standardized lung cancer mortality rates peaked in the 1980s and declined thereafter. Female mortality rates are projected to peak in the first decade of this century. The multiplicative models of age, tobacco exposure and cohort fit the observed data between 1950 and 1999 reasonably well, and time-series models yield plausible past trends of relevant variables. Despite a significant reduction in tobacco consumption and average tar content of cigarettes sold over the past few decades, the effect on lung cancer mortality is affected by the time lag between exposure and established disease. As a result, the burden of lung cancer among females is only just reaching, or soon will reach, its peak but has been declining for 1 to 2 decades in men. Future sex differences in lung cancer mortality are likely to be greater in North America than Australia and the UK due to differences in exposure patterns between the sexes.

Adult↗

A comparison of statistical models in predicting violence in psychotic illness.

BACKGROUND: The application of statistical modeling techniques, including classification and regression trees, in the prediction of violence has increasingly received attention. METHODS: The predictive performance of logistic regression and classification tree methods in predicting violence was explored in a sample of patients with psychotic illness. RESULTS: Of 2 logistic regression models, the forward stepwise method produced a simpler model than the full model, but the latter performed better. The performance of the classification tree appeared to be high before cross-validation, but reduced when cross-validated. The standard logistic model was the most robust model. A simplified tree with extra weight given to violent cases was a reasonable competitor and was simple to apply. CONCLUSION: Although classification trees can be suitable for routine clinical practice, because of the simplicity of their decision-making processes, their robustness and therefore clinical utility was problematic in this sample. Further research is required to compare such models in large prospective epidemiologic studies of other psychiatric populations.

Adolescent↗

Identification of patients with evolving coronary syndromes by using statistical models with data from the time of presentation.

OBJECTIVE: To derive statistical models for the diagnosis of acute coronary syndromes by using clinical and ECG information at presentation and to assess performance, portability, and calibration of these models, as well as how they may be used with cardiac marker proteins. DESIGN AND METHODS: Data from 3462 patients in four UK teaching hospitals were used. Inputs for 8, 14, 25, and 43 factor logistic regression models were selected by using log10 likelihood ratios (log10 LRs). Performance was analysed by receiver operating characteristic curves. RESULTS: A 25 factor model derived from 1253 patients from one centre was selected for further study. On training data, 98.2% of ST elevation myocardial infarctions (STEMIs) and 96.2% of non-ST elevation myocardial infarctions (non-STEMIs) were correctly classified, whereas only 2.1% of non-cardiac cases were incorrectly classified. On data from three other centres, 97.3% of STEMIs and 91.9% of non-STEMIs were correctly classified. Differences in log10 LRs for individual inputs from different centres accounted for the decline in performance when models were applied to unseen data. Classification was improved when output was combined with either clinical opinion or marker proteins. CONCLUSIONS: Logistic regression models based on data available at presentation can classify patients with chest pain with a high degree of accuracy, particularly when combined with clinical opinion or marker proteins.

Adolescent↗

A comparative study of two statistical models for the analysis of binary data from longitudinal studies.

This study extensively compares two statistical models for the analysis of binary data from longitudinal studies. The first model was proposed by Zeger, Liang, and Self, which was abbreviated as ZLS model and another model was proposed by Origasa. The comparison focuses on both analytical and statistical view-points. The first discusses a type of the models and the second evaluates the effect from model misspecification by stimulation, assuming that the ZLS model is true.

Computer Simulation↗

Visual and statistical modeling of facial movement in patients with cleft lip and palate.

OBJECTIVE: To analyze and display facial movement data from noncleft subjects and from patients with cleft lip and palate by using a new dynamic approach. The hypothesis was that there are differences in facial movement between the patients with cleft lip and palate and the noncleft subjects. SETTING: Subjects were recruited from the University of North Carolina School of Dentistry Orthodontic and Craniofacial Clinics. PATIENTS, PARTICIPANTS: Sixteen patients with cleft lip and palate and eight noncleft "control" subjects. INTERVENTIONS: Video recordings and measurements in three dimensions of facial movement. MAIN OUTCOME MEASURES: Principal component (PC) scores for each of six animations or movements and dynamic modeling of mean animations. STATISTICS: Multivariate statistics were used to test for significant differences in the PC mean scores between the patient groups and the noncleft groups. RESULTS: No statistically significant differences were found in PC mean scores between the patient groups and the noncleft groups; however, the variability of the effect of clefting on the soft tissues during animation was noted when the noncleft data were used to establish a "normal" scale of movement. Compensatory movements were seen in some of the patients with cleft lip and palate, and the compensation was not unidirectional. CONCLUSION: Measures of mean movement differences as summarized by PC scores between patients with cleft lip and palate and noncleft subjects may be misleading because of extreme variations about the mean in the patient group that may neutralize group differences. It may be more appropriate to compare patients to a noncleft normal scale of movement.

Adolescent↗

A statistical model for the classification of imipramine response in depressed inpatients.

We present a statistical model, recently developed for application in mathematical economics, that yields empirical evidence for the existence of two distinct subtypes of depression. We reanalyze previously reported data on 65 depressed patients treated with imipramine and repeatedly rated on the Hamilton Rating Scale (HRS). The estimated model parameters suggest two underlying response processes. Patients in the first subgroup were initially more severely depressed (as measured by the total HRS score), but exhibited a rapid rate of symptomatic response over time. In contrast, patients in the second subgroup were initially less severely depressed, yet showed a much slower rate of improvement. These findings recommend a refinement of the clinical definitions of endogenous depression. While this model suggests that there are two underlying response processes, it does not classify individual patients. Subject classification was made possible by fitting an item-response model to the data. This model relates the 17 individual symptom ratings, at baseline, to the total post-treatment HRS scores. The results of this analysis suggest that the previously described relationship between high initial severity of depression and more rapid improvement over time is characteristic of subjects who exhibit initial motor retardation and decreased sexual interest. Graphical analysis clearly indicates that subjects who exhibit both motor retardation and decreased sexual interest at baseline have higher total HRS scores at baseline and show more pronounced improvement in total HRS scores in response to treatment with imipramine. Patients not exhibiting motor retardation and decreased sexual interest were less severely depressed initially and show virtually no clinical response over time.

Depressive Disorder↗

Have sperm counts been reduced 50 percent in 50 years? A statistical model revisited.

OBJECTIVE: To reanalyze data that were used in a linear model to predict that mean sperm counts have been reduced globally by approximately 50% in the last 50 years. DESIGN: The mean sperm counts and their temporal distribution were reanalyzed via several different statistical models (quadratic, spline fit, and stairstep). CONCLUSION: There are several reasons why a published linear regression model is inappropriate to infer a 50% reduction in mean sperm counts in the last 50 years. These include [1] the potential selection biases that may have occurred with the 61 assembled studies such that they are not representative of their underlying populations; [2] the likely variability in collection methods, in particular, the lack of adherence to a minimum prescribed abstinence period, as has been stated for the largest study, which contained 29.7% of all the subjects included in the analysis; [3] the paucity of data in the first 30 years of the 50-year trend analysis; [4] the fact that if the last 20 years of data are examined, which contains 78.7% of all the studies and 88.1% of the total number of subjects, there is no decrease in sperm counts, in fact, sperm counts were observed to have increased; [5] the conflicting data from a large individual laboratory, which was not prone to the collection variability that likely occurred between the 61 studies, that did not suggest a decline in mean sperm count or seminal volume during a comparable time period, even though this laboratory published the data that were largely responsible for the high historical values in the linear model; and, most importantly, [6] the variety of other mathematical models that perform statistically better at describing the recent data than the linear model and thus offer substantially different hypotheses. The data are only robust during the last 20 years of the analysis, in which all the models, except the linear model, suggest constant or slightly increasing sperm counts.

Female↗

A test of several parametic statistical models for estimating success rate in the treatment of carcinoma cervix uteri.

The parametric statistical models discussed include all those which have previously been described in the literature (Boag, 1948-lognormal; Berkson and Gage, 1952-negative exponential; Haybittle, 1959-extrapolated actuarial) and the basic data used to test the models comprised some 3000 case histories of patients treated between 1945 and 1962. The histories were followed up during the period treated between 1945 and 1962. The histories were followed up during the period 1969-71 and thus provided adequate information to validate long-term survival fractions predicted using short-term follow-up data. The results with the log-normal model showed that for series of staged carcinoma cervix patients treated during a 5-year period, satisfactory estimates of long-term survival fractions could be predicted after a minimum waiting period of 3 years for stages I and II, and 2 years for stage III. The model should be used with a value assumed for the lognormal paramater S in the range S = 0.35 to S = 0.40. Although alternative models often gave adequate predictions, the lognormal proved to be the most consistent model. This model may therefore now be used with more confidence for prospective studies on carcinoma cervix series and can provide good estimates of long-term survival fractions several years earlier than would otherwise be possible.

Female↗

A statistical model to compare road mortality in OECD countries.

The objective of this paper is to compare safety levels and trends in OECD countries from 1980 to 1994 with the help of a statistical model and to launch international discussion and further research about international comparisons. Between 1980 and 1994, the annual number of fatalities decreased drastically in all the selected countries except Japan (+ 12%), Greece (+ 56%) and ex-East Germany (+ 50%). The highest decreases were observed in ex-West Germany (- 48%), Switzerland (- 44%), Australia (- 40%), and UK (- 39%). In France, the decrease in fatalities over the same period reached 34%. The fatality rate, an indicator of risk, decreased in the selected countries from 1980 to 1994 except in the east-European countries during the motorization boom in the late 1980s. As fatality rates are not sufficient for international comparisons, a statistical multiple regression model is set up to compare road safety levels in 21 OECD countries over 15 years. Data were collected from IRTAD (International Road Traffic and Accident Database) and other OECD statistical sources. The number of fatalities is explained by seven exogenous (to road safety) variables. The model, pooling cross-sectional and time series data, supplies estimates of elasticity to the fatalities for each variable: 0.96 for the population; 0.28 for the vehicle fleet per capita; -0.16 for the percentage of buses and coaches in the motorised vehicle fleet; 0.83 for the percentage of youngsters in the population; - 0.41 for the percentage of urban population; 0.39 for alcohol consumption per capita; and 0.39 for the percentage of employed people. The model also supplies a rough estimate of the safety performance of a country: the regression residuals are supposed to contain the effects of essentially endogenous and unobserved variables, independent to the exogenous variables. These endogenous variables are safety performance variables (safety actions, traffic safety policy, network improvements and social acceptance). A new indicator, better than the mortality rate, is then set upon the residuals. Mean estimates of this indicator for the years 1980-1982 and the years 1992-1994 rank the countries in the beginning and at the end of the study period. Countries showing the best ranks (and thus the best performance) in 1980 and 1994 are Sweden, the Netherlands and Norway. The UK and Switzerland reach the top 5 in 1994. Greece, Belgium, Portugal and Spain are the last countries in the classification along with, surprisingly, the USA. France was ranked 18th in 1980 and 15th in 1994 but is ranked amongst the five countries that most improved from 1980 to 1994. This model remains non definitive because it is not able to distinguish between safety performance and unobserved exogenous variables although these exogenous variables could explain more about the differences in levels and trends between the countries. More complex models, particularly highly sophisticated models regarding the number of fatalities with breakdowns by road users or road classes would be needed to give a precise and profound ranking of safety levels and safety improvements between countries.

Accidents, Traffic↗

A statistical model estimating the number of African-American physicians in the United States.

Using mark-recaptured methodology and network sampling procedures, a statistical model was developed to estimate the number of African-American physicians in the United States. A sample (stratified by geographic region, medical specialty and an age surrogate) was selected from the National Medical Association's Masterfile of Black Physicians (NMAMBP). Respondents were asked to list the names of five black physicians who resided or practiced in their immediate geographic area. Data also were collected about citizenry as well as other demographic and professional information. The NMAMBP was used mathematically as a "marked" group that could then be "recaptured," allowing mark-recapture methodology to be used as the nucleus of the statistical estimation procedure. The results revealed that in 1991, the total number of US African-American physicians (black US citizens) was estimated to be 16,282 with a conservative standard error of 764 and an approximate 95% confidence interval, yielding a range of 14,754 to 17,810 physicians. This estimate is from 17% to about 32% lower than the 21,538 black doctors reported by the 1990 Bureau of the Census and has important implications for attempts to reform the health-care system and policies designed to produce more African-American physicians.

Black or African American↗

A statistical model providing comprehensive predictions for the mRNA differential display.

MOTIVATION: Differential display (DD) or arbitrarily primed fingerprinting serves to identify differentially expressed genes, but these techniques cannot determine how many of the theoretically available genes have been uncovered. Previous mathematical models are unsatisfying as they are not suitable to analyze experimental data. RESULTS: In the present study, we provide a statistical model based on the redundancy of cDNA fragments amplified during DD experiments. This model is applicable to any DD and predicts (1) the total number of genes expressed in a sample cell type or tissue, (2) the number of differentially expressed genes, (3) the coverage obtained with any given number of primer combinations. In a DD experiment comparing two developmental stages of the post natal rat inner ear, we estimated the total number of differentially expressed genes accessible by DD to be 445, and the number of primer combinations required to uncover 90% of these to be 127. AVAILABILITY: The algorithms were implemented in Matlab (The Mathworks, Inc., Natick, MA) environment and are available at www.physiologie.uni-freiburg.de/download.html CONTACT: ellen.reisinger@physiologie.uni-freiburg.de.

Computer Simulation↗

Statistical models of acute mountain sickness.

Acute mountain sickness (AMS) is caused by exposure to altitudes exceeding 2500 m and often resolves by acclimatization without further ascent. Statistical models of AMS score and the probability of an AMS diagnosis were developed to allow the combination of dissimilar exposures for simultaneous analysis. The study population was 302 trekkers from a previous investigation who provided self-reported symptoms upon arrival at 3840 m during hikes through altitudes of 1500 to 6200 m. AMS score (Hackett scale) was estimated by linear regression and the probability of an AMS diagnosis (Lake Louise criteria) by logistic regression. AMS score or probability was significantly associated with exposure day and altitude. Increased altitude over the prior 3 days resulted in higher estimated AMS score or probability and decreased altitude in lower score or probability. The odds ratio (OR) of AMS was 3.6 if not on acetazolamide. Females appeared slightly more susceptible than males (1.5 OR). The approach offers the advantages of (1) improved statistical power by combining exposures, (2) insight into the dose-response relationship of altitude exposure and AMS risk, (3) quantitative tests for the significance of factors that might affect AMS susceptibility, and (4) practical tools to track individual climbers and plan operational ascents.

Acclimatization↗