PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Models, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

A fast and flexible statistical model for large-scale population genotype data: applications to inferring missing genotypes and haplotypic phase.

We present a statistical model for patterns of genetic variation in samples of unrelated individuals from natural populations. This model is based on the idea that, over short regions, haplotypes in a population tend to cluster into groups of similar haplotypes. To capture the fact that, because of recombination, this clustering tends to be local in nature, our model allows cluster memberships to change continuously along the chromosome according to a hidden Markov model. This approach is flexible, allowing for both "block-like" patterns of linkage disequilibrium (LD) and gradual decline in LD with distance. The resulting model is also fast and, as a result, is practicable for large data sets (e.g., thousands of individuals typed at hundreds of thousands of markers). We illustrate the utility of the model by applying it to dense single-nucleotide-polymorphism genotype data for the tasks of imputing missing genotypes and estimating haplotypic phase. For imputing missing genotypes, methods based on this model are as accurate or more accurate than existing methods. For haplotype estimation, the point estimates are slightly less accurate than those from the best existing methods (e.g., for unrelated Centre d'Etude du Polymorphisme Humain individuals from the HapMap project, switch error was 0.055 for our method vs. 0.051 for PHASE) but require a small fraction of the computational cost. In addition, we demonstrate that the model accurately reflects uncertainty in its estimates, in that probabilities computed using the model are approximately well calibrated. The methods described in this article are implemented in a software package, fastPHASE, which is available from the Stephens Lab Web site.

Calibration↗

A statistical model for the determination of the optimal metric in factor analysis of medical image sequences (FAMIS).

A statistical model is added to the conventional physical model underlying factor analysis of medical image sequences (FAMIS). It allows a derivation of the optimal metric to be used for the orthogonal decomposition involved in FAMIS. The oblique analysis of FAMIS is extended to take this optimal metric into account. The case of scintigraphic image sequences is used. We derive in this case that the optimal decomposition is obtained by correspondence analysis. A scintigraphic dynamic study illustrates the practical consequences of the use of the optimal metric in FAMIS.

Factor Analysis, Statistical↗

Predictive value of statistical models.

A review is given of different ways of estimating the error rate of a prediction rule based on a statistical model. A distinction is drawn between apparent, optimum and actual error rates. Moreover it is shown how cross-validation can be used to obtain an adjusted predictor with smaller error rate. A detailed discussion is given for ordinary least squares, logistic regression and Cox regression in survival analysis. Finally, the splitsample approach is discussed and demonstrated on two data sets.

Calibration↗

Empirical evaluation of statistical models for counts or rates.

We consider methods for selecting the joint specification of the mean and variance functions in statistical models for rates or counts. Based on analyses of diagnosis-specific hospital discharge rates in Michigan, we show that a Poisson model with an extra variance component for the systematic variation is superior to several other probability models with regard to specification of the error structure. Further, the deviance residual appears superior to the Pearson residual. The proper specification of such variation is crucial for many types of analyses, such as identification of outliers and regression analyses designed to explain the systematic component of the variation.

Analysis of Variance↗

Use of brain biopsy for diagnostic evaluation of patients with suspected herpes simplex encephalitis: a statistical model and its clinical implications. NIAID Collaborative Antiviral Study Group.

Using the decision analysis technique and multivariate regression methods, a statistical model was established to define the utility of brain biopsy for diagnostic evaluation of patients with suspected herpes simplex encephalitis (HSE). Two strategies were compared: strategy I, brain biopsy with acyclovir (ACV) treatment for 10 days in biopsy-positive patients, and strategy II, ACV therapy without brain biopsy. Strategy I resulted in a greater 6-month survival rate when the likelihood of patients having HSE was less than 70%. Based on the current estimated prevalence of HSE (for patients with suspected HSE) of 35%, strategy I showed a slight advantage of a 3.2% increase in 6-month survival rate. An individual patient's chance of a positive brain biopsy can be predicted using a mathematical equation based on several important clinical assessments. This equation in conjunction with the decision analysis is a useful guide for the clinical management of patients with regard to brain biopsy.

Acyclovir↗

Statistical model to determine the relationship of response and survival in patients with advanced ovarian cancer treated with chemotherapy.

BACKGROUND: A statistically appropriate analysis of the association between survival and response measures in patients with ovarian cancer could help to define the role of response rate in planning, monitoring, and interpreting the results of clinical trials. PURPOSE: This study was designed to investigate the relationship between antitumor response determined by clinical or pathological means and survival in patients with advanced ovarian cancer with no previous treatment. We focused on avoiding the limitations of the usual approach of comparing durations of survival for patients responding to therapy with those for nonresponders. METHODS: A new meta-analytic statistical model we developed was used to analyze data from 26 randomized clinical trials published between 1975 and 1989. Our model incorporates intra-study and inter-study sources of variability in the estimates of response and survival. The study also addresses the methodological problems of evaluating response as a surrogate end point and the relevance of this association to clinical decision making and the design of clinical trials. RESULTS: For 13 studies in which response was pathologically assessed, an improvement in surgically documented complete response rate was associated with an increase in median survival. A similar but apparently smaller effect was found for the association between objective clinical response and median survival in the 25 studies reporting these data. CONCLUSIONS: These results suggest that therapeutic measures must produce large improvements in clinical response rates to achieve meaningful effects on median survival. Improvement in surgically documented complete response rate appears to be more strongly associated with increased median survival and, hence, might be used for interim monitoring in clinical trials, but the role of second-look procedures in clinical management is controversial.

Antineoplastic Combined Chemotherapy Protocols↗

Performance of a statistical model to predict stroke outcome in the context of a large, simple, randomized, controlled trial of feeding.

BACKGROUND AND PURPOSE: Statistical models to predict the outcome of stroke patients have several uses. Their utility depends on their predictive accuracy in patients other than those on whom they were developed (ie, external validity). We sought to test the external validity of some recently described models in patients enrolled in the FOOD (Feed Or Ordinary Diet) trial: a large randomized trial evaluating feeding policies in patients with stroke. METHODS: The predictive variables were collected during a telephone call to randomize the patient a median of 5 days after stroke onset. Patients were followed up 6 months later to establish their survival, functional status, and residence. Charts were plotted to demonstrate the discrimination and calibration of the models. RESULTS: The models performed well in the first 2955 patients enrolled and followed up in the FOOD trial. The area under the receiver operating characteristic curves varied between 0.78 and 0.81 (with 0.5 indicating no discrimination and 1.0 indicating perfect discrimination). The discrimination was marginally better for patients enrolled within the first day of stroke than later. The models tended to provide rather pessimistic predictions in all groups except those predicted to have a high likelihood of surviving free of dependency. CONCLUSIONS: As one might predict, the discriminatory power in the selected cohort of trial patients was marginally less good than in previously studied unselected cohorts used to test their external validity. These models provide a well-tested tool for stratification in trials, comparing outcomes in different cohorts and examining the additional predictive power of novel factors.

Aged↗

Statistical models for trisomic phenotypes.

Certain genetic disorders are rare in the general population but more common in individuals with specific trisomies, which suggests that the genes involved in the etiology of these disorders may be located on the trisomic chromosome. As with all aneuploid syndromes, however, a considerable degree of variation exists within each phenotype so that any given trait is present only among a subset of the trisomic population. We have previously presented a simple gene-dosage model to explain this phenotypic variation and developed a strategy to map genes for such traits. The mapping strategy does not depend on the simple model but works in theory under any model that predicts that affected individuals have an increased likelihood of disomic homozygosity at the trait locus. This paper explores the robustness of our mapping method by investigating what kinds of models give an expected increase in disomic homozygosity. We describe a number of basic statistical models for trisomic phenotypes. Some of these are logical extensions of standard models for disomic phenotypes, and some are more specific to trisomy. Where possible, we discuss genetic mechanisms applicable to each model. We investigate which models and which parameter values give an expected increase in disomic homozygosity in individuals with the trait. Finally, we determine the sample sizes required to identify the increased disomic homozygosity under each model. Most of the models we explore yield detectable increases in disomic homozygosity for some reasonable range of parameter values, usually corresponding to smaller trait frequencies. It therefore appears that our mapping method should be effective for a wide variety of moderately infrequent traits, even though the exact mode of inheritance is unlikely to be known.

Abortion, Spontaneous↗

Statistical model building and model criticism for human circadian data.

Mathematical models have played an important role in the analysis of circadian systems. The models include simulation of differential equation systems to assess the dynamic properties of a circadian system and the use of statistical models, primarily harmonic regression methods, to assess the static properties of the system. The dynamical behaviors characterized by the simulation studies are the response of the circadian pacemaker to light, its rate of decay to its limit cycle, and its response to the rest-activity cycle. The static properties are phase, amplitude, and period of the intrinsic oscillator. Formal statistical methods are not routinely employed in simulation studies, and therefore the uncertainty in inferences based on the differential equation models and their sensitivity to model specification and parameter estimation error cannot be evaluated. The harmonic regression models allow formal statistical analysis of static but not dynamical features of the circadian pacemaker. The authors present a paradigm for analyzing circadian data based on the Box iterative scheme for statistical model building. The paradigm unifies the differential equation-based simulations (direct problem) and the model fitting approach using harmonic regression techniques (inverse problem) under a single schema. The framework is illustrated with the analysis of a core-temperature data series collected under a forced desynchrony protocol. The Box iterative paradigm provides a framework for systematically constructing and analyzing models of circadian data.

Adult↗

Statistical modelling as an aid to the design of retail sampling plans for mycotoxins in food.

A study has been carried out to assess appropriate statistical models for use in evaluating retail sampling plans for the determination of mycotoxins in food. A compound gamma model was found to be a suitable fit. A simulation model based on the compound gamma model was used to produce operating characteristic curves for a range of parameters relevant to retail sampling. The model was also used to estimate the minimum number of increments necessary to minimize the overall measurement uncertainty. Simulation results showed that measurements based on retail samples (for which the maximum number of increments is constrained by cost) may produce fit-for-purpose results for the measurement of ochratoxin A in dried fruit, but are unlikely to do so for the measurement of aflatoxin B1 in pistachio nuts. In order to produce a more accurate simulation, further work is required to determine the degree of heterogeneity associated with batches of food products. With appropriate parameterization in terms of physical and biological characteristics, the systems developed in this study could be applied to other analyte/matrix combinations.

Aflatoxin B1↗

Statistical models of outcome in malpractice lawsuits involving death or neurologically impaired infants.

The objective was to determine whether factors could be identified in medical and legal records that are associated with the successful defense of obstetrical malpractice cases involving the death or neurological impairment of infants. Obstetrical claims (169) closed by PROMUTUAL between January 1, 1990, and December 31, 1994, were retrospectively abstracted and analyzed to identify associations between medical and legal factors, and the medicolegal outcome. Multivariable analysis identifies that the use of pitocin, diagnosis of asphyxia, a delay in delivery, and the use of multiple defense expert witnesses decreased the chances of a successful defense. Two statistical models explaining indemnity payment were developed. The first, based on medical outcome, showed an increased indemnity payment when a case involved major neurological deficits, diagnosis of asphyxia, newborn seizures, later year of delivery, and participation of a particular defense firm. Perinatal or childhood death and the use of pitocin were indicators of a decrease in payment. The second model was based on long-term care requirements. In this model, indicators of increased indemnity payment were: nonreassuring intrapartum fetal heart rate tracing, later year of delivery, intensity of long-term care required, and participation of a particular defense law firm. Perinatal or childhood death, the use of pitocin, and settlement date increasingly removed from the occurrence date were the determinants of decreased payments in this model. Finally, the presence of major neurological deficits, the prolongation of a case, and the involvement of multiple law firms and defense witnesses increased the expense charged to and paid by the insurance company. Using the medical, legal, and financial data relevant to 169 obstetrical cases closed by one malpractice insurance carrier between 1990 and 1994, statistical models with potential predictive values for future malpractice claims involving neurologically impaired infants were constructed. These models may help determine in advance the chance a future case has for successful defense and the likely amount of expense and indemnity dollars that will be paid out to settle and defend it.

Adolescent↗

Statistical modelling of general practice medicine for computer assisted data entry in electronic medical record systems.

Electronic medical record (EMR) systems have much potential, however, there are still a number of issues that need to be resolved before EMRs are widely accepted. One of these issues is the data input task, a potentially serious practical barrier to on-line medical computer usage. This paper reports the empirical modelling of data input requirements for physicians who use a problem-orientated medical record system. Three statistical models (Bayesian conditional probability, multiple linear regression and discriminant analysis) to predict drug treatment given problem diagnoses are derived from EMRs of 2500 general Practice encounters. Two metrics are used to measure the predictive power of the models considering both the number of drugs correctly predicted and the strength with which the models predict them. The models are tested on 500 unseen records from the same patient-physician population and the data used to build the models. The Bayesian model produces the best predictions on unseen data and is also the easiest model to compute. A prototype interface that enables new patient cases to be entered is constructed to demonstrate how the predictive power of the model can translate into benefits in the data entry task.

Bayes Theorem↗

A physician-based architecture for the construction and use of statistical models.

Physicians need specially tailored computer tools to take advantage of published research results. We present a knowledge-based computer framework--the physician-based (PB) architecture--for constructing such tools, and we use the problem of physicians' interpretation of two-arm parallel randomized clinical trials (TAPRCT) as a working example. Statistical models are represented by influence diagrams. The interpretation of influence-diagram elements are mapped into users' language in a domain-specific, physician-based user interface, called a patient-flow diagram. Statistical-model transformations that maintain the semantic relationships of the model and that embody clinical-epidemiological knowledge are encoded in a mediating structure called the cohort-state diagram. The algorithm that coordinates the interactions among the knowledge representations uses modular actions called construction steps. This architecture has been implemented in a Bayesian system, called THOMAS, that supports physician decision making in light of TAPRCT data. This support entails assessing clinical significance, prior beliefs, and methodological concerns. We suggest that the PB architecture applies to a wide range of statistical tools and users.

Algorithms↗

A multi-variate statistical model integrating passive sampler and meteorology data to predict the frequency distributions of hourly ambient ozone (O3) concentrations.

A multi-variate, non-linear statistical model is described to simulate passive O3 sampler data to mimic the hourly frequency distributions of continuous measurements using climatologic O3 indicators and passive sampler measurements. The main meteorological parameters identified by the model were, air temperature, relative humidity, solar radiation and wind speed, although other parameters were also considered. Together, air temperature, relative humidity and passive sampler data by themselves could explain 62.5-67.5% (R(2)) of the corresponding variability of the continuously measured O3 data. The final correlation coefficients (r) between the predicted hourly O3 concentrations from the passive sampler data and the true, continuous measurements were 0.819-0.854, with an accuracy of 92-94% for the predictive capability. With the addition of soil moisture data, the model can lead to the first order approximation of atmospheric O3 flux and plant stomatal uptake. Additionally, if such data are coupled to multi-point plant response measurements, meaningful cause-effect relationships can be derived in the future.

Air Pollutants↗

Tropical geometry of statistical models.

This article presents a unified mathematical framework for inference in graphical models, building on the observation that graphical models are algebraic varieties. From this geometric viewpoint, observations generated from a model are coordinates of a point in the variety, and the sum-product algorithm is an efficient tool for evaluating specific coordinates. Here, we address the question of how the solutions to various inference problems depend on the model parameters. The proposed answer is expressed in terms of tropical algebraic geometry. The Newton polytope of a statistical model plays a key role. Our results are applied to the hidden Markov model and the general Markov model on a binary tree.

Algorithms↗

Efficient statistical modelling of longitudinal data.

A new class of statistical models is proposed for the analysis of longitudinal data, especially those from growth studies. The models are all derived from a simple univariate two-level polynomial model. It is shown that they make efficient use of available data, and can handle a very wide range of problems. They have several important advantages over existing procedures.

Age Factors↗

Anatomical statistical models and their role in feature extraction.

A detailed model of the shape of anatomical structures can significantly improve the ability to segment such structures from medical images. Statistical models representing the variation of shape and appearance can be constructed from suitably annotated training sets. Such models can be used to synthesize images of anatomy, and to search new images to accurately locate the structures of interest, even in the presence of noise and clutter. In this paper we summarize recent work on constructing and using such models, and demonstrate their application to several domains.

Brain↗