PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Models, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 667 records · Page 37Linked to original sources

Teenage smoking: a longitudinal analysis.

This paper examines the binary recurrent outcome "teenage smoking" within a statistical modelling paradigm. The proposed statistical modelling relates smoking to a set of explanatory variables, which include subjective as well as objective measures. In order to assess the degree to which explanatory variables influence smoking, an adequate statistical model must handle the possibility that substantial variation between respondents will be due to omitted variables, multicollinearity and past behaviour. An earlier paper, using a secondary cross-sectional data source, concluded that an investigation of smoking needs to be based on longitudinal data using appropriate statistical modelling. The same data source provided observations on young adults over a period of 2 years. For comparison purposes, the same cross-sectional model was fitted to the longitudinal data. The results suggest there may be substantial heterogeneity due to omitted variables in the data and complex inter-relationships between observed explanatory variables leading to underestimation. Longitudinal data provide additional flexibility to control for omitted variables and are necessary to investigate dynamic social processes such as smoking. The results from our analysis suggest that the effects of variables reported in the literature on teenage smoking may be overestimated. For example, the role of peer pressure may not be as clear as it has been made out to be.

Adolescent↗

The natural history of depression in general practice: stochastic models.

Three statistical models are presented to describe different aspects of the natural history of depression (as recognized by a general practitioner) during a 20-year study of a single general practice. After controlling for age, the prevalence of recognized depression in men increased during the 20 years of the study (1957-76), but there was little change in women. In any given year, however, women are both more likely to become depressed than men, and are less likely to recover. The changes in prevalence are shown to be due to changes in observed incidence and recovery rates. Taking into account attendances over the previous 5 years at which a diagnosis of depression was made, the models enable one to predict, for example incidence rates for 'first episodes' of recognized depression, and recovery rates for 'chronic' patients. In all situations there is a significant sex difference (women always being more likely to be recognized as depressed), but this difference is smaller at the end than at the beginning of the 20 years.

Adult↗

A toxicity estimation model.

A statistical model has been developed for estimation of acute toxicity. The model, currently operational for rat oral LD50, permits the estimation of rat oral LD50 for untested chemical compounds. Only the chemical structure, partition coefficient, and molecular weight for a compound are needed for estimation purposes. The chemical structure is partitioned into substructural fragments using the CIDS fragment keys. A regression model is developed on the basis of 425 compounds. A test of the regression equation with 100 compounds not used in its design shows that 56 percent of the compounds are predicted with less than 0.4 log unit deviation between extimated and measured LD50. This toxicity estimation model can be readily adapted to other species and to other measures of toxicity by the use of suitable design data bases. The model also identifies the contribution to toxicity of the fragments and physical characteristics. The use of this model can materially reduce the amount of toxicological testing for new compounds. It also permits the ranking of potentially toxic compounds to allow the most likely candidates to be tested. The method may also prove applicable to the determination of optimum dosages for new drugs.

Animals↗

Modeling controlled nutrient release from a population of polymer coated fertilizers: statistically based model for diffusion release.

A statistically based model for describing the release from a population of polymer coated controlled release fertilizer (CRF) granules by the diffusion mechanism was constructed. The model is based on a mathematical-mechanistic description of the release from a single granule of a coated CRF accounting for its complex and nonlinear nature. The large variation within populations of coated CRFs poses the need for a statistically based approach to integrate over the release from the individual granules within a given population for which the distribution and range of granule radii and coating thickness are known. The model was constructed and verified using experimentally determined parameters and release curves of polymer-coated CRFs. A sensitivity analysis indicated the importance of water permeability in controlling the lag period and that of solute permeability in governing the rate of linear release and the total duration of the release. Increasing the mean values of normally distributed granule radii or coating thickness, increases the lag period and the period of linear release. The variation of radii and coating thickness, within realistic ranges, affects the release only when the standard deviation is very large or when water permeability is reduced without affecting solute permeability. The model provides an effective tool for designing and improving agronomic and environmental effectiveness of polymer-coated CRFs.

Diffusion↗

Fast and robust parameter estimation for statistical partial volume models in brain MRI.

Due to the finite spatial resolution of imaging devices, a single voxel in a medical image may be composed of mixture of tissue types, an effect known as partial volume effect (PVE). Partial volume estimation, that is, the estimation of the amount of each tissue type within each voxel, has received considerable interest in recent years. Much of this work has been focused on the mixel model, a statistical model of PVE. We propose a novel trimmed minimum covariance determinant (TMCD) method for the estimation of the parameters of the mixel PVE model. In this method, each voxel is first labeled according to the most dominant tissue type. Voxels that are prone to PVE are removed from this labeled set, following which robust location estimators with high breakdown points are used to estimate the mean and the covariance of each tissue class. Comparisons between different methods for parameter estimation based on classified images as well as expectation--maximization-like (EM-like) procedure for simultaneous parameter and partial volume estimation are reported. The robust estimators based on a pruned classification as presented here are shown to perform well even if the initial classification is of poor quality. The results obtained are comparable to those obtained using the EM-like procedure, but require considerably less computation time. Segmentation results of real data based on partial volume estimation are also reported. In addition to considering the parameter estimation problem, we discuss differences between different approximations to the complete mixel model. In summary, the proposed TMCD method allows for the accurate, robust, and efficient estimation of partial volume model parameters, which is crucial to a variety of brain MRI data analysis procedures such as the accurate estimation of tissue volumes and the accurate delineation of the cortical surface.

Artifacts↗

Visualization of the variability of 3D statistical shape models by animation.

Models of the 3D shape of anatomical objects and the knowledge about their statistical variability are of great benefit in many computer assisted medical applications like images analysis, therapy or surgery planning. Statistical model of shapes have successfully been applied to automate the task of image segmentation. The generation of 3D statistical shape models requires the identification of corresponding points on two shapes. This remains a difficult problem, especially for shapes of complicated topology. In order to interpret and validate variations encoded in a statistical shape model, visual inspection is of great importance. This work describes the generation and interpretation of statistical shape models of the liver and the pelvic bone.

Humans↗

Dirichlet mixtures: a method for improved detection of weak but significant protein sequence homology.

We present a method for condensing the information in multiple alignments of proteins into a mixture of Dirichlet densities over amino acid distributions. Dirichlet mixture densities are designed to be combined with observed amino acid frequencies to form estimates of expected amino acid probabilities at each position in a profile, hidden Markov model or other statistical model. These estimates give a statistical model greater generalization capacity, so that remotely related family members can be more reliably recognized by the model. This paper corrects the previously published formula for estimating these expected probabilities, and contains complete derivations of the Dirichlet mixture formulas, methods for optimizing the mixtures to match particular databases, and suggestions for efficient implementation.

Algorithms↗

A minimum description length approach to statistical shape modeling.

We describe a method for automatically building statistical shape models from a training set of example boundaries/surfaces. These models show considerable promise as a basis for segmenting and interpreting images. One of the drawbacks of the approach is, however, the need to establish a set of dense correspondences between all members of a set of training shapes. Often this is achieved by locating a set of "landmarks" manually on each training image, which is time consuming and subjective in two dimensions and almost impossible in three dimensions. We describe how shape models can be built automatically by posing the correspondence problem as one of finding the parameterization for each shape in the training set. We select the set of parameterizations that build the "best" model. We define "best" as that which minimizes the description length of the training set, arguing that this leads to models with good compactness, specificity and generalization ability. We show how a set of shape parameterizations can be represented and manipulated in order to build a minimum description length model. Results are given for several different training sets of two-dimensional boundaries, showing that the proposed method constructs better models than other approaches including manual landmarking-the current gold standard. We also show that the method can be extended straightforwardly to three dimensions.

Algorithms↗

A statistical multiprobe model for analyzing cis and trans genes in genetical genomics experiments with short-oligonucleotide arrays.

Short-oligonucleotide arrays typically contain multiple probes per gene. In genetical genomics applications a statistical model for the individual probe signals can help in separating "true" differential mRNA expression from "ghost" effects caused by polymorphisms, misdesigned probes, and batch effects. It can also help in detecting alternative splicing, start, or termination.

Animals↗

Factors affecting alkalinity generation by successive alkalinity-producing systems: regression analysis.

Use of successive alkalinity-producing systems (SAPS) for treatment of acidic mine drainage (AMD) has grown in recent years. However, inconsistent performance has hampered widespread acceptance of this technology. This research was conducted to determine the influence of system design and influent AMD chemistry on net alkalinity generation by SAPS. Monthly observations were obtained from eight SAPS cells in southern West Virginia and southwestern Virginia. Analysis of these data revealed strong, positive correlations between net alkalinity generation and three variables: the natural log of limestone residence time, influent dissolved Fe concentration, and influent non-Mn acidity. A statistical model was constructed to describe SAPS performance. Subsequent analysis of data obtained from five systems in western Pennsylvania (calibration data set) was used to reevaluate the model form, and the statistical model was adjusted using the combined data sets. Limestone residence time exhibited a strong, positive logarithmic correlation with net alkalinity generation, indicating net alkalinity generation occurs most rapidly within the first few hours of AMD-limestone contact and additional residence time yields diminishing gains in treatment. Influent Fe and non-Mn acidity concentrations both show strong positive linear relationships with net alkalinity generation, reflecting the increased solubility of limestone under acidic conditions. These relationships were present in the original and the calibration data sets, separately, and in the statistical model derived from the combined data set. In the combined data set, these three factors accounted for 68% of the variability in SAPS systems performance.

Calibration↗

Statistical properties of a DNA sample under the finite-sites model.

Statistical properties of a DNA sample from a random-mating population of constant size are studied under the finite-sites model. It is assumed that there is no migration and no recombination occurs within the locus. A Markov process model is used for nucleotide substitution, allowing for multiple substitutions at a single site. The evolutionary rates among sites are treated as either constant or variable. The general likelihood calculation using numerical integration involves intensive computation and is feasible for three or four sequences only, it may be used for validating approximate algorithms. Methods are developed to approximate the probability distribution of the number of segregating sites in a random sample of n sequences, with either constant or variable substitution rates across sites. Calculations using parameter estimates obtained for human D-loop mitochondrial DNAs show that among-site rate variation has a major effect on the distribution of the number of segregating sites; the distribution under the finite-sites model with variable rates among sites is quite different from that under the infinite-sites model.

Animals↗

Analysis of PIN1 WW domain through a simple statistical mechanics model.

We have applied a simple statistical mechanics Go-like model to the analysis of the PIN1 WW domain, resorting to mean field and Monte Carlo techniques to characterize its thermodynamics, and comparing the results with the wealth of available experimental data. PIN1 WW domain is a 39-residue protein fragment which folds on an antiparallel beta-sheet, thus representing an interesting model system to study the behavior of these secondary structure elements. Results show that the model correctly reproduces the two-state behavior of the protein, and also the trends of the experimental phi(T) values. Moreover, there is a good agreement between Monte Carlo results and the mean field ones, which can be obtained with a substantially smaller computational effort.

Models, Statistical↗

Spontaneous speech recognition using a statistical coarticulatory model for the vocal-tract-resonance dynamics.

A statistical coarticulatory model is presented for spontaneous speech recognition, where knowledge of the dynamic, target-directed behavior in the vocal tract resonance is incorporated into the model design, training, and in likelihood computation. The principal advantage of the new model over the conventional HMM is the use of a compact, internal structure that parsimoniously represents long-span context dependence in the observable domain of speech acoustics without using additional, context-dependent model parameters. The new model is formulated mathematically as a constrained, nonstationary, and nonlinear dynamic system, for which a version of the generalized EM algorithm is developed and implemented for automatically learning the compact set of model parameters. A series of experiments for speech recognition and model synthesis using spontaneous speech data from the Switchboard corpus are reported. The promise of the new model is demonstrated by showing its consistently superior performance over a state-of-the-art benchmark HMM system under controlled experimental conditions. Experiments on model synthesis and analysis shed insight into the mechanism underlying such superiority in terms of the target-directed behavior and of the long-span context-dependence property, both inherent in the designed structure of the new dynamic model of speech.

Algorithms↗

Evaluation of 3D correspondence methods for model building.

The correspondence problem is of high relevance in the construction and use of statistical models. Statistical models are used for a variety of medical application, e.g. segmentation, registration and shape analysis. In this paper, we present comparative studies in three anatomical structures of four different correspondence establishing methods. The goal in all of the presented studies is a model-based application. We have analyzed both the direct correspondence via manually selected landmarks as well as the properties of the model implied by the correspondences, in regard to compactness, generalization and specificity. The studied methods include a manually initialized subdivision surface (MSS) method and three automatic methods that optimize the object parameterization: SPHARM, MDL and the covariance determinant (DetCov) method. In all studies, DetCov and MDL showed very similar results. The model properties of DetCov and MDL were better than SPHARM and MSS. The results suggest that for modeling purposes the best of the studied correspondence method are MDL and DetCov.

Algorithms↗

A statistical mechanical model of the pre- and subtransitions of lecithin membranes.

A statistical mechanical model with experimentally proved facts as starting points is presented. This model explains on molecular level, the pre- and subtransitions appearing in lipid membranes. The model describes the main features of the transitions, the hysteresis of the subtransition and the mobility changes of the heads and chains at these transitions. The model was expanded for phosphatidylcholine homologues with arbitrary chain lengths, and a qualitative agreement in the case of pretransition as far as a quantitative one for the subtransition were found.

Chemical Phenomena↗

Statistical face models for the rediction of soft-tissue deformations after orthognathic osteotomies.

This paper describes a technique to approximately predict the facial morphology after standardized orthognathic ostoetomies. The technique only relies on the outer facial morphology represented as a set of surface points and does not require computed tomography (CT) images as input. Surface points may either be taken from 3D surface scans or from 3D positions palpated on the face using a tracking system. The method is based on a statistical model generated from a set of pre- and postoperative 3D surface scans of patients that underwent the same standardized surgery. The model contains both the variability of preoperative facial morphologies and the corresponding postoperative deformations. After fitting the preoperative part to 3D data from a new patient the preoperative face is approximated by the model and the preiction of the postoperative morphology can be extracted at the same time. We built a model based on a set of 15 patient data sets and tested the predictive power in leave-one-out tests for a set of relevant cephalometric landmarks. The average prediction error was found to be between 0.3 and 1.2 mm at all important facial landmarks in the relevant areas of upper and lower jaw. Thus the technique provides an easy and powerful way of prediction which avoids time, cost and radiation required by other prediction techniques such as those based on CT scans.

Computer Simulation↗

Using statistical image models for objective evaluation of spot detection in two-dimensional gels.

Protein spot detection is central to the analysis of two-dimensional electrophoresis gel images. There are many commercially available packages, each implementing a protein spot detection algorithm. Despite this, there have been relatively few studies comparing the performance characteristics of the different packages. This is in part due to the fact that different packages employ different sets of user-adjustable parameters. It is also partly due to the fact that the images are complex. To carry out an evaluation, "ground truth" data specifying spot position, shape and intensities needs to be defined subjectively on selected test images. We address this problem by proposing a method of evaluation using synthetic images with unambiguous interpretation. The characteristics of the spots in the synthetic images are determined from statistical models of the shape, intensity, size, spread and location of real spot data. The distribution of parameters is described using a Gaussian mixture model obtained from training images. The synthetic images allow us to investigate the effects of individual image properties, such as signal-to-noise ratios and degree of spot overlap, by measuring quantifiable outcomes, e.g. accuracy of spot position, false positive and false negative detection. We illustrate the approach by carrying out quantitative evaluations of spot detection on a number of widely used analysis packages.

Algorithms↗