PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Models, Statistical”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Angioarchitecture associated with haemorrhage in cerebral arteriovenous malformations: a prognostic statistical model.

The overall haemorrhagic risk of a cerebral arteriovenous malformation (cAVM) is 2-4% per year. However, the individual risk of haemorrhage has never been determined. This study was undertaken to assess the haemorrhage risk of an individual cAVM. Neuroangiographic findings of 160 cAVM were analysed retrospectively, looking at 30 angiographic features. A statistical model was established by logistic regression to evaluate the risk of an individual cAVM. We statistically correlated 15 parameters with the haemorrhage risk. The statistical model includes five independent parameters. Four are unfavourable: exclusively deep drainage, venous stenoses, venous reflux and the radio of afferent to efferent systems; one is favourable: venous recruitment. This model quantifies the individual risk of haemorrhage. When this model is applied to the population studied, the error rate is 5%. This model can contribute to therapeutic strategy, and to a better understanding of the natural history of cAVM.

Adolescent↗

Predictive accuracy study: comparing a statistical model to clinicians' estimates of outcomes after coronary bypass surgery.

BACKGROUND: The purpose of this study was to compare clinicians' prior probability estimates of operative mortality (OM) and prolonged intensive care unit stay (ICU) length of stay greater than 48 hours after coronary artery bypass graft surgery (CABG) with estimates derived from statistical models alone. METHODS: Nine clinicians estimated the predicted probability of OM and ICU stay greater than 48 hours from an abstract of information for each of 100 patients selected from the 1996 to 1997 database of 1,904 patients who underwent isolated CABG. Logistic regression models were used to calculate the predicted probability of OM and ICU stay greater than 48 hours for each patient. The study sample was split into two parts; clinicians were randomly given access to a predictive rule to guide their judgements for one part of the study. RESULTS: Clinicians' estimates were similar with or without access to the rule, and both parts of the study were therefore pooled. Clinicians significantly overestimated the probability of OM (model 6.3% +/- 1%, clinicians 7.6% +/- 3%, p = 0.0001) and ICU stay greater than 48 hours (model 25% +/- 2%, clinicians 28% +/- 1%, p = 0.0012). Clinicians' estimates of OM were not significantly higher than the model's for nonsurvivors (0.8% +/- 0.7%, p = 0.2), but were significantly higher for survivors (1.4% +/- 0.3%, p = 0.039). CONCLUSIONS: Clinicians trusted their own empiric estimates rather than a predictive rule and overestimated the probability of OM and ICU stay greater than 48 hours.

Aged↗

Statistical modelling for clinical mastitis in the dairy cow: problems and solutions.

Modelling case occurrence and risk factors for clinical mastitis, as a key multifactorial disease in the dairy cow, requires statistical models. The type of model used depends on the choice of perception or the study level: herd, lactation, animal, udder and quarter. The validity of the tests that are performed through these models is especially ensured when hypotheses of independence between statistical units are respected, and when the model adjustments do not involve overdispersion faced with the observed data. In the article, the main sources of overdispersion are identified according to the different levels of perception of mastitis risk. Then, the proposed solutions to control for overdispersion at each study level are discussed and the difficulty to compare the study results is highlighted through a variety of methodological choices of the authors. Two main categories of models are used for modelling clinical mastitis, i.e. generalist exploratory models and explanatory designed models. The contribution of the explanatory models to improve modelling accuracy and relevance is documented through the two main published methodological approaches, the first one being based on a states model, and the second on a survival model. The integration and optimisation of such explanatory modelling methods should be possible in the future in order to develop a more global explanatory model including herd risk factors, which could pertinently predict udder infections (both clinical and subclinical) at the cow, lactation, or even udder and quarter levels.

Animals↗

Validity of linear regression in method comparison studies: is it limited by the statistical model or the quality of the analytical input data?

We compared the application of ordinary linear regression, Deming regression, standardized principal component analysis, and Passing-Bablok regression to real-life method comparison studies to investigate whether the statistical model of regression or the analytical input data have more influence on the validity of the regression estimates. We took measurements of serum potassium as an example for comparisons that cover a narrow data range and measurements of serum estradiol-17beta as an example for comparisons that cover a wide data range. We demonstrate that, in practice, it is not the statistical model but the quality of the analytical input data that is crucial for interpretation of method comparison studies. We show the usefulness of ordinary linear regression, in particular, because it gives a better estimate of the standard deviation of the residuals than the other procedures. The latter is important for distinguishing whether the observed spread across the regression line is caused by the analytical imprecision alone or whether sample-related effects also contribute. We further demonstrate the usefulness of linear correlation analysis as a first screening test for the validity of linear regression data. When ordinary linear regression (in combination with correlation analysis) gives poor estimates, we recommend investigating the analytical reason for the poor performance instead of assuming that other linear regression procedures add substantial value to the interpretation of the study. This investigation should address whether (a) the x and y data are linearly related; (b) the total analytical imprecision (s(a,tot)) is responsible for the poor correlation; (c) sample-related effects are present (standard deviation of the residuals >> s(a,tot)); (d) the samples are adequately distributed over the investigated range; and (e) the number of samples used for the comparison is adequate.

Chromatography, Ion Exchange↗

PyEvolve: a toolkit for statistical modelling of molecular evolution.

BACKGROUND: Examining the distribution of variation has proven an extremely profitable technique in the effort to identify sequences of biological significance. Most approaches in the field, however, evaluate only the conserved portions of sequences - ignoring the biological significance of sequence differences. A suite of sophisticated likelihood based statistical models from the field of molecular evolution provides the basis for extracting the information from the full distribution of sequence variation. The number of different problems to which phylogeny-based maximum likelihood calculations can be applied is extensive. Available software packages that can perform likelihood calculations suffer from a lack of flexibility and scalability, or employ error-prone approaches to model parameterisation. RESULTS: Here we describe the implementation of PyEvolve, a toolkit for the application of existing, and development of new, statistical methods for molecular evolution. We present the object architecture and design schema of PyEvolve, which includes an adaptable multi-level parallelisation schema. The approach for defining new methods is illustrated by implementing a novel dinucleotide model of substitution that includes a parameter for mutation of methylated CpG's, which required 8 lines of standard Python code to define. Benchmarking was performed using either a dinucleotide or codon substitution model applied to an alignment of BRCA1 sequences from 20 mammals, or a 10 species subset. Up to five-fold parallel performance gains over serial were recorded. Compared to leading alternative software, PyEvolve exhibited significantly better real world performance for parameter rich models with a large data set, reducing the time required for optimisation from approximately 10 days to approximately 6 hours. CONCLUSION: PyEvolve provides flexible functionality that can be used either for statistical modelling of molecular evolution, or the development of new methods in the field. The toolkit can be used interactively or by writing and executing scripts. The toolkit uses efficient processes for specifying the parameterisation of statistical models, and implements numerous optimisations that make highly parameter rich likelihood functions solvable within hours on multi-cpu hardware. PyEvolve can be readily adapted in response to changing computational demands and hardware configurations to maximise performance. PyEvolve is released under the GPL and can be downloaded from http://cbis.anu.edu.au/software.

Animals↗

A statistical model for high-resolution mapping of quantitative trait loci determining HIV dynamics.

Are there specific genes that control the pathogenesis of HIV infection? This question, which is of fundamental importance in designing personalized strategies of gene therapy to control HIV infection, can be examined by genetic mapping approaches. In this article, we present a new statistical model for unravelling the genetic mechanisms for the dynamic change of HIV that causes AIDS by marker-based linkage disequilibrium (LD) analyses. This new model is the extension of our functional mapping theory to integrate viral load trajectories within a genetic mapping framework. Earlier studies of HIV dynamics have led to various mathematical functions for modelling the kinetic curves of plasma virions and CD4 lymphocytes in HIV patients. Through incorporating these functions into the LD-based mapping procedure, we can identify and map individual quantitative trait loci (or QTL) responsible for viral pathogenesis. We derive a closed-form solution for estimating QTL allele frequency and marker-QTL linkage disequilibrium in the context of EM algorithm and implement the simplex algorithm to estimate the mathematical parameters describing the curve shapes of HIV pathogenesis. We performed different simulation scenarios based on currently used clinical designs in AIDS/HIV research to illustrate the utility and power of our model for genetic mapping of HIV dynamics. The implications of our model for genetic and genomic research into AIDS pathogenesis are discussed.

Acquired Immunodeficiency Syndrome↗

[Variability in scabies mites Sarcoptes scabiei De Geer (Acariformes, Sarcoptidae) in relation to scabies epidemiology. 1. A statistical model of female variability].

Individual variability of 235 grain mite females from Moscow and the Moscow Province was studied. The data were processed on a computer. Body proportions were used as a representative dimension index. The size of proterosomal scutellum was stable. A map of the chaetoid cover of notum is given, and its variability was studied. Statistical analysis of all signs was made. Frontal chaetoid and sejugal sections and the number of caudal chaetoids are stable. Naked, middle and abdominal chaetoid sections are the most variable. Linear sizes of naked section and the number of chaetoid sets on middle and abdominal sections correlate with the number of undeveloped chaetoids. The statistical model of variability may be useful for populational analysis of Sarcoptes forms in connection with scab epidemiology.

Animals↗

Statistical modelling of the determinants of historical exposure to bitumen and polycyclic aromatic hydrocarbons among paving workers.

INTRODUCTION: An industrial hygiene database has been constructed for the exposure assessment in a study of cancer risk among asphalt workers. AIM: To create models of bitumen and polycyclic aromatic hydrocarbons (PAH) exposure intensity among paving workers. METHODS: Individual exposure measurements from pavers (N = 1581) were collected from 8 countries. Correlation patterns between exposure measures were examined and factors affecting exposure were identified using statistical modelling. RESULTS: Inhalable dust appeared to be a good proxy of bitumen fume exposure. Bitumen fume and vapour levels were not correlated. Benzo(a)pyrene level appeared to be a good indicator of PAH exposure. All exposures steadily declined over the last 20 years. Mastic laying, re-paving, surface dressing, oil gravel paving and asphalt temperature were significant determinants of bitumen exposure. Coal tar use dictated PAH exposure levels. DISCUSSION: Bitumen fume, vapour and PAH have different determinants of exposure. For paving workers, exposure intensity can be assessed on the basis of time period and production characteristics.

Databases, Factual↗

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3·8 × 10-10 to 2·4 × 10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7·87 million to $287 000) and 29-fold (from $1·98 million to $69 100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0·01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans↗

A unifying statistical model for QTL mapping of genotype x sex interaction for developmental trajectories.

Most organisms display remarkable differences in morphological, anatomical, and developmental features between the two sexes. It has been recognized that these sex-dependent differences are controlled by an array of specific genetic factors, mediated through various environmental stimuli. In this paper, we present a unifying statistical model for mapping quantitative trait loci (QTL) that are responsible for sexual differences in growth trajectories during ontogenetic development. This model is derived within the maximum likelihood context, incorporated by sex-stimulated differentiation in growth form that is described by mathematical functions. A typical structural model is implemented to approximate time-dependent covariance matrices for longitudinal traits. This model allows for a number of biologically meaningful hypothesis tests regarding the effects of QTL on overall growth trajectories or particular stages of development. It is particularly powerful to test whether and how the genetic effects of QTL are expressed differently in different sexual backgrounds. Our model has been employed to map QTL affecting body mass growth trajectories in both male and female mice of an F2 population derived from the large (LG/J) and small (SM/J) mouse strains. We detected four growth QTL on chromosomes 6, 7, 11, and 15, two of which trigger different effects on growth curves between the two sexes. All the four QTL display significant genotype-sex interaction effects on the timing of maximal growth rate in the ontogenetic growth of mice. The implications of our model for studying the genetic architecture of growth trajectories and its extensions to some more general situations are discussed.

Animals↗

A statistical model to analyse quantitative trait locus interactions for HIV dynamics from the virus and human genomes.

Viruses can be considered 'parasites' because they cannot survive outside of a host. The progression rate to AIDS caused by human immunodeficiency virus type-1 (HIV-1) is therefore a consequence of HIV-host cell interactions. In this article, we present an innovative statistical model for detecting the effects of genetic interactions on HIV-1 dynamics triggered by different quantitative trait loci (QTL) from the HIV and human genomes. Our model integrates the principles of functional mapping for longitudinal traits and of linkage disequilibrium analysis for high-resolution mapping of QTL within the maximum likelihood context and is implemented with the EM algorithm. We performed Monte Carlo simulation studies to investigate the impacts of different heritability levels and sample sizes on the power to detect interacting QTL. Our model allows for the tests of a number of clinically meaningful hypotheses and provides a powerful tool for unravelling the genetic architecture of HIV-1 dynamics and therefore AIDS progression rate.

Algorithms↗

Statistical model for prediction of retrospective exposure to ethylene oxide in an occupational mortality study.

Since direct measures of individual exposure seldom exist for the entire period of an occupational mortality study, retrospective exposure estimates are necessary. This is often done in a subjective manner involving a consensus of opinion from a panel of epidemiologists and industrial hygienists. An alternative method utilizing a statistical model provides a more objective procedure for retrospective exposure assessment. The development of a weighted multiple regression model is presented for estimation of exposure levels to ethylene oxide (ETO) for inclusion in a cohort mortality study of workers in the sterilization industry. Three steps in development of the model are described: (1) data acquisition and assessment, (2) model building, and (3) evaluation of the model. The final model explained a remarkable 85% of the variability in 205 average measurements of ETO levels. Exposure factors included in the model were exposure category, product type, size of the sterilization unit, selected engineering controls, days after sterilization, and calendar year. The model was evaluated in two ways: against a set of measurement data not used to develop the model and a panel of 11 industrial hygienists representing the sterilization industry. The model predicted ETO exposures within 1.1 ppm of the validation data set with a standard deviation of 3.7 ppm. The arithmetic and geometric means of the 46 measurements in the validation data set were 4.6 and 2.2 ppm, respectively. The model also outperformed the panel of industrial hygienists relative to the validation data in terms of both bias and precision.

Ethylene Oxide↗

A hierarchical statistical model for estimating population properties of quantitative genes.

BACKGROUND: Earlier methods for detecting major genes responsible for a quantitative trait rely critically upon a well-structured pedigree in which the segregation pattern of genes exactly follow Mendelian inheritance laws. However, for many outcrossing species, such pedigrees are not available and genes also display population properties. RESULTS: In this paper, a hierarchical statistical model is proposed to monitor the existence of a major gene based on its segregation and transmission across two successive generations. The model is implemented with an EM algorithm to provide maximum likelihood estimates for genetic parameters of the major locus. This new method is successfully applied to identify an additive gene having a large effect on stem height growth of aspen trees. The estimates of population genetic parameters for this major gene can be generalized to the original breeding population from which the parents were sampled. A simulation study is presented to evaluate finite sample properties of the model. CONCLUSIONS: A hierarchical model was derived for detecting major genes affecting a quantitative trait based on progeny tests of outcrossing species. The new model takes into account the population genetic properties of genes and is expected to enhance the accuracy, precision and power of gene detection.

Algorithms↗

A statistical model that takes into account patient heterogeneity in decision making.

Statistical evaluation of clinical treatments or preventive medicine has profoundly contributed to decision making in medical fields such as with the acceptance of new treatment methods and health promotion policies. It is crucial in such decision making to find a correct statistical model to treat a surprisingly large variety of patients, or a heterogeneous group of patients, even with the same diagnosis. In diseases such as cancer, cardiovascular disease or diabetes, patients are often followed up to certain endpoints and these data are frequently analyzed by logrank tests or Cox-models to evaluate treatment effects. Although these methods have been widely accepted and extensively studied, we are sometimes faced with problems in applying these methods when the heterogeneity of patients is large and a lot of prognostic factors affecting the endpoints have to be considered. Based on the results of the analyses of survival data from more than 6,000 gastric cancer patients, it is revealed that the stratified logrank test may suffer serious power loss, even though primary prognostic factors are used as stratified factors. A so-called 'piecewise linear Cox regression method' for properly treating the heterogeneity of patients is introduced and extensively studied. This method is shown to be appropriate for patient groups with a high degree of heterogeneity such as the gastric cancer patients. The same method is, in principle, applicable to patients of other diseases, too, using statistical software such as SAS, BMDP and etc.

Algorithms↗

Statistical models of shape for the analysis of protein spots in two-dimensional electrophoresis gel images.

In image analysis of two-dimensional electrophoresis gels, individual spots need to be identified and quantified. Two classes of algorithms are commonly applied to this task. Parametric methods rely on a model, making strong assumptions about spot appearance, but are often insufficiently flexible to adequately represent all spots that may be present in a gel. Nonparametric methods make no assumptions about spot appearance and consequently impose few constraints on spot detection, allowing more flexibility but reducing robustness when image data is complex. We describe a parametric representation of spot shape that is both general enough to represent unusual spots, and specific enough to introduce constraints on the interpretation of complex images. Our method uses a model of shape based on the statistics of an annotated training set. The model allows new spot shapes, belonging to the same statistical distribution as the training set, to be generated. To represent spot appearance we use the statistically derived shape convolved with a Gaussian kernel, simulating the diffusion process in spot formation. We show that the statistical model of spot appearance and shape is able to fit to image data more closely than the commonly used spot parameterizations based solely on Gaussian and diffusion models. We show that improvements in model fitting are gained without degrading the specificity of the representation.

Computer Simulation↗

A statistical modelling approach to community prevalence data.

Sample survey techniques are often used to assess the prevalence of illness in a community and to determine any variation with environmental, social and demographic factors. Analysis of survey data is often carried out using several elementary statistical procedures. The formulation of a statistical model is an effective way of conducting a unified analysis. The model provides a concise description of the study population and is a most effective way of summarizing community prevalence data. The testing of statistical hypotheses is equivalent to model simplification and is conveniently performed using general procedures. Two alternative statistical models are given for an investigation into the extent of minor psychiatric morbidity in Perth, Western Australia.

Adolescent↗

Epidemiology of seasonal influenza: use of surveillance data and statistical models to estimate the burden of disease.

The US Centers for Disease Control and Prevention (CDC) uses a 7-component national surveillance system for influenza that includes virologic, influenza-like illness, hospitalization, and mortality data. In addition, some states and health organizations collect additional influenza surveillance data that complement the CDC's surveillance system. Current surveillance data from these programs, together with national hospitalization and mortality data, have been used in statistical models to estimate the annual burden of disease associated with influenza in the United States for many years. National influenza surveillance data also have been used in suitable models to estimate the possible impact of future pandemics. As part of the public health response to the 2003-2004 influenza season, which was noteworthy for its severe effect among children, new US surveillance activities were undertaken. Further improvements in national influenza surveillance systems will be needed to collect and analyze data in a timely manner during the next pandemic.

Adolescent↗

Two-dimensional statistical model for regularized backprojection in SPECT.

In SPECT, both the noise affecting the data and the discretization of the inverse Radon transform are responsible for the ill-posed nature of the reconstruction. To constrain the problem, we propose a regularized backprojection method (RBP) which takes advantage of the relationships existing between the continuity properties of the projections and those of the reconstructed object. The RBP method involves two stages: first, a statistical model (the fixed-effect model) is used to estimate the noise-free part of the projections. Then, the filtered projections are reconstructed using a backprojection algorithm (spline filtered backprojection) which ensures that the reconstructed object belongs to a space consistent with that containing the projections. The method is illustrated using analytical simulations, and the RBP approach is compared to the conventional filtered backprojection. The effect on the reconstructed slices of the parameters involved in RBP is studied in terms of spatial resolution, homogeneity in uniform regions and quantification. It is shown that appropriate combinations of these parameters yield a better compromise between homogeneity and spatial resolution than conventional FBP, with similar quantification performances.

Algorithms↗