PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 379 records · Page 21Linked to original sources

Spatial statistical modeling of disease outbreaks with particular reference to the UK foot and mouth disease (FMD) epidemic of 2001.

In this paper we examine issues relating to the analysis of spatially-referenced disease data. Initially, we discuss the use of exploratory statistical tools such as density estimation and nonparametric regression. We then consider the need for descriptive epidemic models in space, time, and space-time models for epidemic dynamics. Implicitly space-time must be considered in any analysis of the spatial structure of epidemics. The use of Bayesian models for disease spread is discussed and applied to the recent foot and mouth outbreak in the UK.

Animals↗

Evaluation of nonlinear regression approaches to estimation of insulin sensitivity by the minimal model with reference to Bayesian hierarchical analysis.

Minimal model analysis of intravenous glucose tolerance test (IVGTT) glucose and insulin concentrations offers a validated approach to measuring insulin sensitivity, but model identification is not always successful. Improvements may be achieved by using alternative settings in the modeling process, although results may differ according to setting, and care must be exercised in combining results. IVGTT data (12 samples, regular test) from 533 men without diabetes was modeled by the traditional nonlinear regression (NLR) approach, using five different permutations of settings. Results were evaluated with reference to the more robust Bayesian hierarchical (BH) approach to model identification and to the proportion of variance they explained in known correlates of insulin sensitivity (age, BMI, blood pressure, fasting glucose and insulin, serum triglyceride, HDL cholesterol, and uric acid concentration). BH analysis was successful in all cases. With NLR analysis, between 17 and 35 IVGTTs were associated with parameter coefficients of variation (PCVs) for minimal model parameters S(I) (insulin sensitivity) and S(G) (glucose effectiveness) of >100%. Systematic use of each different approach in combination reduced this number to five. Mean (interquartile range) S(I)(NLR) was then 3.14 (2.29-4.63) min(-1).mU(-1).l x 10(-4) and 2.56 (1.74-3.83) min(-1).mU(-1).l x 10(-4) for S(I)(BH) (correlation 0.86, P < 0.0001). S(I)(NLR) explained, on average, 10.6% of the variance in known correlates of insulin sensitivity, whereas S(I)(BH) explained 8.5%. In a large body of data, which BH analysis demonstrated could be fully identified, use of alternative modeling settings in NLR analysis could substantially reduce the number of analyses with PCVs >100%. S(I)(NLR) compared favorably with S(I)(BH) in the proportion of variance explained in known correlates of insulin sensitivity.

Bayes Theorem↗

Inferring pathways and networks with a Bayesian framework.

Numerous mathematical methods have been adapted and developed to quantitatively reverse engineer biological networks, for example, signal transduction pathways, from experimental micro-array data. Compared with stochastic methods, such as Boolean networks, and deterministic methods, such as thermodynamic or differential equation-based models, Bayesian network analysis has the ability to assess, with scoring metrics, causal relations based on conditional probabilities and thus permit hypothesis testing. The goal of this paper is to illustrate the integration of several Bayesian based techniques into a unified Bayesian framework that can infer hepatocellular networks from metabolic data. Reverse engineering of pathways and networks provides a framework for predictive modeling and hypotheses testing to gain deeper insight into living organisms, disease mechanisms, and targeted therapeutics. Evaluating this methodology initially against the known biochemical network provides confidence in the networks that are uncovered from the experimental data using this framework. From the metabolic data we inferred the known sub-networks, such as the tricarboxylic acid (TCA) and urea cycles. In addition, we combined the relationships learned from the data and our current knowledge of the biological system to postulate several alternative metabolic sub-network models that can predict a particular cellular function, such as intracellular triglyceride accumulation.

Algorithms↗

Predictive Bayesian neural network models of MHC class II peptide binding.

We used Bayesian regularized neural networks to model data on the MHC class II-binding affinity of peptides. Training data consisted of sequences and binding data for nonamer (nine amino acid) peptides. Independent test data consisted of sequences and binding data for peptides of length </=25. We assumed that MHC class II-binding activity of peptides depends only on the highest ranked embedded nonamer and that reverse sequences of active nonamers are inactive. We also internally validated the models by using 30% of the training data in an internal test set. We obtained robust models, with near identical statistics for multiple training runs. We determined how predictive our models were using statistical tests and area under the Receiver Operating Characteristic (ROC) graphs (A(ROC)). Most models gave training A(ROC) values close to 1.0 and test set A(ROC) values >0.8. We also used both amino acid indicator variables (bin20) and property-based descriptors to generate models for MHC class II-binding of peptides. The property-based descriptors were more parsimonious than the indicator variable descriptors, making them applicable to larger peptides, and their design makes them able to generalize to unknown peptides outside of the training space. None of the external test data sets contained any of the nonamer sequences in the training sets. Consequently, the models attempted to predict the activity of truly unknown peptides not encountered in the training sets. Our models were well able to tackle the difficult problem of correctly predicting the MHC class II-binding activities of a majority of the test set peptides. Exceptions to the assumption that nonamer motif activities were invariant to the peptide in which they were embedded, together with the limited coverage of the test data, and the fuzziness of the classification procedure, are likely explanations for some misclassifications.

Amino Acid Sequence↗

Bayesian nonparametric inference on the dose level with specified response rate.

The richness of nonparametric Bayesian models has attracted many different applications. Its application in dose-finding studies has been hindered due to lack of methodologies on the nonparametric Bayesian inference on percentiles. The primary interest in dose-finding studies focuses inference on the unknown toxicity or efficacy dose level corresponding to a prespecified rate. This paper shows how this problem may generally be handled by deriving inference on percentiles of a distribution following a Dirichlet process prior. In particular, theoretical results are derived to obtain the nonparametric Bayesian inference of the unknown dose level. This is followed by a description of the numerical implementation of that theory. The method also allows efficient estimation of the entire potency curve. Finally, the usefulness of the approach is demonstrated via an experimental data example.

Animals↗

PGMC: a framework for probabilistic graphic model combination.

Decision making in biomedicine often involves incorporating new evidences into existing or working models reflecting the decision problems at hand. We propose a new framework that facilitates effective and incremental integration of multiple probabilistic graphical models. The proposed framework aims to minimize time and effort required to customize and extend the original models through preserving the conditional independence relationships inherent in two types of probabilistic graphical models: Bayesian networks and influence diagrams. We present a four-step algorithm to systematically combine the qualitative and the quantitative parts of the different models; we also describe three heuristic methods for target variable generation to reduce the complexity of the integrated models. Preliminary results from a case study in heart disease diagnosis demonstrate the feasibility and potential for applying the proposed framework in real applications.

Algorithms↗

Comparing the performance of two indices for spatial model selection: application to two mortality data.

The statistical analysis of spatially correlated data has become an important scientific research topic lately. The analysis of the mortality or morbidity rates observed at different areas may help to decide if people living in certain locations are considered at higher risk than others. Once the statistical model for the data of interest has been chosen, further effort can be devoted to identifying the areas under higher risks. Many scientists, including statisticians, have tried the conditional autoregressive (CAR) model to describe the spatial autocorrelation among the observed data. This model has greater smoothing effect than the exchangeable models, such as the Poisson gamma model for spatial data. This paper focuses on comparing the two types of models using the index LG, the ratio of local to global variability. Two applications, Taiwan asthma mortality and Scotland lip cancer, are considered and the use of LG is illustrated. The estimated values for both data sets are small, implying a Poisson gamma model may be favoured over the CAR model. We discuss the implications for the two applications respectively. To evaluate the performance of the index LG, we also compute the Bayes factor, a Bayesian model selection criterion, to see which model is preferred for the two applications and simulation data. To derive the value of LG, we estimate its posterior mode based on samples derived from the BUGS program, while for Bayes factor we use the double Laplace-Metropolis method, Schwarz criterion, and a modified harmonic mean for approximations. The results of LG and Bayes factor are consistent. We conclude that LG is fairly accurate as an index for selection between Poisson gamma and CAR model. When easy and fast computation is of concern, we recommend using LG as the first and less costly index.

Asthma↗

The development and application of a multilevel decision analysis model for the remediation of contaminated groundwater under uncertainty.

A study was initiated which combined elements of stochastic hydrology, risk assessment, simulation modeling, cost analysis and decision making to define the optimum remediation choice(s) for a Superfund site in the southern United States. The effort focused upon the premise that groundwater remediation is inherently complex due to uncertainties in the geological matrix as well as in contaminant concentrations at points of compliance and/or exposure. The technical analyst should supply the decision maker with estimates of these uncertainties as well as the cost penalties required to reduce them to manageable levels. Monte Carlo transport modeling was employed to define the probability of contaminant excursions from the site, while geostatistical simulation identified a joint plume configuration and its attendant probability. Bayesian modeling was used to define the worth of additional data. These individual components were combined within a Decision Model to identify optimum remediation configurations for a given levels of risk tolerance which could be supplied by the decision maker or affected community. Sensitivity analyses were conducted to define ranges over which the decision would not be affected by variation in the respective decision parameter.

Decision Making↗

Time squared: repeated measures on phylogenies.

Studies of gene expression profiles in response to external perturbation generate repeated measures data that generally follow nonlinear curves. To explore the evolution of such profiles across a gene family, we introduce phylogenetic repeated measures (PR) models. These models draw strength from 2 forms of correlation in the data. Through gene duplication, the family's evolutionary relatedness induces the first form. The second is the correlation across time points within taxonic units, individual genes in this example. We borrow a Brownian diffusion process along a given phylogenetic tree to account for the relatedness and co-opt a repeated measures framework to model the latter. Through simulation studies, we demonstrate that repeated measures models outperform the previously available approaches that consider the longitudinal observations or their differences as independent and identically distributed by using deviance information criteria as Bayesian model selection tools; PR models that borrow phylogenetic information also perform better than nonphylogenetic repeated measures models when appropriate. We then analyze the evolution of gene expression in the yeast kinase family using splines to estimate nonlinear behavior across 3 perturbation experiments. Again, the PR models outperform previous approaches and afford the prediction of ancestral expression profiles. To demonstrate PR model applicability more generally, we conclude with a short examination of variation in brain development across 4 primate species.

Algorithms↗

Bootstrap investigation of the stability of disease mapping of Bayesian cancer relative risk estimations.

BACKGROUND: Bayesian approaches to disease mapping of relative risks are useful for rare disease when geographical units have very different population sizes. As Bayesian approaches may induce very different estimations, it is useful to consider the stability of the estimations as a criterion for evaluating the quality of the results. MATERIAL: Cancer incidence data, from the Isere cancer registry (France) over the 1985-1994 period, have been used to check the proposed method: the study is based on 22 cancer sites among males and 24 among females. METHOD: A bootstrap approach has been retained to evaluate the stability of the estimations. The coefficient of variation was chosen as an indicator of stability. Three Bayesian models corresponding to global, local and combined smoothing techniques, have been considered. The stability analysis has taken account of the results of spatial autocorrelation and heterogeneity tests. RESULTS: Bayesian approaches do not necessarily lead to stable estimations. The local smoothing approach induces estimations that are often unstable. The global smoothing approach is the most stable, but is conservative. Combined smoothing appears to be a good compromise if significant spatial variations and heterogeneity of relative risks exist. CONCLUSION: Bayesian estimations of relative risks may be very unstable. However, when results of spatial autocorrelation and heterogeneity tests are taken into account to choose between the different Bayesian approaches, instability becomes negligible.

Algorithms↗

Development and Bayesian evaluation of an ELISA to detect specific antibodies to Sarcoptes scabiei var suis in the meat juice of pigs.

Samples of ear scrapings, serum and diaphragmatic muscle were collected from 271 fattening pigs at the slaughterhouse. The scrapings were examined for the presence of mites, and tests for specific antibodies to Sarcoptes scabiei var suis in the serum and meat juice were made with an experimental ELISA. The cut-off value for the meat-juice ELISA was estimated at an optical density of 0.5 by receiver operating characteristic curve analysis, on the basis of the cut-off value for the serum ELISA of 0.4. The results of the three tests were used in a Bayesian model to estimate the characteristics of each test. The specificity of the tests of the ear scrapings was considered to be 1 and their sensitivity was estimated by Bayesian analysis to be 0.86, with a 95 per cent confidence interval (CI) of 0.73 to 0.99. The sensitivity of the meat juice ELISA (0.71, 95 per cent CI 0.6 to 0.8) and its specificity (0.77, 95 per cent CI 0.66 to 0.89) were comparable with the sensitivity (0.73, 95 per cent CI 0.6 to 0.8) and specificity (0.81, 95 per cent CI 0.69 to 0.95) of the serum ELISA.

Abattoirs↗

Bayesian detection and modeling of spatial disease clustering.

Many current statistical methods for disease clustering studies are based on a hypothesis testing paradigm. These methods typically do not produce useful estimates of disease rates or cluster risks. In this paper, we develop a Bayesian procedure for drawing inferences about specific models for spatial clustering. The proposed methodology incorporates ideas from image analysis, from Bayesian model averaging, and from model selection. With our approach, we obtain estimates for disease rates and allow for greater flexibility in both the type of clusters and the number of clusters that may be considered. We illustrate the proposed procedure through simulation studies and an analysis of the well-known New York leukemia data.

Bayes Theorem↗

Disease mapping models: an empirical evaluation. Disease Mapping Collaborative Group.

The analysis of small area disease incidence has now developed to a degree where many methods have been proposed. However, there are few studies of the relative merits of the methods available. While many Bayesian models have been examined with respect to prior sensitivity, it is clear that wider comparisons of methods are largely missing from the literature. In this paper we present some preliminary results concerning the goodness-of-fit of a variety of disease mapping methods to simulated data for disease incidence derived from a range of models. These simulated models cover simple risk gradients to more complex true risk structures, including spatial correlation. The main general results presented here show that the gamma-Poisson exchangeable model and the Besag, York and Mollie (BYM) model are most robust across a range of diverse models. Mixture models are less robust. Non-parametric smoothing methods perform badly in general. Linear Bayes methods display behaviour similar to that of the gamma-Poisson methods.

Algorithms↗

The analysis of survival data with a non-susceptible fraction and dual censoring mechanisms.

It is known that the ages of onset of many diseases are determined by both a genetic predisposition to disease as well as environmental risk factors that are capable of either triggering or hastening the onset of disease. Difficulties in modelling onset ages arise when a large fraction fail to inherit the disease-causing gene, and multiple reasons for censoring result in unobserved onset ages. We present a parametric Bayesian model that includes subjects with missing age information, non-susceptible subjects and allows for regression on risk factor information. The model is fit using Markov chain Monte Carlo simulation from the posterior distribution, and allows the simultaneous estimation of the proportion of the population at risk of disease, the mean onset age of disease, survival after disease onset, and the association of risk factors with susceptibility, onset age and survival after onset. An example employing Huntington's disease data is presented.

Age of Onset↗

Estimating allelic number and identity in state of QTLs in interconnected families.

When multiple related families derived from inbred lines are jointly analysed to detect quantitative trait loci (QTLs), the analysis should estimate allelic effects as accurately as possible and estimate the probability that different parents carry alleles that are identical in state. Analyses exist that assume that all parents carry unique alleles or that all parents but one carry the same allele. In practice, many configurations are possible that group different parents according to their identity-in-state condition at a putative QTL allele. Here, we propose a variable model Bayesian analysis that selects among possible identity-in-state configurations and jointly estimates the allelic effects of identical-in-state parents. We contrast this analysis with a fixed model analysis that estimates unique allelic effects for all parents. We analyse two simulated mating designs: an experimental design in which three inbred parents were crossed to generate two families of 150 doubled haploid lines; and a breeding design in which 20 inbred parents were crossed to generate 60 families of 20 doubled haploid lines, with each parent contributing to six families. In all cases where some parents were simulated to carry alleles of identical effect (that is, they were identical in state), the variable analysis estimated allelic effects with lower mean-squared error than the fixed analysis. The variable analysis showed that, unless each family contains many individuals (more than 100), there is insufficient information in DNA-marker and phenotypic data to determine with high probability the QTL allelic number.

Bayes Theorem↗

Computational modeling of the Plasmodium falciparum interactome reveals protein function on a genome-wide scale.

Many thousands of proteins encoded by the genome of Plasmodium falciparum, the causal organism of the deadliest form of human malaria, are of unknown function. It is of utmost importance that these proteins be characterized if we are to develop combative strategies against malaria based on the biology of the parasite. In an attempt to infer protein function on a genome-wide scale, we computationally modeled the P. falciparum interactome, elucidating local and global functional relationships between gene products. The resulting interaction network, reconstructed by integrating in silico and experimental functional genomics data within a Bayesian framework, covers approximately 68% of the parasite genome and provides functional inferences for more than 2000 uncharacterized proteins, based on their associations. Network reconstruction involved the use of a novel strategy, where we incorporated continuously updated, uniform reference priors in our Bayesian model. This method for generating interaction maps is thus also well suited for application to other genomes, where pre-existing interactome knowledge is sparse. Additionally, we superimposed this map on genomes of three apicomplexan pathogens--Plasmodium yoelii, Toxoplasma gondii, and Cryptosporidium parvum--describing relationships between these organisms based on retained functional linkages. This comparison provided a glimpse of the highly evolved nature of P. falciparum; for instance, a deficit of nearly 26% in terms of predicted interactions is observed against P. yoelii, because of missing ortholog partners in pairs of functionally linked proteins.

Animals↗

How the probability of a false positive affects the value of DNA evidence.

Errors in sample handling or test interpretation may cause false positives in forensic DNA testing. This article uses a Bayesian model to show how the potential for a false positive affects the evidentiary value of DNA evidence and the sufficiency of DNA evidence to meet traditional legal standards for conviction. The Bayesian analysis is contrasted with the "false positive fallacy," an intuitively appealing but erroneous alternative interpretation. The findings show the importance of having accurate information about both the random match probability and the false positive probability when evaluating DNA evidence. It is argued that ignoring or underestimating the potential for a false positive can lead to serious errors of interpretation, particularly when the suspect is identified through a "DNA dragnet" or database search, and that ignorance of the true rate of error creates an important element of uncertainty about the value of DNA evidence.

Bayes Theorem↗

A statistical model for locating regulatory regions in genomic DNA.

In addition to genes, chromosomal DNA contains sequences that serve as signals for turning on and off gene expression. These signals are thought to be distributed as clusters in the regulatory regions of genes. We develop a Bayesian model that views locating regulatory regions in genomic DNA as a change-point problem, with the beginning of regulatory and non-regulatory regions corresponding to the change points. The model is based on a hidden Markov chain. The data consist of nucleotide positions of protein-binding elements in a genomic DNA sequence. These positions are identified using a reference catalogue containing elements that interact with transcription factors implicated in controlling the expression of protein-encoding genes. Among the protein-binding elements in a genomic DNA sequence, the statistical model automatically selects those that tend to predict regulatory regions. We test the model using viral sequences that include known regulatory regions and provide the results obtained for human genomic DNA corresponding to the beta globin locus on chromosome 11.

Adenoviridae↗