PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18Linked to original sources

Exorcising protocol-induced spirits: making the clinical trial relevant for economics.

Economic evaluation frequently depends on estimates from clinical trials of both effectiveness of treatment and resource utilization accompanying it. Protocol-driven events in the trial among other influences often imply that both estimates will be inaccurate. This paper indicates how one may supplement a trial with additional data to connect the artificial trial to the real world of clinical practice. It also shows that data required for this model may be estimated from other sources (via Bayesian modeling) if they are not directly available. The required data include (for example) the proportion of patients with disease who would have presented with clinical signs if they had not been part of a trial that allowed early detection and treatment based on subclinical testing mandated by trial protocol. Those presenting with clinical signs would use additional resources for treatment and/or confirmatory diagnostics. Those with subclinical disease either 1) would never use resources in the case where they never developed clinical manifestations or 2) would use resources of a different type or intensity and perhaps have different outcomes by virtue of their disease's being discovered at a later point in time. Basing resource use and ultimate effectiveness on this revised measure of outcome rather than the one in the trial should lead to more accurate predictions for economic purposes.

Anticoagulants↗

A spatial statistical model for landscape genetics.

Landscape genetics is a new discipline that aims to provide information on how landscape and environmental features influence population genetic structure. The first key step of landscape genetics is the spatial detection and location of genetic discontinuities between populations. However, efficient methods for achieving this task are lacking. In this article, we first clarify what is conceptually involved in the spatial modeling of genetic data. Then we describe a Bayesian model implemented in a Markov chain Monte Carlo scheme that allows inference of the location of such genetic discontinuities from individual geo-referenced multilocus genotypes, without a priori knowledge on populational units and limits. In this method, the global set of sampled individuals is modeled as a spatial mixture of panmictic populations, and the spatial organization of populations is modeled through the colored Voronoi tessellation. In addition to spatially locating genetic discontinuities, the method quantifies the amount of spatial dependence in the data set, estimates the number of populations in the studied area, assigns individuals to their population of origin, and detects individual migrants between populations, while taking into account uncertainty on the location of sampled individuals. The performance of the method is evaluated through the analysis of simulated data sets. Results show good performances for standard data sets (e.g., 100 individuals genotyped at 10 loci with 10 alleles per locus), with high but also low levels of population differentiation (e.g., FST<0.05). The method is then applied to a set of 88 individuals of wolverines (Gulo gulo) sampled in the northwestern United States and genotyped at 10 microsatellites.

Animals↗

Non-parametric maximum likelihood estimators for disease mapping.

A Non-Parametric Maximum Likelihood approach to the estimation of relative risks in the context of disease mapping is discussed and a NPML approximation to conditional autoregressive models is proposed. NPML estimates have been compared to other proposed solutions (Maximum Likelihood via Monte Carlo Scoring, Hierarchical Bayesian models) using real examples. Overall, the NPML autoregressive estimates (with weighted term) were closer to the Bayesian estimates. The exchangeable NPML model ranked immediately after, even if it implied a greater shrinkage, while the truncated auto-Poisson showed inadequate for disease mapping. The coefficients of the autoregressive term for the different mixtures have clear interpretations: in the breast cancer example, the larger cities in the region showed high rates and very low correlation with the neighbouring areas, while the less populated rural areas with low rates were strongly positively correlated each other. This pattern is expected since breast cancer is strongly correlated with parity and age at first birth, and the female population of the rural areas experienced a decline in fertility much later than those living in the larger cities. The leukemia example highlighted the failure of the Poisson-Gamma model and other general overdispersion tests to detect high risk areas under specific conditions. The NPML approach in Aitkin is very general, simple and flexible. However the user should be warned against the possibility of local maxima and the difficulty in detecting the optimal number of components. Special software (such as CAMAN or DismapWin) had been developed and should be recommended mainly to not experienced users.

Algorithms↗

A novel approach for clustering proteomics data using Bayesian fast Fourier transform.

MOTIVATION: Bioinformatics clustering tools are useful at all levels of proteomic data analysis. Proteomics studies can provide a wealth of information and rapidly generate large quantities of data from the analysis of biological specimens. The high dimensionality of data generated from these studies requires the development of improved bioinformatics tools for efficient and accurate data analyses. For proteome profiling of a particular system or organism, a number of specialized software tools are needed. Indeed, significant advances in the informatics and software tools necessary to support the analysis and management of these massive amounts of data are needed. Clustering algorithms based on probabilistic and Bayesian models provide an alternative to heuristic algorithms. The number of clusters (diseased and non-diseased groups) is reduced to the choice of the number of components of a mixture of underlying probability. The Bayesian approach is a tool for including information from the data to the analysis. It offers an estimation of the uncertainties of the data and the parameters involved. RESULTS: We present novel algorithms that can organize, cluster and derive meaningful patterns of expression from large-scaled proteomics experiments. We processed raw data using a graphical-based algorithm by transforming it from a real space data-expression to a complex space data-expression using discrete Fourier transformation; then we used a thresholding approach to denoise and reduce the length of each spectrum. Bayesian clustering was applied to the reconstructed data. In comparison with several other algorithms used in this study including K-means, (Kohonen self-organizing map (SOM), and linear discriminant analysis, the Bayesian-Fourier model-based approach displayed superior performances consistently, in selecting the correct model and the number of clusters, thus providing a novel approach for accurate diagnosis of the disease. Using this approach, we were able to successfully denoise proteomic spectra and reach up to a 99% total reduction of the number of peaks compared to the original data. In addition, the Bayesian-based approach generated a better classification rate in comparison with other classification algorithms. This new finding will allow us to apply the Fourier transformation for the selection of the protein profile for each sample, and to develop a novel bioinformatic strategy based on Bayesian clustering for biomarker discovery and optimal diagnosis.

Algorithms↗

Site-specific updating and aggregation of Bayesian belief network models for multiple experts.

A method for combining multiple expert opinions that are encoded in a Bayesian Belief Network (BBN) model is presented and applied to a problem involving the cleanup of hazardous chemicals at a site with contaminated groundwater. The method uses Bayes Rule to update each expert model with the observed evidence, then uses it again to compute posterior probability weights for each model. The weights reflect the consistency of each model with the observed evidence, allowing the aggregate model to be tailored to the particular conditions observed in the site-specific application of the risk model. The Bayesian update is easy to implement, since the likelihood for the set of evidence (observations for selected nodes of the BBN model) is readily computed by sequential execution of the BBN model. The method is demonstrated using a simple pedagogical example and subsequently applied to a groundwater contamination problem using an expert-knowledge BBN model. The BBN model in this application predicts the probability that reductive dechlorination of the contaminant trichlorethene (TCE) is occurring at a site--a critical step in the demonstration of the feasibility of monitored natural attenuation for site cleanup--given information on 14 measurable antecedent and descendant conditions. The predictions for the BBN models for 21 experts are weighted and aggregated using examples of hypothetical and actual site data. The method allows more weight for those expert models that are more reflective of the site conditions, and is shown to yield an aggregate prediction that differs from that of simple model averaging in a potentially significant manner.

Algorithms↗

Bootstrap model averaging in time series studies of particulate matter air pollution and mortality.

The consensus from time series studies that have investigated the mortality effects of particulate matter air pollution (PM) is that increases in PM are associated with increases in daily mortality. However, recently concerns have been raised that the observed positive association between PM and mortality may be an artefact of model selection due to multiple hypothesis testing. This problem arises when a number of models are investigated, but only the "best" model is reported and all subsequent inference is based on this model, ignoring the model selection process. In this paper, we introduce the use of the bootstrap as a means of addressing the problems of model selection in PM mortality time series studies. Using the bootstrap to perform inference about the effect of PM on mortality is a process based on a set of models rather than on a single model. It is shown that using the bootstrap to overcome the problems of model selection is competitive with the existing methodology of Bayesian model averaging.

Air Pollutants↗

Identifying biological concepts from a protein-related corpus with a probabilistic topic model.

BACKGROUND: Biomedical literature, e.g., MEDLINE, contains a wealth of knowledge regarding functions of proteins. Major recurring biological concepts within such text corpora represent the domains of this body of knowledge. The goal of this research is to identify the major biological topics/concepts from a corpus of protein-related MEDLINE titles and abstracts by applying a probabilistic topic model. RESULTS: The latent Dirichlet allocation (LDA) model was applied to the corpus. Based on the Bayesian model selection, 300 major topics were extracted from the corpus. The majority of identified topics/concepts was found to be semantically coherent and most represented biological objects or concepts. The identified topics/concepts were further mapped to the controlled vocabulary of the Gene Ontology (GO) terms based on mutual information. CONCLUSION: The major and recurring biological concepts within a collection of MEDLINE documents can be extracted by the LDA model. The identified topics/concepts provide parsimonious and semantically-enriched representation of the texts in a semantic space with reduced dimensionality and can be used to index text.

Abstracting and Indexing↗

[Probabilistic approach to age estimation of children by dental maturation].

Two probabilist methods of age prediction in children are proposed: they are both based on the radiological presence of erupted teeth or germs. Using an apprenticeship sample of known age and sex, we established several discriminant models (+/- 13, +/- 16, +/- 18 years old). We also evaluated a Bayesian model with the following age groups: < 13, [13-16[, [16-18[, > or = 18 years old, or [X and Y] years old. When applied on a known test sample, Fisher's linear functions presented a success rate greater than 90%, above 13 years threshold, and below 16 and 18 years thresholds, and Bayesian approach, greater than 85%. Therefore, these methods provide an interesting alternative for children age determination that can be applied in biological and forensic anthropology, too.

Adolescent↗

The importance of proper model assumption in bayesian phylogenetics.

We studied the importance of proper model assumption in the context of Bayesian phylogenetics by examining >5,000 Bayesian analyses and six nested models of nucleotide substitution. Model misspecification can strongly bias bipartition posterior probability estimates. These biases were most pronounced when rate heterogeneity was ignored. The type of bias seen at a particular bipartition appeared to be strongly influenced by the lengths of the branches surrounding that bipartition. In the Felsenstein zone, posterior probability estimates of bipartitions were biased when the assumed model was underparameterized but were unbiased when the assumed model was overparameterized. For the inverse Felsenstein zone, however, both underparameterization and overparameterization led to biased bipartition posterior probabilities, although the bias caused by overparameterization was less pronounced and disappeared with increased sequence length. Model parameter estimates were also affected by model misspecification. Underparameterization caused a bias in some parameter estimates, such as branch lengths and the gamma shape parameter, whereas overparameterization caused a decrease in the precision of some parameter estimates. We caution researchers to assure that the most appropriate model is assumed by employing both a priori model choice methods and a posteriori model adequacy tests.

Bayes Theorem↗

Differential and trajectory methods for time course gene expression data.

MOTIVATION: The issue of high dimensionality in microarray data has been, and remains, a hot topic in statistical and computational analysis. Efficient gene filtering and differentiation approaches can reduce the dimensions of data, help to remove redundant genes and noises, and highlight the most relevant genes that are major players in the development of certain diseases or the effect of drug treatment. The purpose of this study is to investigate the efficiency of parametric (including Bayesian and non-Bayesian, linear and non-linear), non-parametric and semi-parametric gene filtering methods through the application of time course microarray data from multiple sclerosis patients being treated with interferon-beta-1a. The analysis of variance with bootstrapping (parametric), class dispersion (semi-parametric) and Pareto (non-parametric) with permutation methods are presented and compared for filtering and finding differentially expressed genes. The Bayesian linear correlated model, the Bayesian non-linear model the and non-Bayesian mixed effects model with bootstrap were also developed to characterize the differential expression patterns. Furthermore, trajectory-clustering approaches were developed in order to investigate the dynamic patterns and inter-dependency of drug treatment effects on gene expression. RESULTS: Results show that the presented methods performed significant differently but all were adequate in capturing a small number of the potentially relevant genes to the disease. The parametric method, such as the mixed model and two Bayesian approaches proved to be more conservative. This may because these methods are based on overall variation in expression across all time points. The semi-parametric (class dispersion) and non-parametric (Pareto) methods were appropriate in capturing variation in expression from time point to time point, thereby making them more suitable for investigating significant monotonic changes and trajectories of changes in gene expressions in time course microarray data. Also, the non-linear Bayesian model proved to be less conservative than linear Bayesian correlated growth models to filter out the redundant genes, although the linear model showed better fit than non-linear model (smaller DIC). We also report the trajectories of significant genes-since we have been able to isolate trajectories of genes whose regulations appear to be inter-dependent.

Computer Simulation↗

Intensity-based hierarchical Bayes method improves testing for differentially expressed genes in microarray experiments.

BACKGROUND: The small sample sizes often used for microarray experiments result in poor estimates of variance if each gene is considered independently. Yet accurately estimating variability of gene expression measurements in microarray experiments is essential for correctly identifying differentially expressed genes. Several recently developed methods for testing differential expression of genes utilize hierarchical Bayesian models to "pool" information from multiple genes. We have developed a statistical testing procedure that further improves upon current methods by incorporating the well-documented relationship between the absolute gene expression level and the variance of gene expression measurements into the general empirical Bayes framework. RESULTS: We present a novel Bayesian moderated-T, which we show to perform favorably in simulations, with two real, dual-channel microarray experiments and in two controlled single-channel experiments. In simulations, the new method achieved greater power while correctly estimating the true proportion of false positives, and in the analysis of two publicly-available "spike-in" experiments, the new method performed favorably compared to all tested alternatives. We also applied our method to two experimental datasets and discuss the additional biological insights as revealed by our method in contrast to the others. The R-source code for implementing our algorithm is freely available at http://eh3.uc.edu/ibmt. CONCLUSION: We use a Bayesian hierarchical normal model to define a novel Intensity-Based Moderated T-statistic (IBMT). The method is completely data-dependent using empirical Bayes philosophy to estimate hyperparameters, and thus does not require specification of any free parameters. IBMT has the strength of balancing two important factors in the analysis of microarray data: the degree of independence of variances relative to the degree of identity (i.e. t-tests vs. equal variance assumption), and the relationship between variance and signal intensity. When this variance-intensity relationship is weak or does not exist, IBMT reduces to a previously described moderated t-statistic. Furthermore, our method may be directly applied to any array platform and experimental design. Together, these properties show IBMT to be a valuable option in the analysis of virtually any microarray experiment.

Animals↗

A critique of Oaksford, Chater, and Larkin's (2000) conditional probability model of conditional reasoning.

M. Oaksford, N. Chater, and J. Larkin (2000) proffered a Bayesian model in which conditional inferences are a direct function of conditional probabilities. In the current article, the authors first considered this model regarding the processing of negatives in conditional reasoning. Its predictions were evaluated against a large-scale meta-analysis (W. J. Schroyens, W. Schaeken, & G. d'Ydewalle, 2001b). This evaluation shows that the model is flawed: The relative size of the negative effects does not match predictions. Next, the authors evaluated the model in relation to inferences about affirmative conditionals, again considering the results of a meta-analysis (W. J. Schroyens, W. Schaeken, & G. d'Ydewalle, 2001a). The conditional probability model is countered by the data reported in literature; a mental models based model produces a better fit. The authors conclude that a purely probabilistic model is deficient and incomplete and cannot do without algorithmic processing assumptions if it is to advance toward a descriptively adequate psychological theory.

Conditioning, Psychological↗

Small area estimation of incidence of cancer around a known source of exposure with fine resolution data.

OBJECTIVES: To describe the small area system developed in Finland. To illustrate the use of the system with analyses of incidence of lung cancer around an asbestos mine. To compare the performance of different spatial statistical models when applied to sparse data. METHODS: In the small area system, cancer and population data are available by sex, age, and socioeconomic status in adjacent "pixels", squares of size 0.5 km x 0.5 km. The study area was partitioned into sub-areas based on estimated exposure. The original data at the pixel level were used in a spatial random field model. For comparison, standardised incidence ratios were estimated, and full bayesian and empirical bayesian models were fitted to aggregated data. Incidence of lung cancer around a former asbestos mine was used as an illustration. RESULTS: The spatial random field model, which has been used in former small area studies, did not converge with present fine resolution data. The number of neighbouring pixels used in smoothing had to be enlarged, and informative distributions for hyperparameters were used to stabilise the unobserved random field. The ordered spatial random field model gave lower estimates than the Poisson model. When one of the three effects of area were fixed, the model gave similar estimates with a narrower interval than the Poisson model. CONCLUSIONS: The use of fine resolution data and socioeconomic status as a means of controlling for confounding related to lifestyle is useful when estimating risk of cancer around point sources. However, better statistical methods are needed for spatial modelling of fine resolution data.

Adolescent↗

Bayesian inference for model-based segmentation of computed radiographs of the hand.

We present a method for medical image understanding by computer that uses model-based, hierarchical Bayesian inference to accurately segment imaged anatomy. A first application is a prototype system that automatically segments and measures symptoms of arthridities in hand radiographs. This is potentially useful in radiological diagnosis and tracking of arthridities. Key steps of the model-based, Bayesian inference approach are: (1) prediction of imagery features from 3D models of anatomy, parameterized by population statistics, (2) local image feature extraction in predicted sub-regions, and (3) the use of a probabilistic calculus to accrue results of image processing and image feature matching procedures in support or denial of hypotheses about the imaged anatomy. The prototype system for hand radiograph analysis accurately segments normal and somewhat degenerated hand anatomy. Results are shown of the ability of the automated system to 'fail soft', recognizing when segmentation is inadequate for accurate measurement. This self evaluation capability improves reliability of measurements for potential clinical use.

Arthritis↗

Modelling geographically referenced survival data with a cure fraction.

The emergence of geographical information systems and related softwares nowadays enables medical databases to incorporate the geographical information on patients, allowing studies in spatial associations. Public health administrators and researchers are often interested in detecting variation in survival patterns by region or county in order to understand the possible factors that contribute towards such spatial discrepancies. These issues have led statisticians to develop survival models that account for spatial clustering and variation. Additionally, with rapid developments in medical and health sciences, researchers increasingly encounter data sets where a substantial portion of patients are cured. Models accounting for cure in the population assist in the prognosis of potentially terminal diseases. This article proposes a Bayesian modelling framework that models spatial associations for areally referenced survival data using a general class of cure models proposed by Cooner et al. The special models we outline are alternatives to the traditional proportional hazards models and can be fitted using standard Bayesian software such as WinBUGS.

Bayes Theorem↗

Bayesian methods in meta-analysis and evidence synthesis.

This paper reviews the use of Bayesian methods in meta-analysis. Whilst there has been an explosion in the use of meta-analysis over the last few years, driven mainly by the move towards evidence-based healthcare, so too Bayesian methods are being used increasingly within medical statistics. Whilst in many meta-analysis settings the Bayesian models used mirror those previously adopted in a frequentist formulation, there are a number of specific advantages conferred by the Bayesian approach. These include: full allowance for all parameter uncertainty in the model, the ability to include other pertinent information that would otherwise be excluded, and the ability to extend the models to accommodate more complex, but frequently occurring, scenarios. The Bayesian methods discussed are illustrated by means of a meta-analysis examining the evidence relating to electronic fetal heart rate monitoring and perinatal mortality in which evidence is available from a variety of sources.

Bayes Theorem↗

Genetic analysis of the age at menopause by using estimating equations and Bayesian random effects models.

Multi-wave self-report data on age at menopause in 2182 female twin pairs (1355 monozygotic and 827 dizygotic pairs), were analysed to estimate the genetic, common and unique environmental contribution to variation in age at menopause. Two complementary approaches for analysing correlated time-to-onset twin data are considered: the generalized estimating equations (GEE) method in which one can estimate zygosity-specific dependence simultaneously with regression coefficients that describe the average population response to changing covariates; and a subject-specific Bayesian mixed model in which heterogeneity in regression parameters is explicitly modelled and the different components of variation may be estimated directly. The proportional hazards and Weibull models were utilized, as both produce natural frameworks for estimating relative risks while adjusting for simultaneous effects of other covariates. A simple Markov chain Monte Carlo method for covariate imputation of missing data was used and the actual implementation of the Bayesian model was based on Gibbs sampling using the freeware package BUGS.

Adult↗

Flexibility versus generalizability in model selection.

Which quantitative method should be used to choose among competing mathematical models of cognition? Massaro, Cohen, Campbell, and Rodriguez (2001) favor root mean squared deviation (RMSD), choosing the model that provides the best fit to the data. Their simulation results appear to legitimize its use for comparing two models of information integration because it performed just as well as Bayesian model selection (BMS), which had previously been shown by Myung and Pitt (1997) to be a superior alternative selection method because it considers a model's complexity in addition to its fit. In the present study, after contrasting the theoretical approaches to model selection espoused by Massaro et al. and Myung and Pitt, we discuss the cause of the inconsistencies by expanding on the simulations of Massaro et al. Findings demonstrate that the results from model recovery simulations can be misleading if they are not interpreted relative to the data on which they were evaluated, and that BMS is a more robust selection method.

Bayes Theorem↗