PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Differential and trajectory methods for time course gene expression data.

MOTIVATION: The issue of high dimensionality in microarray data has been, and remains, a hot topic in statistical and computational analysis. Efficient gene filtering and differentiation approaches can reduce the dimensions of data, help to remove redundant genes and noises, and highlight the most relevant genes that are major players in the development of certain diseases or the effect of drug treatment. The purpose of this study is to investigate the efficiency of parametric (including Bayesian and non-Bayesian, linear and non-linear), non-parametric and semi-parametric gene filtering methods through the application of time course microarray data from multiple sclerosis patients being treated with interferon-beta-1a. The analysis of variance with bootstrapping (parametric), class dispersion (semi-parametric) and Pareto (non-parametric) with permutation methods are presented and compared for filtering and finding differentially expressed genes. The Bayesian linear correlated model, the Bayesian non-linear model the and non-Bayesian mixed effects model with bootstrap were also developed to characterize the differential expression patterns. Furthermore, trajectory-clustering approaches were developed in order to investigate the dynamic patterns and inter-dependency of drug treatment effects on gene expression. RESULTS: Results show that the presented methods performed significant differently but all were adequate in capturing a small number of the potentially relevant genes to the disease. The parametric method, such as the mixed model and two Bayesian approaches proved to be more conservative. This may because these methods are based on overall variation in expression across all time points. The semi-parametric (class dispersion) and non-parametric (Pareto) methods were appropriate in capturing variation in expression from time point to time point, thereby making them more suitable for investigating significant monotonic changes and trajectories of changes in gene expressions in time course microarray data. Also, the non-linear Bayesian model proved to be less conservative than linear Bayesian correlated growth models to filter out the redundant genes, although the linear model showed better fit than non-linear model (smaller DIC). We also report the trajectories of significant genes-since we have been able to isolate trajectories of genes whose regulations appear to be inter-dependent.

Computer Simulation↗

A critique of Oaksford, Chater, and Larkin's (2000) conditional probability model of conditional reasoning.

M. Oaksford, N. Chater, and J. Larkin (2000) proffered a Bayesian model in which conditional inferences are a direct function of conditional probabilities. In the current article, the authors first considered this model regarding the processing of negatives in conditional reasoning. Its predictions were evaluated against a large-scale meta-analysis (W. J. Schroyens, W. Schaeken, & G. d'Ydewalle, 2001b). This evaluation shows that the model is flawed: The relative size of the negative effects does not match predictions. Next, the authors evaluated the model in relation to inferences about affirmative conditionals, again considering the results of a meta-analysis (W. J. Schroyens, W. Schaeken, & G. d'Ydewalle, 2001a). The conditional probability model is countered by the data reported in literature; a mental models based model produces a better fit. The authors conclude that a purely probabilistic model is deficient and incomplete and cannot do without algorithmic processing assumptions if it is to advance toward a descriptively adequate psychological theory.

Conditioning, Psychological↗

Small area estimation of incidence of cancer around a known source of exposure with fine resolution data.

OBJECTIVES: To describe the small area system developed in Finland. To illustrate the use of the system with analyses of incidence of lung cancer around an asbestos mine. To compare the performance of different spatial statistical models when applied to sparse data. METHODS: In the small area system, cancer and population data are available by sex, age, and socioeconomic status in adjacent "pixels", squares of size 0.5 km x 0.5 km. The study area was partitioned into sub-areas based on estimated exposure. The original data at the pixel level were used in a spatial random field model. For comparison, standardised incidence ratios were estimated, and full bayesian and empirical bayesian models were fitted to aggregated data. Incidence of lung cancer around a former asbestos mine was used as an illustration. RESULTS: The spatial random field model, which has been used in former small area studies, did not converge with present fine resolution data. The number of neighbouring pixels used in smoothing had to be enlarged, and informative distributions for hyperparameters were used to stabilise the unobserved random field. The ordered spatial random field model gave lower estimates than the Poisson model. When one of the three effects of area were fixed, the model gave similar estimates with a narrower interval than the Poisson model. CONCLUSIONS: The use of fine resolution data and socioeconomic status as a means of controlling for confounding related to lifestyle is useful when estimating risk of cancer around point sources. However, better statistical methods are needed for spatial modelling of fine resolution data.

Adolescent↗

Bayesian inference for model-based segmentation of computed radiographs of the hand.

We present a method for medical image understanding by computer that uses model-based, hierarchical Bayesian inference to accurately segment imaged anatomy. A first application is a prototype system that automatically segments and measures symptoms of arthridities in hand radiographs. This is potentially useful in radiological diagnosis and tracking of arthridities. Key steps of the model-based, Bayesian inference approach are: (1) prediction of imagery features from 3D models of anatomy, parameterized by population statistics, (2) local image feature extraction in predicted sub-regions, and (3) the use of a probabilistic calculus to accrue results of image processing and image feature matching procedures in support or denial of hypotheses about the imaged anatomy. The prototype system for hand radiograph analysis accurately segments normal and somewhat degenerated hand anatomy. Results are shown of the ability of the automated system to 'fail soft', recognizing when segmentation is inadequate for accurate measurement. This self evaluation capability improves reliability of measurements for potential clinical use.

Arthritis↗

Bayesian methods in meta-analysis and evidence synthesis.

This paper reviews the use of Bayesian methods in meta-analysis. Whilst there has been an explosion in the use of meta-analysis over the last few years, driven mainly by the move towards evidence-based healthcare, so too Bayesian methods are being used increasingly within medical statistics. Whilst in many meta-analysis settings the Bayesian models used mirror those previously adopted in a frequentist formulation, there are a number of specific advantages conferred by the Bayesian approach. These include: full allowance for all parameter uncertainty in the model, the ability to include other pertinent information that would otherwise be excluded, and the ability to extend the models to accommodate more complex, but frequently occurring, scenarios. The Bayesian methods discussed are illustrated by means of a meta-analysis examining the evidence relating to electronic fetal heart rate monitoring and perinatal mortality in which evidence is available from a variety of sources.

Bayes Theorem↗

Genetic analysis of the age at menopause by using estimating equations and Bayesian random effects models.

Multi-wave self-report data on age at menopause in 2182 female twin pairs (1355 monozygotic and 827 dizygotic pairs), were analysed to estimate the genetic, common and unique environmental contribution to variation in age at menopause. Two complementary approaches for analysing correlated time-to-onset twin data are considered: the generalized estimating equations (GEE) method in which one can estimate zygosity-specific dependence simultaneously with regression coefficients that describe the average population response to changing covariates; and a subject-specific Bayesian mixed model in which heterogeneity in regression parameters is explicitly modelled and the different components of variation may be estimated directly. The proportional hazards and Weibull models were utilized, as both produce natural frameworks for estimating relative risks while adjusting for simultaneous effects of other covariates. A simple Markov chain Monte Carlo method for covariate imputation of missing data was used and the actual implementation of the Bayesian model was based on Gibbs sampling using the freeware package BUGS.

Adult↗

Flexibility versus generalizability in model selection.

Which quantitative method should be used to choose among competing mathematical models of cognition? Massaro, Cohen, Campbell, and Rodriguez (2001) favor root mean squared deviation (RMSD), choosing the model that provides the best fit to the data. Their simulation results appear to legitimize its use for comparing two models of information integration because it performed just as well as Bayesian model selection (BMS), which had previously been shown by Myung and Pitt (1997) to be a superior alternative selection method because it considers a model's complexity in addition to its fit. In the present study, after contrasting the theoretical approaches to model selection espoused by Massaro et al. and Myung and Pitt, we discuss the cause of the inconsistencies by expanding on the simulations of Massaro et al. Findings demonstrate that the results from model recovery simulations can be misleading if they are not interpreted relative to the data on which they were evaluated, and that BMS is a more robust selection method.

Bayes Theorem↗

Identification of transcription factor binding sites with variable-order Bayesian networks.

MOTIVATION: We propose a new class of variable-order Bayesian network (VOBN) models for the identification of transcription factor binding sites (TFBSs). The proposed models generalize the widely used position weight matrix (PWM) models, Markov models and Bayesian network models. In contrast to these models, where for each position a fixed subset of the remaining positions is used to model dependencies, in VOBN models, these subsets may vary based on the specific nucleotides observed, which are called the context. This flexibility turns out to be of advantage for the classification and analysis of TFBSs, as statistical dependencies between nucleotides in different TFBS positions (not necessarily adjacent) may be taken into account efficiently--in a position-specific and context-specific manner. RESULTS: We apply the VOBN model to a set of 238 experimentally verified sigma-70 binding sites in Escherichia coli. We find that the VOBN model can distinguish these 238 sites from a set of 472 intergenic 'non-promoter' sequences with a higher accuracy than fixed-order Markov models or Bayesian trees. We use a replicated stratified-holdout experiment having a fixed true-negative rate of 99.9%. We find that for a foreground inhomogeneous VOBN model of order 1 and a background homogeneous variable-order Markov (VOM) model of order 5, the obtained mean true-positive (TP) rate is 47.56%. In comparison, the best TP rate for the conventional models is 44.39%, obtained from a foreground PWM model and a background 2nd-order Markov model. As the standard deviation of the estimated TP rate is approximately 0.01%, this improvement is highly significant.

Algorithms↗

Interpreting posterior relative risk estimates in disease-mapping studies.

There is currently much interest in conducting spatial analyses of health outcomes at the small-area scale. This requires sophisticated statistical techniques, usually involving Bayesian models, to smooth the underlying risk estimates because the data are typically sparse. However, questions have been raised about the performance of these models for recovering the "true" risk surface, about the influence of the prior structure specified, and about the amount of smoothing of the risks that is actually performed. We describe a comprehensive simulation study designed to address these questions. Our results show that Bayesian disease-mapping models are essentially conservative, with high specificity even in situations with very sparse data but low sensitivity if the raised-risk areas have only a moderate (less than 2-fold) excess or are not based on substantial expected counts (> 50 per area). Semiparametric spatial mixture models typically produce less smoothing than their conditional autoregressive counterpart when there is sufficient information in the data (moderate-size expected count and/or high true excess risk). Sensitivity may be improved by exploiting the whole posterior distribution to try to detect true raised-risk areas rather than just reporting and mapping the mean posterior relative risk. For the widely used conditional autoregressive model, we show that a decision rule based on computing the probability that the relative risk is above 1 with a cutoff between 70 and 80% gives a specific rule with reasonable sensitivity for a range of scenarios having moderate expected counts (approximately 20) and excess risks (approximately 1.5- to 2-fold). Larger (3-fold) excess risks are detected almost certainly using this rule, even when based on small expected counts, although the mean of the posterior distribution is typically smoothed to about half the true value.

Bayes Theorem↗

Feature selection in Bayesian classifiers for the prognosis of survival of cirrhotic patients treated with TIPS.

The transjugular intrahepatic portosystemic shunt (TIPS) is a treatment for cirrhotic patients with portal hypertension. A subgroup of patients dies in the first 6 months and another subgroup lives a long period of time. Nowadays, no risk factors have been identified in order to determine how long a patient will survive. An empirical study for predicting the survival rate within the first 6 months after TIPS placement is conducted using a clinical database with 107 cases and 77 variables. Applications of Bayesian classification models, based on Bayesian networks, to medical problems have become popular in the last years. Feature subset selection is useful due to the heterogeneity of the medical databases where not all the variables are required to perform the classification. In this paper, filter and wrapper approaches based on the feature subset selection are adapted to induce Bayesian classifiers (naive Bayes, selective naive Bayes, semi naive Bayes, tree augmented naive Bayes, and k-dependence Bayesian classifier) and are applied to distinguish between the two subgroups of cirrhotic patients. The estimated accuracies obtained tally with the results of previous studies. Moreover, the medical significance of the subset of variables selected by the classifiers along with the comprehensibility of Bayesian models is greatly appreciated by physicians.

Bayes Theorem↗

Finding scientific topics.

A first step in identifying the content of a document is determining which topics that document addresses. We describe a generative model for documents, introduced by Blei, Ng, and Jordan [Blei, D. M., Ng, A. Y. & Jordan, M. I. (2003) J. Machine Learn. Res. 3, 993-1022], in which each document is generated by choosing a distribution over topics and then choosing each word in the document from a topic selected according to this distribution. We then present a Markov chain Monte Carlo algorithm for inference in this model. We use this algorithm to analyze abstracts from PNAS by using Bayesian model selection to establish the number of topics. We show that the extracted topics capture meaningful structure in the data, consistent with the class designations provided by the authors of the articles, and outline further applications of this analysis, including identifying "hot topics" by examining temporal dynamics and tagging abstracts to illustrate semantic content.

Databases, Factual↗

Bayesian image processing in magnetic resonance imaging.

In the past several years, image processing techniques based on Bayesian models have received considerable attention. In our earlier work, we developed a novel Bayesian approach which was primarily aimed at the processing and reconstruction of images in positron emission tomography. In this paper, we describe how the technique has been adopted to process magnetic resonance images in order to reduce noise and artifacts, thereby improving image quality. In this framework, the image is assumed to be a statistical variable whose posterior probability density conditional on the observed image is modeled by the product of the likelihood function of the observed data with a prior density based our prior knowledge. A Gibbs random field incorporating local continuity information and with edge-detection capability is used as the prior model. Based on the formalism of the posterior density, we can compute an estimate of the image using an iterative technique. We have implemented this technique and applied it to phantom and clinical images. Our results indicate that the approach works reasonably well for reducing noise, enhancing edges, and removing ringing artifact.

Algorithms↗

[Meta-analysis of the Italian studies on short-term effects of air pollution--MISA 1996-2002].

INTRODUCTION: the Italian Meta-analysis of short-term effects of air pollution for the period 1996-2002 (MISA-2) is a planned study on 15 Italian cities, among the larger country towns summing up 9 millions and one hundred thousand inhabitants at 2001 census. HEALTH OUTCOMES DATA: mortality for all natural causes (362254 deaths), for respiratory causes (22317) and cardiovascular causes (146830), and hospital admissions for acute conditions, respiratory (278028 admissions), cardiac (455540) and cerebrovascular (60960), have been considered. Mortality data came from Regional or Local Health Unit Registries, while hospital admissions data have been selected from Regional or Hospital Archives (exclusion percentages range for all admissions between 45% and 82%). For each participating city daily series averaged about 4.3 years, with a minimum of three consecutive years. AIR POLLUTANTS DATA: daily pollutants concentration series (SO2, NO2, CO, PM10, O3) came from air quality monitoring networks of Regional Environmental Protection Agencies, of Environmental Offices of Provinces or Municipalities. Monitors' selection has been done by a working group composed by representatives of monitoring network Agencies. The selection criteria are the representativeness of general population exposure for each specific pollutant, avoiding as possible monitors close to high traffic roads; and the number, quality and location of monitors, selecting around 3-4 monitors with continuous data flow in the period (at least 75% of valid hourly data). The final series has been created averaging over monitors and imputing missing values under proportionality assumptions. Median of Pearson correlation coefficients between pairs of monitors of the each city was 0.62, interquartile range 0.42-0.77. STATISTICAL METHODS: A generalized linear model on daily counts of health events has been fitted for each city. Linear pollutant effect has been specified and bi-pollutant models have been fitted for PM10+NO2 and PMO+O3. Temperature has been modelled parametrically using a change point at 21 degrees C and lagged effects. Humidity, day of the week, national holidays and influenza epidemics (using data from the National Surveillance Programs from 1999) are the other considered confounders. An age-specific natural cubic spline on season has been specified with 5 degree of freedom (on average) per year for mortality and 7 degree of freedom per year for hospital admission data. The base model is age-stratified (0-64, 65-74, 75+ years). Gender, age, season specific models have been fitted, too. Five sensitivity analyses have been done, varying the degree of freedom for the seasonality spline and specifying non parametric functions on temperature. Constrained distributed lag models have been fitted on mortality data to study potential harvesting effects. City-specific results have been meta-analyzed by random effects hierarchical Bayesian model. Four different models have been fitted in the sensitivity analyses, assuming different priors on heterogeneity variance and outlier-resistant prior on city-specific effects. Bayesian meta-regressions have been fitted on base model, bi-pollutant and season-specific city-specific results. Attributable deaths have been estimated by Monte Carlo methods using effect, pollutant, baseline rate distributions. Fourteen different scenarios have been considered for PM10 and ten for NO2 and CO, using meta-analitic and posterior city-specific effect estimates RESULTS: Pollutants effects are reported as percent increase on mortality or hospital admissions for an increase of 10 microg/m3 of SO2, NO2 and PM10, and 1 mg/m3 of CO. We found an increase on mortality for all natural causes associated to increase of air pollutants concentration (for NO2 0.6% 95%CrI 0.3,0.9; CO 1.2% 0.6,1.7; PM10 0.31% -0.2,0.7). Similar findings were found for cardiorespiratory mortality and hospital admissions for respiratory and cardiac diseases. We found no difference by gender. There was a weak evidence of greater effect size in extreme age groups (0-24 months and over 85 years where we found a percent increase in mortality for all natural causes for PM10 of 0.39% CrI95% 0.0,0.8). There was a strong evidence for each pollutant of greater effects in the warm season (1st May-30th September) on mortality and hospital admissions (we found a percent increase in mortality for all natural causes for PM10 in the warm season of 1.95% CrI95% 0.6,3.3). The associations between pollutants concentration and health events were present at different time lags, depending on outcome and exposure. For mortality, the excess risk peaked within few days from the exposure increase (two days for PM10, up to four days for NO2 and CO). Mortality displacement was minor and ended within two weeks. Cumulative effects at fifteen days showed higher risks for respiratory diseases (PM10 1.65% CI95% 0.3,3.0). The results of meta-regressions showed associations between PM10 effects on mortality and hospital admissions, and mortality for all causes (SMR) and PM10/NO2 ratio. The effect modification of temperature was very consistent, and also using bi-pollutant models. Such effect modification was greater during the cold season. We found and overall impact on mortality for all natural causes in the period 1996-2002 between 1.4% and 4.1% of all deaths for gaseous pollutants (NO2 and CO). The estimates were more imprecise for PM10, due to the variability among cities of the effect estimates (0.1%; 3.3%). The limits stated in the European Union directives for 2010 would have been saved about 900 deaths (1.4%) for PM10 or 1400 deaths for NO2 (1.7%) among all the MISA cities, applying posterior city-specific effect estimates.

Adolescent↗

Statistical limitations in functional neuroimaging. I. Non-inferential methods and statistical models.

Functional neuroimaging (FNI) provides experimental access to the intact living brain making it possible to study higher cognitive functions in humans. In this review and in a companion paper in this issue, we discuss some common methods used to analyse FNI data. The emphasis in both papers is on assumptions and limitations of the methods reviewed. There are several methods available to analyse FNI data indicating that none is optimal for all purposes. In order to make optimal use of the methods available it is important to know the limits of applicability. For the interpretation of FNI results it is also important to take into account the assumptions, approximations and inherent limitations of the methods used. This paper gives a brief overview over some non-inferential descriptive methods and common statistical models used in FNI. Issues relating to the complex problem of model selection are discussed. In general, proper model selection is a necessary prerequisite for the validity of the subsequent statistical inference. The non-inferential section describes methods that, combined with inspection of parameter estimates and other simple measures, can aid in the process of model selection and verification of assumptions. The section on statistical models covers approaches to global normalization and some aspects of univariate, multivariate, and Bayesian models. Finally, approaches to functional connectivity and effective connectivity are discussed. In the companion paper we review issues related to signal detection and statistical inference.

Bayes Theorem↗

Bayesian extrapolation of space-time trends in cancer registry data.

We apply a full Bayesian model framework to a dataset on stomach cancer mortality in West Germany. The data are stratified by age group, year, and district. Using an age-period-cohort model with an additional spatial component, our goal is to investigate whether there is evidence for space-time interactions in these data. Furthermore, we will determine whether a period-space or a cohort-space interaction model is more appropriate to predict future mortality rates. The setup will be fully Bayesian based on a series of Gaussian Markov random field priors for each of the components. Statistical inference is based on efficient algorithms to block update Gaussian Markov random fields, which have recently been proposed in the literature.

Bayes Theorem↗

Robust full Bayesian learning for radial basis networks.

We propose a hierarchical full Bayesian model for radial basis networks. This model treats the model dimension (number of neurons), model parameters, regularization parameters, and noise parameters as unknown random variables. We develop a reversible-jump Markov chain Monte Carlo (MCMC) method to perform the Bayesian computation. We find that the results obtained using this method are not only better than the ones reported previously, but also appear to be robust with respect to the prior specification. In addition, we propose a novel and computationally efficient reversible-jump MCMC simulated annealing algorithm to optimize neural networks. This algorithm enables us to maximize the joint posterior distribution of the network parameters and the number of basis function. It performs a global search in the joint space of the parameters and number of parameters, thereby surmounting the problem of local minima to a large extent. We show that by calibrating the full hierarchical Bayesian prior, we can obtain the classical Akaike information criterion, Bayesian information criterion, and minimum description length model selection criteria within a penalized likelihood framework. Finally, we present a geometric convergence theorem for the algorithm with homogeneous transition kernel and a convergence theorem for the reversible-jump MCMC simulated annealing method.

Journal Article↗

Understanding tuberculosis epidemiology using structured statistical models.

Molecular epidemiological studies can provide novel insights into the transmission of infectious diseases such as tuberculosis. Typically, risk factors for transmission are identified using traditional hypothesis-driven statistical methods such as logistic regression. However, limitations become apparent in these approaches as the scope of these studies expand to include additional epidemiological and bacterial genomic data. Here we examine the use of Bayesian models to analyze tuberculosis epidemiology. We begin by exploring the use of Bayesian networks (BNs) to identify the distribution of tuberculosis patient attributes (including demographic and clinical attributes). Using existing algorithms for constructing BNs from observational data, we learned a BN from data about tuberculosis patients collected in San Francisco from 1991 to 1999. We verified that the resulting probabilistic models did in fact capture known statistical relationships. Next, we examine the use of newly introduced methods for representing and automatically constructing probabilistic models in structured domains. We use statistical relational models (SRMs) to model distributions over relational domains. SRMs are ideally suited to richly structured epidemiological data. We use a data-driven method to construct a statistical relational model directly from data stored in a relational database. The resulting model reveals the relationships between variables in the data and describes their distribution. We applied this procedure to the data on tuberculosis patients in San Francisco from 1991 to 1999, their Mycobacterium tuberculosis strains, and data on contact investigations. The resulting statistical relational model corroborated previously reported findings and revealed several novel associations. These models illustrate the potential for this approach to reveal relationships within richly structured data that may not be apparent using conventional statistical approaches. We show that Bayesian methods, in particular statistical relational models, are an important tool for understanding infectious disease epidemiology.

Adult↗

Assessing uncertainty in reference intervals via tolerance intervals: application to a mixed model describing HIV infection.

We define the reference interval as the range between the 2.5th and 97.5th percentiles of a random variable. We use reference intervals to compare characteristics of a marker of disease progression between affected populations. We use a tolerance interval to assess uncertainty in the reference interval. Unlike the tolerance interval, the estimated reference interval does not contains the true reference interval with specified confidence (or credibility). The tolerance interval is easy to understand, communicate and visualize. We derive estimates of the reference interval and its tolerance interval for markers defined by features of a linear mixed model. Examples considered are reference intervals for time trends in HIV viral load, and CD4 per cent, in HIV-infected haemophiliac children and homosexual men. We estimate the intervals with likelihood methods and also develop a Bayesian model in which the parameters are estimated via Markov-chain Monte Carlo. The Bayesian formulation naturally overcomes some important limitations of the likelihood model.

Adult↗