PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19Linked to original sources

Identification of transcription factor binding sites with variable-order Bayesian networks.

MOTIVATION: We propose a new class of variable-order Bayesian network (VOBN) models for the identification of transcription factor binding sites (TFBSs). The proposed models generalize the widely used position weight matrix (PWM) models, Markov models and Bayesian network models. In contrast to these models, where for each position a fixed subset of the remaining positions is used to model dependencies, in VOBN models, these subsets may vary based on the specific nucleotides observed, which are called the context. This flexibility turns out to be of advantage for the classification and analysis of TFBSs, as statistical dependencies between nucleotides in different TFBS positions (not necessarily adjacent) may be taken into account efficiently--in a position-specific and context-specific manner. RESULTS: We apply the VOBN model to a set of 238 experimentally verified sigma-70 binding sites in Escherichia coli. We find that the VOBN model can distinguish these 238 sites from a set of 472 intergenic 'non-promoter' sequences with a higher accuracy than fixed-order Markov models or Bayesian trees. We use a replicated stratified-holdout experiment having a fixed true-negative rate of 99.9%. We find that for a foreground inhomogeneous VOBN model of order 1 and a background homogeneous variable-order Markov (VOM) model of order 5, the obtained mean true-positive (TP) rate is 47.56%. In comparison, the best TP rate for the conventional models is 44.39%, obtained from a foreground PWM model and a background 2nd-order Markov model. As the standard deviation of the estimated TP rate is approximately 0.01%, this improvement is highly significant.

Algorithms↗

Perceptual distortions of speed at low luminance: evidence inconsistent with a Bayesian account of speed encoding.

Our perception of speed has been shown to be distorted under a number of viewing conditions. Recently the well-known reduction of perceived speed at low contrast has led to Bayesian models of speed perception that account for these distortions with a slow speed 'prior'. To test the predictive, rather than the descriptive, power of the Bayesian approach we have investigated perceived speed at low luminance. Our results indicate that, for the mesopic and photopic range (0.13-30 cd m(-2)) the perceived speed of lower luminance patterns is virtually unaffected at low speeds (<4 deg s(-1)) but is over-estimated at higher speeds (>4 deg s(-1)). We show here that the results can be accounted for by an extension to a simple ratio model of speed encoding [Hammett, S. T., Champion, R. A., Morland, A. & Thompson, P. G. (2005). A ratio model of perceived speed in the human visual system. Proceedings of Royal Society B, 262, 2351-2356.] that takes account of known changes in neural responses as a function of luminance, contrast and temporal frequency. The results are not consistent with current Bayesian approaches to modelling speed encoding that postulate a slow speed prior.

Bayes Theorem↗

Interpreting posterior relative risk estimates in disease-mapping studies.

There is currently much interest in conducting spatial analyses of health outcomes at the small-area scale. This requires sophisticated statistical techniques, usually involving Bayesian models, to smooth the underlying risk estimates because the data are typically sparse. However, questions have been raised about the performance of these models for recovering the "true" risk surface, about the influence of the prior structure specified, and about the amount of smoothing of the risks that is actually performed. We describe a comprehensive simulation study designed to address these questions. Our results show that Bayesian disease-mapping models are essentially conservative, with high specificity even in situations with very sparse data but low sensitivity if the raised-risk areas have only a moderate (less than 2-fold) excess or are not based on substantial expected counts (> 50 per area). Semiparametric spatial mixture models typically produce less smoothing than their conditional autoregressive counterpart when there is sufficient information in the data (moderate-size expected count and/or high true excess risk). Sensitivity may be improved by exploiting the whole posterior distribution to try to detect true raised-risk areas rather than just reporting and mapping the mean posterior relative risk. For the widely used conditional autoregressive model, we show that a decision rule based on computing the probability that the relative risk is above 1 with a cutoff between 70 and 80% gives a specific rule with reasonable sensitivity for a range of scenarios having moderate expected counts (approximately 20) and excess risks (approximately 1.5- to 2-fold). Larger (3-fold) excess risks are detected almost certainly using this rule, even when based on small expected counts, although the mean of the posterior distribution is typically smoothed to about half the true value.

Bayes Theorem↗

Feature selection in Bayesian classifiers for the prognosis of survival of cirrhotic patients treated with TIPS.

The transjugular intrahepatic portosystemic shunt (TIPS) is a treatment for cirrhotic patients with portal hypertension. A subgroup of patients dies in the first 6 months and another subgroup lives a long period of time. Nowadays, no risk factors have been identified in order to determine how long a patient will survive. An empirical study for predicting the survival rate within the first 6 months after TIPS placement is conducted using a clinical database with 107 cases and 77 variables. Applications of Bayesian classification models, based on Bayesian networks, to medical problems have become popular in the last years. Feature subset selection is useful due to the heterogeneity of the medical databases where not all the variables are required to perform the classification. In this paper, filter and wrapper approaches based on the feature subset selection are adapted to induce Bayesian classifiers (naive Bayes, selective naive Bayes, semi naive Bayes, tree augmented naive Bayes, and k-dependence Bayesian classifier) and are applied to distinguish between the two subgroups of cirrhotic patients. The estimated accuracies obtained tally with the results of previous studies. Moreover, the medical significance of the subset of variables selected by the classifiers along with the comprehensibility of Bayesian models is greatly appreciated by physicians.

Bayes Theorem↗

Finding scientific topics.

A first step in identifying the content of a document is determining which topics that document addresses. We describe a generative model for documents, introduced by Blei, Ng, and Jordan [Blei, D. M., Ng, A. Y. & Jordan, M. I. (2003) J. Machine Learn. Res. 3, 993-1022], in which each document is generated by choosing a distribution over topics and then choosing each word in the document from a topic selected according to this distribution. We then present a Markov chain Monte Carlo algorithm for inference in this model. We use this algorithm to analyze abstracts from PNAS by using Bayesian model selection to establish the number of topics. We show that the extracted topics capture meaningful structure in the data, consistent with the class designations provided by the authors of the articles, and outline further applications of this analysis, including identifying "hot topics" by examining temporal dynamics and tagging abstracts to illustrate semantic content.

Databases, Factual↗

Modeling individual differences in cognition.

Many evaluations of cognitive models rely on data that have been averaged or aggregated across all experimental subjects, and so fail to consider the possibility of important individual differences between subjects. Other evaluations are done at the single-subject level, and so fail to benefit from the reduction of noise that data averaging or aggregation potentially provides. To overcome these weaknesses, we have developed a general approach to modeling individual differences using families of cognitive models in which different groups of subjects are identified as having different psychological behavior. Separate models with separate parameterizations are applied to each group of subjects, and Bayesian model selection is used to determine the appropriate number of groups. We evaluate this individual differences approach in a simulation study and show that it is superior in terms of the key modeling goals of prediction and understanding. We also provide two practical demonstrations of the approach, one using the ALCOVE model of category learning with data from four previously analyzed category learning experiments, the other using multidimensional scaling representational models with previously analyzed similarity data for colors. In both demonstrations, meaningful individual differences are found and the psychological models are able to account for this variation through interpretable differences in parameterization. The results highlight the potential of extending cognitive models to consider individual differences.

Cognition↗

Bayesian image processing in magnetic resonance imaging.

In the past several years, image processing techniques based on Bayesian models have received considerable attention. In our earlier work, we developed a novel Bayesian approach which was primarily aimed at the processing and reconstruction of images in positron emission tomography. In this paper, we describe how the technique has been adopted to process magnetic resonance images in order to reduce noise and artifacts, thereby improving image quality. In this framework, the image is assumed to be a statistical variable whose posterior probability density conditional on the observed image is modeled by the product of the likelihood function of the observed data with a prior density based our prior knowledge. A Gibbs random field incorporating local continuity information and with edge-detection capability is used as the prior model. Based on the formalism of the posterior density, we can compute an estimate of the image using an iterative technique. We have implemented this technique and applied it to phantom and clinical images. Our results indicate that the approach works reasonably well for reducing noise, enhancing edges, and removing ringing artifact.

Algorithms↗

[Meta-analysis of the Italian studies on short-term effects of air pollution--MISA 1996-2002].

INTRODUCTION: the Italian Meta-analysis of short-term effects of air pollution for the period 1996-2002 (MISA-2) is a planned study on 15 Italian cities, among the larger country towns summing up 9 millions and one hundred thousand inhabitants at 2001 census. HEALTH OUTCOMES DATA: mortality for all natural causes (362254 deaths), for respiratory causes (22317) and cardiovascular causes (146830), and hospital admissions for acute conditions, respiratory (278028 admissions), cardiac (455540) and cerebrovascular (60960), have been considered. Mortality data came from Regional or Local Health Unit Registries, while hospital admissions data have been selected from Regional or Hospital Archives (exclusion percentages range for all admissions between 45% and 82%). For each participating city daily series averaged about 4.3 years, with a minimum of three consecutive years. AIR POLLUTANTS DATA: daily pollutants concentration series (SO2, NO2, CO, PM10, O3) came from air quality monitoring networks of Regional Environmental Protection Agencies, of Environmental Offices of Provinces or Municipalities. Monitors' selection has been done by a working group composed by representatives of monitoring network Agencies. The selection criteria are the representativeness of general population exposure for each specific pollutant, avoiding as possible monitors close to high traffic roads; and the number, quality and location of monitors, selecting around 3-4 monitors with continuous data flow in the period (at least 75% of valid hourly data). The final series has been created averaging over monitors and imputing missing values under proportionality assumptions. Median of Pearson correlation coefficients between pairs of monitors of the each city was 0.62, interquartile range 0.42-0.77. STATISTICAL METHODS: A generalized linear model on daily counts of health events has been fitted for each city. Linear pollutant effect has been specified and bi-pollutant models have been fitted for PM10+NO2 and PMO+O3. Temperature has been modelled parametrically using a change point at 21 degrees C and lagged effects. Humidity, day of the week, national holidays and influenza epidemics (using data from the National Surveillance Programs from 1999) are the other considered confounders. An age-specific natural cubic spline on season has been specified with 5 degree of freedom (on average) per year for mortality and 7 degree of freedom per year for hospital admission data. The base model is age-stratified (0-64, 65-74, 75+ years). Gender, age, season specific models have been fitted, too. Five sensitivity analyses have been done, varying the degree of freedom for the seasonality spline and specifying non parametric functions on temperature. Constrained distributed lag models have been fitted on mortality data to study potential harvesting effects. City-specific results have been meta-analyzed by random effects hierarchical Bayesian model. Four different models have been fitted in the sensitivity analyses, assuming different priors on heterogeneity variance and outlier-resistant prior on city-specific effects. Bayesian meta-regressions have been fitted on base model, bi-pollutant and season-specific city-specific results. Attributable deaths have been estimated by Monte Carlo methods using effect, pollutant, baseline rate distributions. Fourteen different scenarios have been considered for PM10 and ten for NO2 and CO, using meta-analitic and posterior city-specific effect estimates RESULTS: Pollutants effects are reported as percent increase on mortality or hospital admissions for an increase of 10 microg/m3 of SO2, NO2 and PM10, and 1 mg/m3 of CO. We found an increase on mortality for all natural causes associated to increase of air pollutants concentration (for NO2 0.6% 95%CrI 0.3,0.9; CO 1.2% 0.6,1.7; PM10 0.31% -0.2,0.7). Similar findings were found for cardiorespiratory mortality and hospital admissions for respiratory and cardiac diseases. We found no difference by gender. There was a weak evidence of greater effect size in extreme age groups (0-24 months and over 85 years where we found a percent increase in mortality for all natural causes for PM10 of 0.39% CrI95% 0.0,0.8). There was a strong evidence for each pollutant of greater effects in the warm season (1st May-30th September) on mortality and hospital admissions (we found a percent increase in mortality for all natural causes for PM10 in the warm season of 1.95% CrI95% 0.6,3.3). The associations between pollutants concentration and health events were present at different time lags, depending on outcome and exposure. For mortality, the excess risk peaked within few days from the exposure increase (two days for PM10, up to four days for NO2 and CO). Mortality displacement was minor and ended within two weeks. Cumulative effects at fifteen days showed higher risks for respiratory diseases (PM10 1.65% CI95% 0.3,3.0). The results of meta-regressions showed associations between PM10 effects on mortality and hospital admissions, and mortality for all causes (SMR) and PM10/NO2 ratio. The effect modification of temperature was very consistent, and also using bi-pollutant models. Such effect modification was greater during the cold season. We found and overall impact on mortality for all natural causes in the period 1996-2002 between 1.4% and 4.1% of all deaths for gaseous pollutants (NO2 and CO). The estimates were more imprecise for PM10, due to the variability among cities of the effect estimates (0.1%; 3.3%). The limits stated in the European Union directives for 2010 would have been saved about 900 deaths (1.4%) for PM10 or 1400 deaths for NO2 (1.7%) among all the MISA cities, applying posterior city-specific effect estimates.

Adolescent↗

Statistical limitations in functional neuroimaging. I. Non-inferential methods and statistical models.

Functional neuroimaging (FNI) provides experimental access to the intact living brain making it possible to study higher cognitive functions in humans. In this review and in a companion paper in this issue, we discuss some common methods used to analyse FNI data. The emphasis in both papers is on assumptions and limitations of the methods reviewed. There are several methods available to analyse FNI data indicating that none is optimal for all purposes. In order to make optimal use of the methods available it is important to know the limits of applicability. For the interpretation of FNI results it is also important to take into account the assumptions, approximations and inherent limitations of the methods used. This paper gives a brief overview over some non-inferential descriptive methods and common statistical models used in FNI. Issues relating to the complex problem of model selection are discussed. In general, proper model selection is a necessary prerequisite for the validity of the subsequent statistical inference. The non-inferential section describes methods that, combined with inspection of parameter estimates and other simple measures, can aid in the process of model selection and verification of assumptions. The section on statistical models covers approaches to global normalization and some aspects of univariate, multivariate, and Bayesian models. Finally, approaches to functional connectivity and effective connectivity are discussed. In the companion paper we review issues related to signal detection and statistical inference.

Bayes Theorem↗

Bayesian extrapolation of space-time trends in cancer registry data.

We apply a full Bayesian model framework to a dataset on stomach cancer mortality in West Germany. The data are stratified by age group, year, and district. Using an age-period-cohort model with an additional spatial component, our goal is to investigate whether there is evidence for space-time interactions in these data. Furthermore, we will determine whether a period-space or a cohort-space interaction model is more appropriate to predict future mortality rates. The setup will be fully Bayesian based on a series of Gaussian Markov random field priors for each of the components. Statistical inference is based on efficient algorithms to block update Gaussian Markov random fields, which have recently been proposed in the literature.

Bayes Theorem↗

Robust full Bayesian learning for radial basis networks.

We propose a hierarchical full Bayesian model for radial basis networks. This model treats the model dimension (number of neurons), model parameters, regularization parameters, and noise parameters as unknown random variables. We develop a reversible-jump Markov chain Monte Carlo (MCMC) method to perform the Bayesian computation. We find that the results obtained using this method are not only better than the ones reported previously, but also appear to be robust with respect to the prior specification. In addition, we propose a novel and computationally efficient reversible-jump MCMC simulated annealing algorithm to optimize neural networks. This algorithm enables us to maximize the joint posterior distribution of the network parameters and the number of basis function. It performs a global search in the joint space of the parameters and number of parameters, thereby surmounting the problem of local minima to a large extent. We show that by calibrating the full hierarchical Bayesian prior, we can obtain the classical Akaike information criterion, Bayesian information criterion, and minimum description length model selection criteria within a penalized likelihood framework. Finally, we present a geometric convergence theorem for the algorithm with homogeneous transition kernel and a convergence theorem for the reversible-jump MCMC simulated annealing method.

Journal Article↗

Understanding tuberculosis epidemiology using structured statistical models.

Molecular epidemiological studies can provide novel insights into the transmission of infectious diseases such as tuberculosis. Typically, risk factors for transmission are identified using traditional hypothesis-driven statistical methods such as logistic regression. However, limitations become apparent in these approaches as the scope of these studies expand to include additional epidemiological and bacterial genomic data. Here we examine the use of Bayesian models to analyze tuberculosis epidemiology. We begin by exploring the use of Bayesian networks (BNs) to identify the distribution of tuberculosis patient attributes (including demographic and clinical attributes). Using existing algorithms for constructing BNs from observational data, we learned a BN from data about tuberculosis patients collected in San Francisco from 1991 to 1999. We verified that the resulting probabilistic models did in fact capture known statistical relationships. Next, we examine the use of newly introduced methods for representing and automatically constructing probabilistic models in structured domains. We use statistical relational models (SRMs) to model distributions over relational domains. SRMs are ideally suited to richly structured epidemiological data. We use a data-driven method to construct a statistical relational model directly from data stored in a relational database. The resulting model reveals the relationships between variables in the data and describes their distribution. We applied this procedure to the data on tuberculosis patients in San Francisco from 1991 to 1999, their Mycobacterium tuberculosis strains, and data on contact investigations. The resulting statistical relational model corroborated previously reported findings and revealed several novel associations. These models illustrate the potential for this approach to reveal relationships within richly structured data that may not be apparent using conventional statistical approaches. We show that Bayesian methods, in particular statistical relational models, are an important tool for understanding infectious disease epidemiology.

Adult↗

Assessing uncertainty in reference intervals via tolerance intervals: application to a mixed model describing HIV infection.

We define the reference interval as the range between the 2.5th and 97.5th percentiles of a random variable. We use reference intervals to compare characteristics of a marker of disease progression between affected populations. We use a tolerance interval to assess uncertainty in the reference interval. Unlike the tolerance interval, the estimated reference interval does not contains the true reference interval with specified confidence (or credibility). The tolerance interval is easy to understand, communicate and visualize. We derive estimates of the reference interval and its tolerance interval for markers defined by features of a linear mixed model. Examples considered are reference intervals for time trends in HIV viral load, and CD4 per cent, in HIV-infected haemophiliac children and homosexual men. We estimate the intervals with likelihood methods and also develop a Bayesian model in which the parameters are estimated via Markov-chain Monte Carlo. The Bayesian formulation naturally overcomes some important limitations of the likelihood model.

Adult↗

Bayes' theorem in ophthalmologic computer diagnosis.

Computers are being investigated as diagnostic aids in many fields of medicine. Models employing Bayes' theorem, a statistical formula, commonly are used to supply valuable information on the likelihood of each disease in the differential diagnosis to help the clinician make the diagnosis. However, knowledge of elementary decision analysis is beneficial to help understand the current and potential uses of these models. We discussed Bayes' theorem as an introduction to decision analysis. Moreover, we described a Bayesian model for the differential diagnosis of leukocoria to illustrate the application of computers to ophthalmologic diagnosis.

Decision Making↗

[Data analysis by statistical models].

The basic idea for the realization of effective statistical data analysis is illustrated with an example. The use of statistical models is explained and the feasibility of objective comparison of the models by an information criterion AIC is demonstrated. Further, the possibility of practical use of Bayesian models for complex data analysis is explained. Finally, the necessity of cooperation between the experts of respective fields and statisticians for further development of statistical data analysis is mentioned.

Adult↗

[Cytotoxic chemotherapy in elderly patients: present and future].

Cancer in elderly people accounts for more than 50% of the malignant tumors treated per year in France and this population of patients has a rather high-life expectancy. Chemotherapy is active in these elderly patients but clearly more toxic than for young ones. The general tendency among the physicians to empirically reduce the doses is due to the known increased risk of unexpected toxicities. That is why there is such a large variety of conflicting opinions in the literature concerning the benefit and toxic effects of cytostatic drugs in the elderly. Therefore, it appears consistent to adjust chemotherapy regimen according to physiological criteria. Among them is biological age which is a better parameter than chronological age to describe the biological heterogeneity of this population of patients. Nakamura et al have published an interesting model for the calculation of biological age by principal component analysis using 11 easily measurable biological and clinical variables in a series of healthy elderly people. This kind of approach is not at present available for cancer patients but it allows to demonstrate that the chronological age is only one among many other age-related variables and is not sufficient to fully describe it. The variations in pharmacokinetic data are more frequent in the elderly than in younger people and this reflects age-related physiological heterogeneity. This factor is well taken into account in recently described population pharmacokinetic models, bayesian fittings and adaptative control which may represent promising approaches of cytostatics dose adjustments. Such models have been successfully developed in young patients receiving doxorubicin, methotrexate, melphalan and teniposide. They require a low number of blood samples to determine individual parameters and further adjust the doses, and are therefore of potential interest in old patients. Prospective studies are warranted in the future in order to recommend their use in the elderly.

Aged↗

Using Bayesian inference to perform meta-analysis.

Bayesian modeling offers an elegant approach to meta-analysis that efficiently incorporates all sources of variability and relevant quantifiable external information. It provides a more informative summary of the likely value of parameters after observing the data than do non-Bayesian approaches. This leads to direct probabilistic inference about model parameters such as the average treatment effect, the between-study variance, and individual study treatment effects. The latter are weighted averages of the common mean and individual study means with weights reflecting the amount of information provided by each study relative to the others. Homogeneity among these posterior study estimates indicates that pooling these studies is appropriate; heterogeneity suggests that some cause of between-study variation should be explored. The author describes the construction of such models and shows how to use them to estimate a common mean and regression slopes. Two examples illustrate the additional inferences available with the Bayesian methodology.

Angiotensin-Converting Enzyme Inhibitors↗

Bayesian inference for hierarchical mixtures-of-experts with applications to regression and classification.

This paper studies the problems of inference and prediction in a class of models known as hierarchical mixtures-of-experts (HME). The statistical model underlying an HME is a mixture model in which both the mixture coefficients and the mixture components are generalized linear models. Bayesian inference regarding an HME's parameters is presented in the contexts of regression and classification using Markov chain Monte Carlo methods. A benefit of this Bayesian approach is the ability to obtain a sample from the posterior distribution of any functional of the parameters of the given model. In this way, more information is obtained than provided by a point estimate. The methods are illustrated on a nonlinear regression problem and on a breast cancer classification problem. The results indicate that the HME showed good prediction performance, and also gave the additional benefit of providing for the opportunity to assess the degree of certainty of the model in its predictions.

Algorithms↗