PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Inference and computation with population codes.

In the vertebrate nervous system, sensory stimuli are typically encoded through the concerted activity of large populations of neurons. Classically, these patterns of activity have been treated as encoding the value of the stimulus (e.g., the orientation of a contour), and computation has been formalized in terms of function approximation. More recently, there have been several suggestions that neural computation is akin to a Bayesian inference process, with population activity patterns representing uncertainty about stimuli in the form of probability distributions (e.g., the probability density function over the orientation of a contour). This paper reviews both approaches, with a particular emphasis on the latter, which we see as a very promising framework for future modeling and experimental work.

Animals↗

Hypermedia and randomized algorithms for medical expert systems.

KNET is an environment for constructing probabilistic, knowledge-intensive systems within the axiomatic framework of decision theory. The KNET architecture defines a complete separation between the hypermedia user interface on the one hand, and the representation and management of expert opinion on the other. KNET offers a choice of algorithms for probabilistic inference. We and our coworkers have used KNET to build consultation systems for lymph-node pathology, bone-marrow transplantation therapy, clinical epidemiology, and alarm management in the intensive-care unit. Most important, KNET contains a randomized approximation scheme (RAS) for the difficult and almost certainly intractable problem of Bayesian inference. Our algorithm can, in many circumstances, perform efficient approximate inference in large and richly interconnected models of medical diagnosis. In this article, we describe the architecture of KNET, construct a randomized algorithm for probabilistic inference, and analyze the algorithm's performance. Finally, we characterize our algorithms' empiric behavior and explore its potential for parallel speedups. From design to implementation, then, KNET demonstrates the crucial interaction between theoretical computer science and medical informatics.

Algorithms↗

Predicting dose-time profiles of solar energetic particle events using Bayesian forecasting methods.

Bayesian inference techniques, coupled with Markov chain Monte Carlo sampling methods, are used to predict dose-time profiles for energetic solar particle events. Inputs into the predictive methodology are dose and dose-rate measurements obtained early in the event. Surrogate dose values are grouped in hierarchical models to express relationships among similar solar particle events. Models assume nonlinear, sigmoidal growth for dose throughout an event. Markov chain Monte Carlo methods are used to sample from Bayesian posterior predictive distributions for dose and dose rate. Example predictions are provided for the November 8, 2000, and August 12, 1989, solar particle events.

Bayes Theorem↗

MDL and the statistical mechanics of protein potentials.

The combination of a wealth of structural data and impressive computational power provides detailed information pertaining to the structure and dynamics of biomacromolecules. A natural inclination is to incorporate this information into models to gain added predictive power on protein folding and stability. There has been considerable recent interest in developing "knowledge-based" potentials to describe internal interactions in proteins. In these approaches, probability distribution functions are inferred from existing knowledge. A common assumption has been the "quasi-chemical approximation" or "Boltzmann device". This method relates statistical mechanical probabilities to observed frequencies. The validity of this approach is discussed in detail from a statistical mechanics perspective. Because statistical mechanics is a form of statistical inference based on a lack of knowledge of the system, the "Boltzmann device" does not have a rigorous theoretical justification. In the present work, a statistical mechanics based on partial knowledge of the system is employed. This statistical mechanical scheme uses the minimum description length (MDL) of phase space as its main tool. With this approach, "knowledge-based" potentials can be derived in a rigorous fashion. In practical calculations, these potentials are best obtained using Bayesian inference methods similar to those used in image reconstruction.

Algorithms↗

The problem of multiple inference in studies designed to generate hypotheses.

Epidemiologic research often involves the simultaneous assessment of associations between many risk factors and several disease outcomes. In such situations, often designed to generate hypotheses, multiple univariate hypothesis-testing is not an appropriate basis for inference. The number of true positive associations in a collection of many associations can be estimated by comparing the observed distribution of p values for the positive associations to a theoretical uniform distribution, or to the observed distribution of negative associations, or to an empiric randomization distribution. None of these approaches, however, will distinguish the true from the false positive associations. Various criteria for selecting a subset of associations to report are considered by the authors, including Bonferoni adjustment of p values, splitting the sample for searching and testing, Bayesian inference, and decision theory. The authors prefer an approach in which all associations in the data are reported, whether significant or not, followed by a ranking in order of priority for investigation using empirical Bayes techniques. Methods are illustrated by application to preliminary data from a study aimed at identifying hitherto unsuspected occupational carcinogens.

Bayes Theorem↗

The 'Ideal Homunculus': decoding neural population signals.

Information processing in the nervous system involves the activity of large populations of neurons. It is possible, however, to interpret the activity of relatively small numbers of cells in terms of meaningful aspects of the environment. 'Bayesian inference' provides a systematic and effective method of combining information from multiple cells to accomplish this. It is not a model of a neural mechanism (neither are alternative methods, such as the population vector approach) but a tool for analysing neural signals. It does not require difficult assumptions about the nature of the dimensions underlying cell selectivity, about the distribution and tuning of cell responses or about the way in which information is transmitted and processed. It can be applied to any parameter of neural activity (for example, firing rate or temporal pattern). In this review, we demonstrate the power of Bayesian analysis using examples of visual responses of neurons in primary visual and temporal cortices. We show that interaction between correlation in mean responses to different stimuli (signal) and correlation in response variability within stimuli (noise) can lead to marked improvement of stimulus discrimination using population responses.

Animals↗

Phylogeny and classification of the Digenea (Platyhelminthes: Trematoda).

Complete small subunit ribosomal RNA gene (ssrDNA) and partial (D1-D3) large subunit ribosomal RNA gene (lsrDNA) sequences were used to estimate the phylogeny of the Digenea via maximum parsimony and Bayesian inference. Here we contribute 80 new ssrDNA and 124 new lsrDNA sequences. Fully complementary data sets of the two genes were assembled from newly generated and previously published sequences and comprised 163 digenean taxa representing 77 nominal families and seven aspidogastrean outgroup taxa representing three families. Analyses were conducted on the genes independently as well as combined and separate analyses including only the higher plagiorchiidan taxa were performed using a reduced-taxon alignment including additional characters that could not be otherwise unambiguously aligned. The combined data analyses yielded the most strongly supported results and differences between the two methods of analysis were primarily in their degree of resolution. The Bayesian analysis including all taxa and characters, and incorporating a model of nucleotide substitution (general-time-reversible with among-site rate heterogeneity), was considered the best estimate of the phylogeny and was used to evaluate their classification and evolution. In broad terms, the Digenea forms a dichotomy that is split between a lineage leading to the Brachylaimoidea, Diplostomoidea and Schistosomatoidea (collectively the Diplostomida nomen novum (nom. nov.)) and the remainder of the Digenea (the Plagiorchiida), in which the Bivesiculata nom. nov. and Transversotremata nom. nov. form the two most basal lineages, followed by the Hemiurata. The remainder of the Plagiorchiida forms a large number of independent lineages leading to the crown clade Xiphidiata nom. nov. that comprises the Allocreadioidea, Gorgoderoidea, Microphalloidea and Plagiorchioidea, which are united by the presence of a penetrating stylet in their cercariae. Although a majority of families and to a lesser degree, superfamilies are supported as currently defined, the traditional divisions of the Echinostomida, Plagiorchiida and Strigeida were found to comprise non-natural assemblages. Therefore, the membership of established higher taxa are emended, new taxa erected and a revised, phylogenetically based classification proposed and discussed in light of ontogeny, morphology and taxonomic history.

Animals↗

Prevalence estimates for paratuberculosis adjusted for test variability using Bayesian analysis.

The ELISA tests that are available to detect an infection with Mycobacterium avium subsp. paratuberculosis (MAP) have a limited validity expressed as the sensitivity (Se) and specificity (Sp). In many studies, the Se and Sp of the tests are treated as constants and this will result in an underestimation of the variability of the true prevalence (TP). Bayesian inference provided a natural framework for using information on the test variability (i.e., the uncertainty) in the estimates of test Se and Sp when estimating the TP. Data from two prevalence studies for MAP using an ELISA in several regions in two locations were available for the analyses. In location 1, all cattle of at least 3 years of age were sampled in approximately 90 randomly sampled herds in each of the four regions of the country. In location 2, in 30 randomly sampled herds in each of three regions, approximately 30 randomly selected cows were sampled. Information about the unknown test Se and Sp and MAP prevalence was incorporated into a Bayesian model by joint prior probability distributions. Posterior estimates were obtained by combining the actual likelihood with the prior distributions using Bayes' formula. The corrected cow-level TP (proportion of infected cows in a herd) was low, 5.8 and 3.6% in locations 1 and 2, respectively. Certain regions within a location differed significantly in herd-level TP (proportion of infected herds). The herd-level TP was 54.3% in location 1 (95% credible interval (CI) 46.1, 63.3%) and 32.9% in location 2 (95% CI: 14.4, 73.3%). The variation in the herd-level TP estimate for location 2 was more than three times as large as the variation in location 1 mainly because of the relatively small number of investigated herds in location 2. In future prevalence studies for MAP, sample size calculations should be based on a very low cow-level prevalence. Approximately 50 and 90% of the herds in the current study had an estimated cow-level TP below 4 and 10%, respectively.

Animals↗

Multiple outputation: inference for complex clustered data by averaging analyses from independent data.

This article applies a simple method for settings where one has clustered data, but statistical methods are only available for independent data. We assume the statistical method provides us with a normally distributed estimate, theta, and an estimate of its variance sigma. We randomly select a data point from each cluster and apply our statistical method to this independent data. We repeat this multiple times, and use the average of the associated theta's as our estimate. An estimate of the variance is given by the average of the sigma2's minus the sample variance of the theta's. We call this procedure multiple outputation, as all "excess" data within each cluster is thrown out multiple times. Hoffman, Sen, and Weinberg (2001, Biometrika 88, 1121-1134) introduced this approach for generalized linear models when the cluster size is related to outcome. In this article, we demonstrate the broad applicability of the approach. Applications to angular data, p-values, vector parameters, Bayesian inference, genetics data, and random cluster sizes are discussed. In addition, asymptotic normality of estimates based on all possible outputations, as well as a finite number of outputations, is proven given weak conditions. Multiple outputation provides a simple and broadly applicable method for analyzing clustered data. It is especially suited to settings where methods for clustered data are impractical, but can also be applied generally as a quick and simple tool.

Acute Disease↗

The contributions of Jerome Cornfield to the theory of statistics.

This paper is a review of the contributions of Jerome Cornfield to the theory of statistics. It discusses several highlights of his theoretical work as well as describing his philosophy relating theory to application. The three areas discussed are: linear programming, urn sampling and its generalizations to the analysis of variance, and Bayesian inference. It is not widely known that Jerome Cornfield was perhaps the first to formulate and approximately solve the linear programming problem in 1941. His formulation was made for the famous "Diet Problem". An early publication introduced the method of indicator random variables in the context of urn sampling. This simple method allowed straightforward calculations of the low order moments for estimates arising from sampling finite populations and was later generalized to the two-way analysis of variance. The application of the urn sampling model to the analysis of variance served to illuminate how one chooses proper error terms for making tests in the analysis of variance table. Jerome Cornfield's philosophy on applications of statistics was dominated by a Bayesian outlook. His theoretical contributions in the past two decades were mainly concerned with the development of Bayesian ideas and methods. A brief survey is made of his main contributions to this area. A particularly noteworthy result was his demonstration that for the two-sample slippage problem of location, the likelihood function under a permutation setting is uninformative for the slippage parameter. However, the posterior distribution differs from the prior distribution despite the fact that the likelihood is uninformative.

Bayes Theorem↗

Bayesian analysis of prevalence with covariates using simulation-based techniques: applications to HIV screening.

Ignoring the limited precision of medical diagnostic tests can incur serious bias in prevalence estimation. Conversely, treating the values of sensitivity and specificity as constants, as in most studies, inevitably underestimates the variability of prevalence estimates. Bayesian inference provides a natural framework with which to integrate the variability in the estimates of sensitivity and specificity with estimation of prevalence. However, the resulting model becomes quite complicated and presents a computational challenge. Recently, Mendoza-Blanco et al. proposed a missing-data approach with simulation-based techniques to deal with the computational difficulties. Although their approach is quite effective in reducing the computational complexity into manageable tasks, their developed methodology is not general enough for modelling the effects of covariates in prevalence estimation. In this paper, we extend their work in this direction by combining their missing-data approach with a latent variable technique for modelling discrete data. The present work also generalizes the methods of Albert and Chib for Bayesian analysis of binary response data with errors in the response. We illustrate the methodology with several real data examples extracted from the literature.

AIDS Serodiagnosis↗

Trial-to-trial variability of cortical evoked responses: implications for the analysis of functional connectivity.

OBJECTIVES: The time series of single trial cortical evoked potentials typically have a random appearance, and their trial-to-trial variability is commonly explained by a model in which random ongoing background noise activity is linearly combined with a stereotyped evoked response. In this paper, we demonstrate that more realistic models, incorporating amplitude and latency variability of the evoked response itself, can explain statistical properties of cortical potentials that have often been attributed to stimulus-related changes in functional connectivity or other intrinsic neural parameters. METHODS: Implications of trial-to-trial evoked potential variability for variance, power spectrum, and interdependence measures like cross-correlation and spectral coherence, are first derived analytically. These implications are then illustrated using model simulations and verified experimentally by the analysis of intracortical local field potentials recorded from monkeys performing a visual pattern discrimination task. To further investigate the effects of trial-to-trial variability on the aforementioned statistical measures, a Bayesian inference technique is used to separate single-trial evoked responses from the ongoing background activity. RESULTS: We show that, when the average event-related potential (AERP) is subtracted from single-trial local field potential time series, a stimulus phase-locked component remains in the residual time series, in stark contrast to the assumption of the common model that no such phase-locked component should exist. Two main consequences of this observation are demonstrated for statistical measures that are computed on the residual time series. First, even though the AERP has been subtracted, the power spectral density, computed as a function of time with a short sliding window, can nonetheless show signs of modulation by the AERP waveform. Second, if the residual time series of two channels co-vary, then their cross-correlation and spectral coherence time functions can also be modulated according to the shape of the AERP waveform. Bayesian estimation of single-trial evoked responses provides further proof that these time-dependent statistical changes are due to remnants of the evoked phase-locked component in the residual time series. CONCLUSIONS: Because trial-to-trial variability of the evoked response is commonly ignored as a contributing factor in evoked potential studies, stimulus-related modulations of power spectral density, cross-correlation, and spectral coherence measures is often attributed to dynamic changes of the connectivity within and among neural populations. This work demonstrates that trial-to-trial variability of the evoked response must be considered as a possible explanation of such modulation.

Animals↗

Plastome evolution and phylogenomic relationships in Ajuga (Lamiaceae, Ajugoideae).

BACKGROUND: Ajuga is currently known to include approximately 69 species, with a combined distribution extending throughout Eurasia, Africa, and Australia. Its popularity and significance are largely based on an extensive history of medicinal and horticultural use. It is divided into two sections based on morphological characters, and this sectional classification is also reflected in pronounced geographic patterns. Although previous studies have largely focused on Ajuga sect. Ajuga in East Asia, A. sect. Chamaepithys, which ranges from the Mediterranean to Central Asia, remains insufficiently sampled, thereby limiting a comprehensive understanding of infrageneric sectional relationships within the genus. Here, we generated complete plastid genomes for 12 species representing both sections of the genus and used these data to characterize plastome structure and infer evolutionary relationships. RESULTS: In this study, 21 Ajuga plastomes were analyzed, including 12 newly sequenced plastomes and 9 previously published plastomes representing 19 species. Comparative analyses showed that all plastomes exhibited a highly conserved quadripartite structure, with genome sizes ranging from 149,963 to 150,740 bp and GC contents varying from 38.2% to 38.3%. Each plastome contained 133 genes, including 88 protein-coding genes, 37 transfer RNA genes, and 8 ribosomal RNA genes. The boundaries between the inverted repeat (IR) and single-copy (SC) regions were also highly conserved across species. In addition, 796 simple sequence repeats (SSRs), 874 long repeat sequences (LRSs), and 12 highly variable regions (ccsA-ndhD, ndhF-rpl32, petA-psbJ, rpl32-trnL-UAG, rps2-rpoC2, trnH-GUG-psbA, trnK-UUU-rps16, trnP-UGG-psaJ, trnT-UGU-trnL-UAA, ycf15-trnL-CAA, ndhF, and ycf1) were identified among the 21 plastomes. Phylogenetic analyses based on four datasets and conducted using Maximum Likelihood and Bayesian Inference recovered two major clades corresponding to the traditionally recognized sectional classification, with one distributed from the Mediterranean to Central Asia and the other in East Asia. CONCLUSION: This study represents the most comprehensive plastome-based sampling of Ajuga to date, including representative species from the Mediterranean, Central Asia, and East Asia. Our results have significantly enhanced our understanding of its infrageneric relationships. The plastome resources generated in this study provide a valuable foundation for future research on species delimitation, phylogeny, and the evolutionary history of Ajuga.

Phylogeny↗

Evolutionary HMMs: a Bayesian approach to multiple alignment.

MOTIVATION: We review proposed syntheses of probabilistic sequence alignment, profiling and phylogeny. We develop a multiple alignment algorithm for Bayesian inference in the links model proposed by Thorne et al. (1991, J. Mol. Evol., 33, 114-124). The algorithm, described in detail in Section 3, samples from and/or maximizes the posterior distribution over multiple alignments for any number of DNA or protein sequences, conditioned on a phylogenetic tree. The individual sampling and maximization steps of the algorithm require no more computational resources than pairwise alignment. METHODS: We present a software implementation (Handel) of our algorithm and report test results on (i) simulated data sets and (ii) the structurally informed protein alignments of BAliBASE (Thompson et al., 1999, Nucleic Acids Res., 27, 2682-2690). RESULTS: We find that the mean sum-of-pairs score (a measure of residue-pair correspondence) for the BAliBASE alignments is only 13% lower for Handelthan for CLUSTALW(Thompson et al., 1994, Nucleic Acids Res., 22, 4673-4680), despite the relative simplicity of the links model (CLUSTALW uses affine gap scores and increased penalties for indels in hydrophobic regions). With reference to these benchmarks, we discuss potential improvements to the links model and implications for Bayesian multiple alignment and phylogenetic profiling. AVAILABILITY: The source code to Handelis freely distributed on the Internet at http://www.biowiki.org/Handel under the terms of the GNU Public License (GPL, 2000, http://www.fsf.org./copyleft/gpl.html).

Algorithms↗

Estimation of population pharmacokinetic parameters in the presence of non-compliance.

In population pharmacokinetic (PK) studies, patients' drug plasma profiles are routinely analyzed assuming that all patients took their drug at the times and in the amounts specified. However, patient non-compliance with the prescribed drug regimen is a leading source of failure to drug therapy. It has been reported that over 30% of patients routinely skip doses regardless of their disease, prognosis, or symptoms. This brings into question the assumption regarding full compliance for population PK analyses. This paper describes the estimation of population PK parameters in the presence and absence of non-compliance while either assuming full compliance or estimating compliance using a hierarchical Bayesian approach. Assessment of compliance for a given dose was limited to one of three possibilities: no dose was taken at the prescribed time, the prescribed dose was taken at the prescribed time, or twice the prescribed dose was taken at the prescribed time. Simulated data sets based on a one-compartment pharmacokinetic model with first order elimination were analyzed using WinBUGS (Bayesian inference Using Gibbs Sampling) software. An initial feasibility simulation experiment, using a simple, but informative PK sampling design with bolus input of drug, was performed. A second simulation study was then carried out using a more realistic sampling design and first-order input of drug. The simulated sampling design included observations after known doses as well as after uncertain doses. Results from the feasibility study revealed that when compliance was estimated instead of being assumed to be 100%, the relative prediction error for clearance (CL) decreased from 0.25 to 0.10 for 60% compliance and from 0.6 to 0.2 for 35% compliance. Estimates of the interoccasion variability of clearance were improved by compliance estimation but still had substantial positive bias. Estimated of interindividual variability were relatively insensitive to compliance estimation. Estimates for volume of distribution (V) and its associated variances were not affected by incorporation of compliance estimates, perhaps due to the specific sampling design that was used. The design was relatively uninformative regarding V. In the more realistic study, estimates for CL, V and the difference between the absorption rate constant and the elimination rate constant (KA-K) were improved by the incorporation of compliance estimation. The median relative errors were reduced from 0.51 to -0.01 for CL, from 0.49 to 0.04 for V, and from 0.49 to -0.02 for Ka-K. The bias in interoccasion variances for V and CL appeared to be reduced by compliance estimation while estimates of interindividual variability were not affected in a systematic fashion. The bias in the residual error variance was decreased from a relative error of about 2 to close to 0. The use of hierarchical Bayesian modeling with the incorporation of compliance estimation decreased the bias in the typical value parameter but the effects on variance parameters were less consistent. The encouraging results of these simulation experiments will hopefully stimulate further evaluation of this methodology for the estimation of population pharmacokinetic parameters in the presence of potential patient noncompliance.

Bayes Theorem↗

Bayesian analysis for a single 2 x 2 table.

The simple comparison of two binomial populations is frequently of interest in epidemiology when the domains are large. For small domains, however, there are no exact methods except Fisher's exact test. A basic problem, therefore, is to compare two populations by assessing the difference between the proportions of individuals who possess a characteristic in the first and second populations. When there is prior information, we take the proportions to have independent conjugate beta distributions with known parameters, thereby facilitating a Bayesian analysis. We consider Bayesian inference on functions of the proportions, and the three most common scalar measures used in epidemiology and health services research, namely relative risk, odds ratio and attributable risk. We develop the highest density regions (both exact and approximate) for relative risk, odds ratio and attributable risk. In addition, we consider the Bayes factor for testing whether the model with a common proportion holds rather than one with distinct proportions. Using data from the population-based Worcester Heart Attack Study, we apply our methodology to study gender differences in the therapeutic management of patients with acute myocardial infarction (AMI) by selected demographic and clinical characteristics. The Bayes factor, the approximate and exact intervals generally suggest that there are no substantial differences in the pharmacologic management of males and females hospitalized with AMI.

Adult↗

Identifying the types of missingness in quality of life data from clinical trials.

This paper discusses methods of identifying the types of missingness in quality of life (QOL) data in cancer clinical trials. The first approach involves collecting information on why the QOL questionnaires were not completed. Based on the reasons provided one may be able to distinguish the mechanisms causing missing data. The second approach is to model the missing data mechanism and perform hypothesis testing to determine the missing data processes. Two methods of testing if missing data are missing completely at random (MCAR) are presented and applied to incomplete longitudinal QOL data obtained from international multi-centre cancer clinical trials. The first method (Ridout, 1991) is based on a logistic regression and the second method (Park and Davis, 1993) is based on an adaptation of weighted least squares. In one application (advanced breast cancer) missing data was not likely to be MCAR. In the second application (adjuvant breast cancer) the missing mechanism was dependent on the QOL scale under study. MCAR and missing at random (MAR) have distinct consequences for data analysis. Therefore it is relevant to distinguish between them. However, if either MCAR or MAR hold, likelihood or Bayesian inferences can be based solely on the observed data, although for MAR, depending on the research question, modelling the dropout mechanism may still be necessary. Distinguishing between MAR and missing not at random (MNAR) is not trivial and relies on fundamentally untestable assumptions.

Clinical Trials as Topic↗

Assessing heterogeneity and correlation of paired failure times with the bivariate frailty model.

We consider bivariate survival times for heterogeneous populations, where heterogeneity induces deviations in an individual's risk of an event as well as associations between survival times. The heterogeneity is characterized by a bivariate frailty model. We measure the heterogeneity effects through deviations associated with hazard functions and an association function defined through the conditional hazard functions: the cross-ratio function proposed by Oakes. We show how the deviation and association measures are determined by the frailty distribution. A Gibbs sampling method is developed for Bayesian inferences on regression coefficients, frailty parameters and the heterogeneity measures. The method is applied to a mental health care data set.

Algorithms↗