PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian inference”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Predicting dose-time profiles of solar energetic particle events using Bayesian forecasting methods.

Bayesian inference techniques, coupled with Markov chain Monte Carlo sampling methods, are used to predict dose-time profiles for energetic solar particle events. Inputs into the predictive methodology are dose and dose-rate measurements obtained early in the event. Surrogate dose values are grouped in hierarchical models to express relationships among similar solar particle events. Models assume nonlinear, sigmoidal growth for dose throughout an event. Markov chain Monte Carlo methods are used to sample from Bayesian posterior predictive distributions for dose and dose rate. Example predictions are provided for the November 8, 2000, and August 12, 1989, solar particle events.

Bayes Theorem↗

MDL and the statistical mechanics of protein potentials.

The combination of a wealth of structural data and impressive computational power provides detailed information pertaining to the structure and dynamics of biomacromolecules. A natural inclination is to incorporate this information into models to gain added predictive power on protein folding and stability. There has been considerable recent interest in developing "knowledge-based" potentials to describe internal interactions in proteins. In these approaches, probability distribution functions are inferred from existing knowledge. A common assumption has been the "quasi-chemical approximation" or "Boltzmann device". This method relates statistical mechanical probabilities to observed frequencies. The validity of this approach is discussed in detail from a statistical mechanics perspective. Because statistical mechanics is a form of statistical inference based on a lack of knowledge of the system, the "Boltzmann device" does not have a rigorous theoretical justification. In the present work, a statistical mechanics based on partial knowledge of the system is employed. This statistical mechanical scheme uses the minimum description length (MDL) of phase space as its main tool. With this approach, "knowledge-based" potentials can be derived in a rigorous fashion. In practical calculations, these potentials are best obtained using Bayesian inference methods similar to those used in image reconstruction.

Algorithms↗

The problem of multiple inference in studies designed to generate hypotheses.

Epidemiologic research often involves the simultaneous assessment of associations between many risk factors and several disease outcomes. In such situations, often designed to generate hypotheses, multiple univariate hypothesis-testing is not an appropriate basis for inference. The number of true positive associations in a collection of many associations can be estimated by comparing the observed distribution of p values for the positive associations to a theoretical uniform distribution, or to the observed distribution of negative associations, or to an empiric randomization distribution. None of these approaches, however, will distinguish the true from the false positive associations. Various criteria for selecting a subset of associations to report are considered by the authors, including Bonferoni adjustment of p values, splitting the sample for searching and testing, Bayesian inference, and decision theory. The authors prefer an approach in which all associations in the data are reported, whether significant or not, followed by a ranking in order of priority for investigation using empirical Bayes techniques. Methods are illustrated by application to preliminary data from a study aimed at identifying hitherto unsuspected occupational carcinogens.

Bayes Theorem↗

The 'Ideal Homunculus': decoding neural population signals.

Information processing in the nervous system involves the activity of large populations of neurons. It is possible, however, to interpret the activity of relatively small numbers of cells in terms of meaningful aspects of the environment. 'Bayesian inference' provides a systematic and effective method of combining information from multiple cells to accomplish this. It is not a model of a neural mechanism (neither are alternative methods, such as the population vector approach) but a tool for analysing neural signals. It does not require difficult assumptions about the nature of the dimensions underlying cell selectivity, about the distribution and tuning of cell responses or about the way in which information is transmitted and processed. It can be applied to any parameter of neural activity (for example, firing rate or temporal pattern). In this review, we demonstrate the power of Bayesian analysis using examples of visual responses of neurons in primary visual and temporal cortices. We show that interaction between correlation in mean responses to different stimuli (signal) and correlation in response variability within stimuli (noise) can lead to marked improvement of stimulus discrimination using population responses.

Animals↗

The contributions of Jerome Cornfield to the theory of statistics.

This paper is a review of the contributions of Jerome Cornfield to the theory of statistics. It discusses several highlights of his theoretical work as well as describing his philosophy relating theory to application. The three areas discussed are: linear programming, urn sampling and its generalizations to the analysis of variance, and Bayesian inference. It is not widely known that Jerome Cornfield was perhaps the first to formulate and approximately solve the linear programming problem in 1941. His formulation was made for the famous "Diet Problem". An early publication introduced the method of indicator random variables in the context of urn sampling. This simple method allowed straightforward calculations of the low order moments for estimates arising from sampling finite populations and was later generalized to the two-way analysis of variance. The application of the urn sampling model to the analysis of variance served to illuminate how one chooses proper error terms for making tests in the analysis of variance table. Jerome Cornfield's philosophy on applications of statistics was dominated by a Bayesian outlook. His theoretical contributions in the past two decades were mainly concerned with the development of Bayesian ideas and methods. A brief survey is made of his main contributions to this area. A particularly noteworthy result was his demonstration that for the two-sample slippage problem of location, the likelihood function under a permutation setting is uninformative for the slippage parameter. However, the posterior distribution differs from the prior distribution despite the fact that the likelihood is uninformative.

Bayes Theorem↗

Bayesian analysis of prevalence with covariates using simulation-based techniques: applications to HIV screening.

Ignoring the limited precision of medical diagnostic tests can incur serious bias in prevalence estimation. Conversely, treating the values of sensitivity and specificity as constants, as in most studies, inevitably underestimates the variability of prevalence estimates. Bayesian inference provides a natural framework with which to integrate the variability in the estimates of sensitivity and specificity with estimation of prevalence. However, the resulting model becomes quite complicated and presents a computational challenge. Recently, Mendoza-Blanco et al. proposed a missing-data approach with simulation-based techniques to deal with the computational difficulties. Although their approach is quite effective in reducing the computational complexity into manageable tasks, their developed methodology is not general enough for modelling the effects of covariates in prevalence estimation. In this paper, we extend their work in this direction by combining their missing-data approach with a latent variable technique for modelling discrete data. The present work also generalizes the methods of Albert and Chib for Bayesian analysis of binary response data with errors in the response. We illustrate the methodology with several real data examples extracted from the literature.

AIDS Serodiagnosis↗

Trial-to-trial variability of cortical evoked responses: implications for the analysis of functional connectivity.

OBJECTIVES: The time series of single trial cortical evoked potentials typically have a random appearance, and their trial-to-trial variability is commonly explained by a model in which random ongoing background noise activity is linearly combined with a stereotyped evoked response. In this paper, we demonstrate that more realistic models, incorporating amplitude and latency variability of the evoked response itself, can explain statistical properties of cortical potentials that have often been attributed to stimulus-related changes in functional connectivity or other intrinsic neural parameters. METHODS: Implications of trial-to-trial evoked potential variability for variance, power spectrum, and interdependence measures like cross-correlation and spectral coherence, are first derived analytically. These implications are then illustrated using model simulations and verified experimentally by the analysis of intracortical local field potentials recorded from monkeys performing a visual pattern discrimination task. To further investigate the effects of trial-to-trial variability on the aforementioned statistical measures, a Bayesian inference technique is used to separate single-trial evoked responses from the ongoing background activity. RESULTS: We show that, when the average event-related potential (AERP) is subtracted from single-trial local field potential time series, a stimulus phase-locked component remains in the residual time series, in stark contrast to the assumption of the common model that no such phase-locked component should exist. Two main consequences of this observation are demonstrated for statistical measures that are computed on the residual time series. First, even though the AERP has been subtracted, the power spectral density, computed as a function of time with a short sliding window, can nonetheless show signs of modulation by the AERP waveform. Second, if the residual time series of two channels co-vary, then their cross-correlation and spectral coherence time functions can also be modulated according to the shape of the AERP waveform. Bayesian estimation of single-trial evoked responses provides further proof that these time-dependent statistical changes are due to remnants of the evoked phase-locked component in the residual time series. CONCLUSIONS: Because trial-to-trial variability of the evoked response is commonly ignored as a contributing factor in evoked potential studies, stimulus-related modulations of power spectral density, cross-correlation, and spectral coherence measures is often attributed to dynamic changes of the connectivity within and among neural populations. This work demonstrates that trial-to-trial variability of the evoked response must be considered as a possible explanation of such modulation.

Animals↗

Plastome evolution and phylogenomic relationships in Ajuga (Lamiaceae, Ajugoideae).

BACKGROUND: Ajuga is currently known to include approximately 69 species, with a combined distribution extending throughout Eurasia, Africa, and Australia. Its popularity and significance are largely based on an extensive history of medicinal and horticultural use. It is divided into two sections based on morphological characters, and this sectional classification is also reflected in pronounced geographic patterns. Although previous studies have largely focused on Ajuga sect. Ajuga in East Asia, A. sect. Chamaepithys, which ranges from the Mediterranean to Central Asia, remains insufficiently sampled, thereby limiting a comprehensive understanding of infrageneric sectional relationships within the genus. Here, we generated complete plastid genomes for 12 species representing both sections of the genus and used these data to characterize plastome structure and infer evolutionary relationships. RESULTS: In this study, 21 Ajuga plastomes were analyzed, including 12 newly sequenced plastomes and 9 previously published plastomes representing 19 species. Comparative analyses showed that all plastomes exhibited a highly conserved quadripartite structure, with genome sizes ranging from 149,963 to 150,740 bp and GC contents varying from 38.2% to 38.3%. Each plastome contained 133 genes, including 88 protein-coding genes, 37 transfer RNA genes, and 8 ribosomal RNA genes. The boundaries between the inverted repeat (IR) and single-copy (SC) regions were also highly conserved across species. In addition, 796 simple sequence repeats (SSRs), 874 long repeat sequences (LRSs), and 12 highly variable regions (ccsA-ndhD, ndhF-rpl32, petA-psbJ, rpl32-trnL-UAG, rps2-rpoC2, trnH-GUG-psbA, trnK-UUU-rps16, trnP-UGG-psaJ, trnT-UGU-trnL-UAA, ycf15-trnL-CAA, ndhF, and ycf1) were identified among the 21 plastomes. Phylogenetic analyses based on four datasets and conducted using Maximum Likelihood and Bayesian Inference recovered two major clades corresponding to the traditionally recognized sectional classification, with one distributed from the Mediterranean to Central Asia and the other in East Asia. CONCLUSION: This study represents the most comprehensive plastome-based sampling of Ajuga to date, including representative species from the Mediterranean, Central Asia, and East Asia. Our results have significantly enhanced our understanding of its infrageneric relationships. The plastome resources generated in this study provide a valuable foundation for future research on species delimitation, phylogeny, and the evolutionary history of Ajuga.

Phylogeny↗

Evolutionary HMMs: a Bayesian approach to multiple alignment.

MOTIVATION: We review proposed syntheses of probabilistic sequence alignment, profiling and phylogeny. We develop a multiple alignment algorithm for Bayesian inference in the links model proposed by Thorne et al. (1991, J. Mol. Evol., 33, 114-124). The algorithm, described in detail in Section 3, samples from and/or maximizes the posterior distribution over multiple alignments for any number of DNA or protein sequences, conditioned on a phylogenetic tree. The individual sampling and maximization steps of the algorithm require no more computational resources than pairwise alignment. METHODS: We present a software implementation (Handel) of our algorithm and report test results on (i) simulated data sets and (ii) the structurally informed protein alignments of BAliBASE (Thompson et al., 1999, Nucleic Acids Res., 27, 2682-2690). RESULTS: We find that the mean sum-of-pairs score (a measure of residue-pair correspondence) for the BAliBASE alignments is only 13% lower for Handelthan for CLUSTALW(Thompson et al., 1994, Nucleic Acids Res., 22, 4673-4680), despite the relative simplicity of the links model (CLUSTALW uses affine gap scores and increased penalties for indels in hydrophobic regions). With reference to these benchmarks, we discuss potential improvements to the links model and implications for Bayesian multiple alignment and phylogenetic profiling. AVAILABILITY: The source code to Handelis freely distributed on the Internet at http://www.biowiki.org/Handel under the terms of the GNU Public License (GPL, 2000, http://www.fsf.org./copyleft/gpl.html).

Algorithms↗

Bayesian analysis for a single 2 x 2 table.

The simple comparison of two binomial populations is frequently of interest in epidemiology when the domains are large. For small domains, however, there are no exact methods except Fisher's exact test. A basic problem, therefore, is to compare two populations by assessing the difference between the proportions of individuals who possess a characteristic in the first and second populations. When there is prior information, we take the proportions to have independent conjugate beta distributions with known parameters, thereby facilitating a Bayesian analysis. We consider Bayesian inference on functions of the proportions, and the three most common scalar measures used in epidemiology and health services research, namely relative risk, odds ratio and attributable risk. We develop the highest density regions (both exact and approximate) for relative risk, odds ratio and attributable risk. In addition, we consider the Bayes factor for testing whether the model with a common proportion holds rather than one with distinct proportions. Using data from the population-based Worcester Heart Attack Study, we apply our methodology to study gender differences in the therapeutic management of patients with acute myocardial infarction (AMI) by selected demographic and clinical characteristics. The Bayes factor, the approximate and exact intervals generally suggest that there are no substantial differences in the pharmacologic management of males and females hospitalized with AMI.

Adult↗

Identifying the types of missingness in quality of life data from clinical trials.

This paper discusses methods of identifying the types of missingness in quality of life (QOL) data in cancer clinical trials. The first approach involves collecting information on why the QOL questionnaires were not completed. Based on the reasons provided one may be able to distinguish the mechanisms causing missing data. The second approach is to model the missing data mechanism and perform hypothesis testing to determine the missing data processes. Two methods of testing if missing data are missing completely at random (MCAR) are presented and applied to incomplete longitudinal QOL data obtained from international multi-centre cancer clinical trials. The first method (Ridout, 1991) is based on a logistic regression and the second method (Park and Davis, 1993) is based on an adaptation of weighted least squares. In one application (advanced breast cancer) missing data was not likely to be MCAR. In the second application (adjuvant breast cancer) the missing mechanism was dependent on the QOL scale under study. MCAR and missing at random (MAR) have distinct consequences for data analysis. Therefore it is relevant to distinguish between them. However, if either MCAR or MAR hold, likelihood or Bayesian inferences can be based solely on the observed data, although for MAR, depending on the research question, modelling the dropout mechanism may still be necessary. Distinguishing between MAR and missing not at random (MNAR) is not trivial and relies on fundamentally untestable assumptions.

Clinical Trials as Topic↗

Assessing heterogeneity and correlation of paired failure times with the bivariate frailty model.

We consider bivariate survival times for heterogeneous populations, where heterogeneity induces deviations in an individual's risk of an event as well as associations between survival times. The heterogeneity is characterized by a bivariate frailty model. We measure the heterogeneity effects through deviations associated with hazard functions and an association function defined through the conditional hazard functions: the cross-ratio function proposed by Oakes. We show how the deviation and association measures are determined by the frailty distribution. A Gibbs sampling method is developed for Bayesian inferences on regression coefficients, frailty parameters and the heterogeneity measures. The method is applied to a mental health care data set.

Algorithms↗

A bayesian analysis for spatial processes with application to disease mapping.

In epidemiology, maps of disease rates and disease risk provide a spatial perspective for researching disease aetiology. For rare diseases or when the population base is small, the rate and risk estimates may be unstable. We propose using a Bayesian analysis based on the conditional autoregressive (CAR) process that will spatially smooth disease rates or risk estimates by allowing each site to 'borrow strength' from its neighbours. Covariates may be included in the model in such a way as to establish a possible association between risk factors and disease incidence. Bayesian inferences are implemented from a direct resampling scheme where large samples are generated from the various posterior distributions. The methodology is demonstrated with a simulation that assesses the effect of sample size and the model parameters on inferences for the parameters. Our approach is also used to spatially smooth district lip cancer rates in Scotland using the CAR model with a covariate that allows for exposure to sunlight.

Bayes Theorem↗

Genetic variance components analysis for binary phenotypes using generalized linear mixed models (GLMMs) and Gibbs sampling.

The common complex diseases such as asthma are an important focus of genetic research, and studies based on large numbers of simple pedigrees ascertained from population-based sampling frames are becoming commonplace. Many of the genetic and environmental factors causing these diseases are unknown and there is often a strong residual covariance between relatives even after all known determinants are taken into account. This must be modelled correctly whether scientific interest is focused on fixed effects, as in an association analysis, or on the covariances themselves. Analysis is straightforward for multivariate Normal phenotypes, but difficulties arise with other types of trait. Generalized linear mixed models (GLMMs) offer a potentially unifying approach to analysis for many classes of phenotype including multivariate Normal traits, binary traits, and censored survival times. Markov Chain Monte Carlo methods, including Gibbs sampling, provide a convenient framework within which such models may be fitted. In this paper, Bayesian inference Using Gibbs Sampling (a generic Gibbs sampler; BUGS) is used to fit GLMMs for multivariate Normal and binary phenotypes in nuclear families. BUGS is easy to use and readily available. We motivate a suitable model structure for Normal phenotypes and show how the model extends to binary traits. We discuss parameter interpretation and statistical inference and show how to circumvent a number of important theoretical and practical problems that we encountered. Using simulated data we show that model parameters seem consistent and appear unbiased in smaller data sets. We illustrate our methods using data from an ongoing cohort study.

Binomial Distribution↗

Modelling the cumulative risk for a false-positive under repeated screening events.

Screening examinations are widely utilized in detecting the presence of medical disorders, for instance, screening mammograms and clinical breast examinations for detection of breast cancer. Such procedures are invaluable in enabling early treatment but produce the possibilities of false-positive and false-negative diagnoses. Focusing on false-positive results, with increasing number of screening events, it is clear that the risk of a false-positive increases. The objective of this paper is to quantify the cumulative risk associated with repeated screening. We provide a very general framework within which to investigate this risk, both at the population and the individual level. The latter allows incorporation of evolving patient medical history to permit individualized assessment of risk. We model cumulative risk in terms of the number of screening events until first false-positive. We develop models which are essentially familiar actuarial models for life table data adding a Cox regression to enable individual level modelling. Because it offers several advantages, we employ a Bayesian inference framework and apply our modelling to the analysis of 9773 screening mammograms collected from 2227 women at an HMO serving nearly 300000 adults in and around Boston, MA.

Adult↗

Variance components analysis for pedigree-based censored survival data using generalized linear mixed models (GLMMs) and Gibbs sampling in BUGS.

Complex human diseases are an increasingly important focus of genetic research. Many of the determinants of these diseases are unknown and there is often a strong residual covariance between relatives even when all known genetic and environmental factors have been taken into account. This must be modeled correctly whether scientific interest is focused on fixed effects, as in an association analysis, or on the covariance structure itself. Analysis is straightforward for multivariate normally distributed traits, but difficulties arise with other types of trait. Generalized linear mixed models (GLMMs) offer a potentially unifying approach to analysis for many classes of phenotype including right censored survival times. This includes age-at-onset and age-at-death data and a variety of other censored traits. Markov chain Monte Carlo (MCMC) methods, including Gibbs sampling, provide a convenient framework within which such GLMMs may be fitted. In this paper, we use BUGS ("Bayesian inference using Gibbs sampling": a readily available, generic Gibbs sampler) to fit GLMMs for right-censored survival times in nuclear and extended families. We discuss parameter interpretation and statistical inference, and show how to circumvent a number of important theoretical and practical problems. Using simulated data, we show that model parameters are consistent. We further illustrate our methods using data from an ongoing cohort study. Finally, we propose that the random effects associated with a genetic component of variance (e.g., sigma(2)(A)) in a GLMM may be regarded as an adjusted "phenotype" and used as input to a conventional model-based or model-free linkage analysis. This provides a simple way to conduct a linkage analysis for a trait reflected in a right-censored survival time while comprehensively adjusting for observed confounders at the level of the individual and latent environmental effects shared across families.

Bayes Theorem↗

Insights Into the Structural Features, Codon Usage Patterns, and Phylogenetic Analysis in Neoniphon argenteus (Teleostei: Holocentriformes) Based on Complete Mitochondrial Genome.

Neoniphon argenteus, a widely distributed nocturnal coral reef fish in the family Holocentridae, plays an important role in maintaining coral reef ecosystem health, yet its phylogenetic position remains poorly resolved. To bridge this gap, we sequenced and analyzed the complete mitochondrial genome of a specimen from the South China Sea to characterize its structural features, codon usage patterns, and phylogenetic relationships. The 16,569 bp mitogenome (GenBank: PP190474.1) encodes 13 protein-coding genes (PCGs), 22 tRNAs, two rRNAs, and two non-coding regions, exhibiting a distinct A + T bias. All tRNAs fold into typical cloverleaf secondary structures except tRNA-Ser (AGN), which lacks the dihydrouridine (DHU) arm. The control region contains palindromic motifs (TACAT/ATGTA) capable of forming hairpin structures and five conserved sequence blocks, whereas the OL region harbors a conserved 5'-GCCGG-3' motif. RSCU analysis revealed 31 frequently used codons (RSCU > 1) with a pronounced preference for A/C-ending codons. The ΔRSCU method identified 10 candidate optimal codons (GCA, CAA, GAA, GGA, AUU, CUA, CCA, CGA, ACA, and GUC). Selection pressure analysis using EasyCodeML and site-specific models indicated that all PCGs are predominantly under purifying selection, with no significant evidence of pervasive positive selection. ND6 exhibited elevated pairwise Ka/Ks ratios (mean = 1.209 ± 0.047), consistent with reduced selective constraint rather than adaptive evolution. Phylogenetic analysis of 19 Holocentriformes species using maximum likelihood and Bayesian inference with partitioned models based on 13 PCGs and two rRNA genes (12S and 16S) assigned all taxa to two well-supported subfamilies (Holocentrinae and Myripristinae). Within Holocentrinae, Neoniphon species form a monophyletic clade nested within a paraphyletic Sargocentron, suggesting that the genus Sargocentron as currently defined is not monophyletic. This study provides useful baseline molecular data for further exploration of the evolutionary history of N. argenteus and other members of Holocentriformes.

Holocentridae↗

Bayesian technique for investigating linearity in event-related BOLD fMRI.

Event-related BOLD fMRI data is modeled as a linear time-invariant system. Together with Bayesian inference techniques, a statistical test is developed for rigorously detecting linearity/nonlinearity in the BOLD response system. The test is applied to data collected from eight subjects using an event-related paradigm with a switching checkerboard as the visual stimulus. Analyzed as a group, the results clearly find the response to be nonlinear. When each subject is analyzed individually, however, the results are predominantly nonlinear, but there is some evidence to suggest that there may be a crossover from a linear to a nonlinear regime and vice versa. This could be important when estimating physiological parameters for individuals. Additionally, estimates of the hemodynamic response function and corresponding response were obtained, but there was no consistent appearance of a poststimulus undershoot in the event-related BOLD response.

Adult↗