PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Fossil calibrations and molecular divergence time estimates in centrarchid fishes (Teleostei: Centrarchidae).

Molecular clock methods allow biologists to estimate divergence times, which in turn play an important role in comparative studies of many evolutionary processes. It is well known that molecular age estimates can be biased by heterogeneity in rates of molecular evolution, but less attention has been paid to the issue of potentially erroneous fossil calibrations. In this study we estimate the timing of diversification in Centrarchidae, an endemic major lineage of the diverse North American freshwater fish fauna, through a new approach to fossil calibration and molecular evolutionary model selection. Given a completely resolved multi-gene molecular phylogeny and a set of multiple fossil-inferred age estimates, we tested for potentially erroneous fossil calibrations using a recently developed fossil cross-validation. We also used fossil information to guide the selection of the optimal molecular evolutionary model with a new fossil jackknife method in a fossil-based model cross-validation. The centrarchid phylogeny resulted from a mixed-model Bayesian strategy that included 14 separate data partitions sampled from three mtDNA and four nuclear genes. Ten of the 31 interspecific nodes in the centrarchid phylogeny were assigned a minimal age estimate from the centrarchid fossil record. Our analyses identified four fossil dates that were inconsistent with the other fossils, and we removed them from the molecular dating analysis. Using fossil-based model cross-validation to determine the optimal smoothing value in penalized likelihood analysis, and six mutually consistent fossil calibrations, the age of the most recent common ancestor of Centrarchidae was 33.59 million years ago (mya). Penalized likelihood analyses of individual data partitions all converged on a very similar age estimate for this node, indicating that rate heterogeneity among data partitions is not confounding our analyses. These results place the origin of the centrarchid radiation at a time of major faunal turnover as the fossil record indicates that the most diverse lineages of the North American freshwater fish fauna originated at the Eocene-Oligocene boundary, approximately 34 mya. This time coincided with major global climate change from warm to cool temperatures and a signature of elevated lineage extinction and origination in the fossil record across the tree of life. Our analyses demonstrate the utility of fossil cross-validation to critically assess individual fossil calibration points, providing the ability to discriminate between consistent and inconsistent fossil age estimates that are used for calibrating molecular phylogenies.

Animals↗

Commentary: practical advantages of Bayesian analysis of epidemiologic data.

In the past decade, there have been enormous advances in the use of Bayesian methodology for analysis of epidemiologic data, and there are now many practical advantages to the Bayesian approach. Bayesian models can easily accommodate unobserved variables such as an individual's true disease status in the presence of diagnostic error. The use of prior probability distributions represents a powerful mechanism for incorporating information from previous studies and for controlling confounding. Posterior probabilities can be used as easily interpretable alternatives to p values. Recent developments in Markov chain Monte Carlo methodology facilitate the implementation of Bayesian analyses of complex data sets containing missing observations and multidimensional outcomes. Tools are now available that allow epidemiologists to take advantage of this powerful approach to assessment of exposure-disease relations.

Bayes Theorem↗

Using Bayesian networks to model expected and unexpected operational losses.

This report describes the use of Bayesian networks (BNs) to model statistical loss distributions in financial operational risk scenarios. Its focus is on modeling "long" tail, or unexpected, loss events using mixtures of appropriate loss frequency and severity distributions where these mixtures are conditioned on causal variables that model the capability or effectiveness of the underlying controls process. The use of causal modeling is discussed from the perspective of exploiting local expertise about process reliability and formally connecting this knowledge to actual or hypothetical statistical phenomena resulting from the process. This brings the benefit of supplementing sparse data with expert judgment and transforming qualitative knowledge about the process into quantitative predictions. We conclude that BNs can help combine qualitative data from experts and quantitative data from historical loss databases in a principled way and as such they go some way in meeting the requirements of the draft Basel II Accord (Basel, 2004) for an advanced measurement approach (AMA).

Journal Article↗

An application of Bayesian QTL mapping to early development in double haploid lines of rainbow trout including environmental effects.

A Bayesian model and variable dimensional parameter estimation based on Markov chain Monte Carlo was applied to map quantitative trait loci (QTLs) in a doubled haploid mapping population of rainbow trout. To increase power, the analysis was performed using the multiple-QTL model, which simultaneously accounted for all the environmental and genetic main effects that influence the expression of early development life history traits. By doing so we obtained the posterior estimated effects for the environmental factors as well as the number, positions, and the effects for the QTLs. The analyses revealed QTLs for time at hatching, embryonic length and weight at swim-up stage. The posterior expectation of the number of QTLs in different linkage groups shows that at least four QTLs are needed to explain the observed differences in early development between the clonal lines. The Bayesian method effectively combined all the information available to accurately position these QTLs in the rainbow trout genome.

Animals↗

Background estimation in experimental spectra

A general probabilistic technique for estimating background contributions to measured spectra is presented. A Bayesian model is used to capture the defining characteristics of the problem, namely, that the background is smoother than the signal. The signal is allowed to have positive and/or negative components. The background is represented in terms of a cubic spline basis. A variable degree of smoothness of the background is attained by allowing the number of knots and the knot positions to be adaptively chosen on the basis of the data. The fully Bayesian approach taken provides a natural way to handle knot adaptivity and allows uncertainties in the background to be estimated. Our technique is demonstrated on a particle induced x-ray emission spectrum from a geological sample and an Auger spectrum from iron, which contains signals with both positive and negative components.

Journal Article↗

Comparative analysis of methodologies for the detection of horizontally transferred genes: a reassessment of first-order Markov models.

With the advent of larger genome databases detection of horizontal gene transfer events has been transformed into an increasingly important issue. Here we present a simple theoretical analysis based on the in silico artificial addition of known foreign genes from different prokaryotic groups into the genome of Escherichia coli K12 MG1655. Using this dataset as a control, we have tested the efficiency of four methodologies commonly employed to detect HTG (Horizontally transferred genes), which are based on (a) the codon adaptation index, codon usage, and GC percentage (CAI/GC); (b) a distributional profile (DP) approach made by a gene search in the closely related phylogenetic genomes; (c) a Bayesian model (BM); and (d) a first-order Markov model (MM). All methods exhibit limitations although, as shown here, the BM and the MM are better approximations. Moreover, the MM has demonstrated a more accurate rate of detections when genes from closely related organisms are evaluated. The application of the MM to detect recently transferred genes in the genomes of E. coli strains K12 MG1655, O157 EDL933, and Salmonella typhimurium, shows that these organisms have undergone a rather significant amount of HTG, most of which appear to be pseudogenes. Few of these sequences that have undergone HGT appear to have well defined functions and may be involved in the organism's adaptation.

Computer Simulation↗

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100× more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study↗

Online updating of space-time disease surveillance models via particle filters.

Online surveillance of disease has become an important issue in public health. In particular, the space-time monitoring of disease plays an important part in any syndromic system. However, methodology for these systems is generally lacking. One approach to space-time monitoring of health data is to consider the space-time model parameters as the focus and to monitor their changes as multivariate time series (Lawson AB. Some considerations in spatial-temporal analysis of public health surveillance data. In Brookmeyer R, Stroup DF eds. Monitoring the Health of Populations. Oxford University Press, 2004; Vidal Rodeiro CL, Lawson AB. Monitoring changes in spatio-temporal maps of disease. Biometrical Journal 2006; to appear). However with complex space-time models, this becomes very time consuming. Some simplifications may be necessary and these can be made in a number of ways. In this article, the focus is on particle filters that can be used to resample the history of the process and thereby reduce computation time. This article describes a particular case of particle filters, the resample-move algorithm, proposed by Gilks and Berzuini (Gilks WR, Berzuini C. Following a moving target--Monte Carlo inference for dynamic Bayesian models. Journal of the Royal Statistical Society, Series B 2001; 63: 127-46), in the context of disease map surveillance. This is followed by an application to a real data set in which a comparison between the use of Markov chain Monte Carlo methods and the resample-move algorithm is carried out.

Algorithms↗

On the nature of over-dispersion in motor vehicle crash prediction models.

Statistical modeling of traffic crashes has been of interest to researchers for decades. Over the most recent decade many crash models have accounted for extra-variation in crash counts--variation over and above that accounted for by the Poisson density. The extra--variation--or dispersion--is theorized to capture unaccounted for variation in crashes across sites. The majority of studies have assumed fixed dispersion parameters in over-dispersed crash models--tantamount to assuming that unaccounted for variation is proportional to the expected crash count. Miaou and Lord [Miaou, S.P., Lord, D., 2003. Modeling traffic crash-flow relationships for intersections: dispersion parameter, functional form, and Bayes versus empirical Bayes methods. Transport. Res. Rec. 1840, 31-40] challenged the fixed dispersion parameter assumption, and examined various dispersion parameter relationships when modeling urban signalized intersection accidents in Toronto. They suggested that further work is needed to determine the appropriateness of the findings for rural as well as other intersection types, to corroborate their findings, and to explore alternative dispersion functions. This study builds upon the work of Miaou and Lord, with exploration of additional dispersion functions, the use of an independent data set, and presents an opportunity to corroborate their findings. Data from Georgia are used in this study. A Bayesian modeling approach with non-informative priors is adopted, using sampling-based estimation via Markov Chain Monte Carlo (MCMC) and the Gibbs sampler. A total of eight model specifications were developed; four of them employed traffic flows as explanatory factors in mean structure while the remainder of them included geometric factors in addition to major and minor road traffic flows. The models were compared and contrasted using the significance of coefficients, standard deviance, chi-square goodness-of-fit, and deviance information criteria (DIC) statistics. The findings indicate that the modeling of the dispersion parameter, which essentially explains the extra-variance structure, depends greatly on how the mean structure is modeled. In the presence of a well-defined mean function, the extra-variance structure generally becomes insignificant, i.e. the variance structure is a simple function of the mean. It appears that extra-variation is a function of covariates when the mean structure (expected crash count) is poorly specified and suffers from omitted variables. In contrast, when sufficient explanatory variables are used to model the mean (expected crash count), extra-Poisson variation is not significantly related to these variables. If these results are generalizable, they suggest that model specification may be improved by testing extra-variation functions for significance. They also suggest that known influences of expected crash counts are likely to be different than factors that might help to explain unaccounted for variation in crashes across sites.

Accidents, Traffic↗

A transformation approach for incorporating monotone or unimodal constraints.

Samples of curves are collected in many applications, including studies of reproductive hormone levels in the menstrual cycle. Many approaches have been proposed for correlated functional data of this type, including smoothing spline methods and other flexible parametric modeling strategies. In many cases, the underlying biological processes involved restrict the curve to follow a particular shape. For example, progesterone levels in healthy women increase during the menstrual cycle to a peak achieved at random location with decreases thereafter. Reproductive epidemiologists are interested in studying the distribution of the peak and the trajectory for women in different groups. Motivated by this application, we propose a simple approach for restricting each woman's mean trajectory to follow an umbrella shape. An unconstrained hierarchical Bayesian model is used to characterize the data, and draws from the posterior distribution obtained using a Gibbs sampler are then mapped to the constrained space. Inferences are based on the resulting quasi-posterior distribution for the peak and individual woman trajectories. The methods are applied to a study comparing progesterone trajectories for conception and nonconception cycles.

Bayes Theorem↗

Discriminating between rate heterogeneity and interspecific recombination in DNA sequence alignments with phylogenetic factorial hidden Markov models.

MOTIVATION: A recently proposed method for detecting recombination in DNA sequence alignments is based on the combination of hidden Markov models (HMMs) with phylogenetic trees. Although this method was found to detect breakpoints of recombinant regions more accurately than most existing techniques, it inherently fails to distinguish between recombination and rate variation. In the present paper, we propose to marry the phylogenetic tree to a factorial HMM (FHMM). The states of the first hidden chain represent tree topologies, whereas the states of the second independent hidden chain represent different global scaling factors of the branch lengths. Inference is done in terms of a hierarchical Bayesian model, where parameters and hidden states are sampled from the posterior distribution with Gibbs sampling. RESULTS: We have tested the proposed model on various synthetic and real-world DNA sequence alignments. The simulation results suggest that as opposed to the standard phylogenetic HMM, the phylogenetic FHMM clearly distinguishes between recombination and rate heterogeneity and thereby avoids the prediction of spurious recombinant regions. AVAILABILITY: The proposed method has been implemented in a MATLAB package that extends Kevin Murphy's HMM toolbox. Software and data used in our study are available from http://www.bioss.sari.ac.uk/~dirk/Supplements

Algorithms↗

Children use categories to maximize accuracy in estimation.

The present study tests a model of category effects upon stimulus estimation in children. Prior work with adults suggests that people inductively generalize distributional information about a category of stimuli and use this information to adjust their estimates of individual stimuli in a way that maximizes average accuracy in estimation (see Huttenlocher, Hedges & Vevea, 2000). However, little is known about the developmental origin of this cognitive process. In the present study, 5- and 7-year-old children viewed stimuli that varied in size and reproduced each from memory. Consistent with the predictions of a Bayesian model of category effects on estimation, responses were adjusted toward the central value of the stimulus distribution. Additionally, the dispersion of the stimulus distribution affected the pattern of bias and variability of responses in a way that is predicted by the model. The results suggest that, like adults, children use categories for increasing average accuracy in estimating inexact stimuli.

Age Factors↗

An illustration of the modelling of cost and efficacy data from a clinical trial.

Health care providers, purchasers and policy makers need to make informed decisions regarding the provision of cost-effective care. When a new health care intervention is to be compared with the current standard, an economic evaluation alongside an evaluation of health benefits provides useful information for the decision making process. We consider the information on cost-effectiveness which arises from an individual clinical trial comparing the two interventions. Recent methods for conducting a cost-effectiveness analysis for a clinical trial have focused on the net benefit parameter. The net benefit parameter, a function of costs and health benefits, is positive if the new intervention is cost-effective compared with the standard. In this paper we describe frequentist and Bayesian approaches to cost-effectiveness analysis which have been suggested in the literature and apply them to data from a clinical trial comparing laparoscopic surgery with open mesh surgery for the repair of inguinal hernias. We extend the Bayesian model to allow the total cost to be divided into a number of different components. The advantages and disadvantages of the different approaches are discussed. In January 2001, NICE issued guidance on the type of surgery to be used for inguinal hernia repair. We discuss our example in the light of this information.

Bayes Theorem↗

A Bayesian approach for the evaluation of six diagnostic assays and the estimation of Cryptosporidium prevalence in dairy calves.

The prevalence of Cryptosporidium in calves and the test properties of six diagnostic assays (microscopy (ME), an immunofluorescence assay (IFA), two ELISA and two PCR assays) were estimated using Bayesian analysis. In a first Bayesian approach, the test results of the four conventional techniques were used: ME, IFA and two ELISA. This four-test approach estimated that the calf prevalence was 17% (95% Probability Interval (PI): 0.1-0.28) and that the specificity estimates of the IFA and ELISA were high compared to ME. A six-test Bayesian model was developed using the test results of the 4 conventional assays and 2 PCR assays, resulting in a higher calf prevalence estimate (58% with a 95% PI: 0.5-0.66) and in a different test evaluation: the sensitivity estimates of the conventional techniques decreased in the six-test approach, due to the inclusion of two PCR assays with a higher sensitivity compared to the conventional techniques. The specificity estimates of these conventional assays were comparable in the four-test and six-test approach. These results both illustrate the potential and the pitfalls of a Bayesian analysis in estimating prevalence and test characteristics, since posterior estimates are variables depending both on the data at hand and prior information included in the analysis. The need for sensitive diagnostic assays in epidemiological studies is demonstrated, especially for the identification of subclinically infected animals since the PCR assays identify these animals with reduced oocyst excretion, which the conventional techniques fail to identify.

Animals↗

Simultaneous detection of linkage disequilibrium and genetic differentiation of subdivided populations.

We propose a new method for simultaneously detecting linkage disequilibrium and genetic structure in subdivided populations. Taking subpopulation structure into account with a hierarchical model, we estimate the magnitude of genetic differentiation and linkage disequilibrium in a metapopulation on the basis of geographical samples, rather than decompose a population into a finite number of random-mating subpopulations. We assume that Hardy-Weinberg equilibrium is satisfied in each locality, but do not assume independence between marker loci. Linkage states remain unknown. Genetic differentiation and linkage disequilibrium are expressed as hyperparameters describing the prior distribution of genotypes or haplotypes. We estimate related parameters by maximizing marginal-likelihood functions and detect linkage equilibrium or disequilibrium by the Akaike information criterion. Our empirical Bayesian model analyzes genotype and haplotype frequencies regardless of haploid or diploid data, so it can be applied to most commonly used genetic markers. The performance of our procedure is examined via numerical simulations in comparison with classical procedures. Finally, we analyze isozyme data of ayu, a severely exploited fish species, and single-nucleotide polymorphisms in human ALDH2.

Animals↗

Bayesian mapping of multiple sclerosis prevalence in the province of Pavia, northern Italy.

The geographical analysis of a disease risk is particularly difficult when the disease is non-frequent and the area units are small. The practical use of the Bayesian modelling, instead of the classical frequentist one, is applied to study the geographical variation of multiple sclerosis (MS) across the province of Pavia, Northern Italy. 464 MS-affected individuals resident in the province of Pavia were identified on December 31st 2000. The overall prevalence was 94 per 100,000 inhabitants. This estimate indicates an increasing MS prevalence in the province, in accordance with the vast majority of the Italian areas where prevalence studies have been repeated. We mapped the geographical variation of MS prevalence across the 190 communes of the province both with a classical approach and a Bayesian approach. The frequentist approach produced an extremely dishomogeneous map, while the Bayesian map was much smoother and more interpretable. Our study underlines the usefulness of Bayesian methods to obtain reliable maps of disease prevalence and to identify possible clusters of disease where to carry out further epidemiological investigations.

Adolescent↗

Cerebrospinal fluid pharmacokinetics and penetration of continuous infusion topotecan in children with central nervous system tumors.

The purpose of this study was to describe the cerebrospinal fluid (CSF) penetration of topotecan in humans, to generate a pharmacokinetic model to simultaneously describe topotecan lactone and total concentrations in the plasma and CSF, and to characterize the CSF and plasma pharmacokinetics of topotecan administered as a continuous infusion (CI). Plasma and CSF samples were collected from 17 patients receiving 5.5 or 7.5 mg/m2 per day as a 24-h CI (5 patients, 7 courses), or 0.5 to 1.25 mg/m2 per day as a 72-h CI (12 patients, 12 courses). CSF samples were obtained from either a ventricular reservoir (VR) or a lumbar puncture (LP). Topotecan lactone and total (lactone plus hydroxy acid) concentrations were determined by HPLC and fluorescence detection. Using MAP-Bayesian modelling, a three-compartment model was fitted simultaneously to topotecan lactone and total concentrations in the plasma and CSF. The penetration of topotecan into the CSF was determined from the ratio of the CSF to the plasma area under the concentration-time curve. The median CSF ventricular lactone concentrations, obtained prior to the end of infusion (EOI), were 0.86, 1.4, 0.73, 5.3, and 4.6 ng/ml for patients receiving 0.5, 1.0, 1.25, 5.5, and 7.5 mg/m2 per day, respectively. EOI CSF lumbar lactone concentrations measured in three patients were 0.44, 1.1, and 1.7 ng/ml for topotecan doses of 1.0, 5.5, and 7.5 mg/m2 per day, respectively. In two patients receiving 1.25 mg/m2 per day, EOI CSF concentrations were obtained simultaneously from a VR and LP; the lumbar lactone concentrations were 30% and 49% lower than the ventricular concentrations. During a 24-h and a 72-h CI, the median CSF penetration of topotecan lactone was 0.29 (range 0.10 to 0.59) and 0.42 (range 0.11 to 0.86), respectively. A three-compartment model adequately described topotecan lactone and total concentrations in the plasma and CSF. Topotecan was therefore found to significantly penetrate into the CSF in humans. The pharmacokinetic model presented may be useful in the design of clinical studies of topotecan to treat CNS tumors.

Adolescent↗

Resolving multisensory conflict: a strategy for balancing the costs and benefits of audio-visual integration.

In order to maintain a coherent, unified percept of the external environment, the brain must continuously combine information encoded by our different sensory systems. Contemporary models suggest that multisensory integration produces a weighted average of sensory estimates, where the contribution of each system to the ultimate multisensory percept is governed by the relative reliability of the information it provides (maximum-likelihood estimation). In the present study, we investigate interactions between auditory and visual rate perception, where observers are required to make judgments in one modality while ignoring conflicting rate information presented in the other. We show a gradual transition between partial cue integration and complete cue segregation with increasing inter-modal discrepancy that is inconsistent with mandatory implementation of maximum-likelihood estimation. To explain these findings, we implement a simple Bayesian model of integration that is also able to predict observer performance with novel stimuli. The model assumes that the brain takes into account prior knowledge about the correspondence between auditory and visual rate signals, when determining the degree of integration to implement. This provides a strategy for balancing the benefits accrued by integrating sensory estimates arising from a common source, against the costs of conflating information relating to independent objects or events.

Acoustic Stimulation↗