PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

How the probability of a false positive affects the value of DNA evidence.

Errors in sample handling or test interpretation may cause false positives in forensic DNA testing. This article uses a Bayesian model to show how the potential for a false positive affects the evidentiary value of DNA evidence and the sufficiency of DNA evidence to meet traditional legal standards for conviction. The Bayesian analysis is contrasted with the "false positive fallacy," an intuitively appealing but erroneous alternative interpretation. The findings show the importance of having accurate information about both the random match probability and the false positive probability when evaluating DNA evidence. It is argued that ignoring or underestimating the potential for a false positive can lead to serious errors of interpretation, particularly when the suspect is identified through a "DNA dragnet" or database search, and that ignorance of the true rate of error creates an important element of uncertainty about the value of DNA evidence.

Bayes Theorem↗

A statistical model for locating regulatory regions in genomic DNA.

In addition to genes, chromosomal DNA contains sequences that serve as signals for turning on and off gene expression. These signals are thought to be distributed as clusters in the regulatory regions of genes. We develop a Bayesian model that views locating regulatory regions in genomic DNA as a change-point problem, with the beginning of regulatory and non-regulatory regions corresponding to the change points. The model is based on a hidden Markov chain. The data consist of nucleotide positions of protein-binding elements in a genomic DNA sequence. These positions are identified using a reference catalogue containing elements that interact with transcription factors implicated in controlling the expression of protein-encoding genes. Among the protein-binding elements in a genomic DNA sequence, the statistical model automatically selects those that tend to predict regulatory regions. We test the model using viral sequences that include known regulatory regions and provide the results obtained for human genomic DNA corresponding to the beta globin locus on chromosome 11.

Adenoviridae↗

2D observers for human 3D object recognition?

In human object recognition, converging evidence has shown that subjects' performance depends on their familiarity with an object's appearance. The extent of such dependence is a function of the inter-object similarity. The more similar the objects are, the stronger this dependence will be and the more dominant the two-dimensional (2D) image-based information will be. However, the degree to which three-dimensional (3D) model-based information is used remains an area of strong debate. Previously the authors showed that all models with independent 2D templates that allowed 2D rotations in the image plane cannot account for human performance in discriminating novel object views. Here the authors derive an analytic formulation of a Bayesian model that gives rise to the best possible performance under 2D affine transformations and demonstrate that this model cannot account for human performance in 3D object discrimination. Relative to this model, human statistical efficiency is higher for novel views than for learned views, suggesting that human observers have used some 3D structural information.

Computer Simulation↗

The analysis of repeated-measures data on schizophrenic reaction times using mixture models.

Reaction times for schizophrenic individuals in a simple visual tracking experiment can be substantially more variable than for non-schizophrenic individuals. Current psychological theory suggests that at least some of this extra variability arises from an attentional lapse that delays some, but not all, of each schizophrenic's reaction times. Based on this theory, we pursue models in which measurements from non-schizophrenics arise from a normal linear model with a separate mean for each individual, whereas measurements from schizophrenics arise from a mixture of (i) a component analogous to the distribution of response times for non-schizophrenics and (ii) a mean-shifted component. We fit four mixture models within this framework, where the distinctions between models arise from assumptions about the variance of the shifted observations and the exchangeability of schizophrenic individuals. Some of these models can be fit by maximum likelihood using the EM algorithm, and all can be fit using the ECM algorithm, where the covariance matrices associated with the parameters are calculated by the SEM and SECM algorithms, respectively. Bayesian model monitoring using posterior predictive checks is invoked to discard models that fail to reproduce certain observed features of the data and to stimulate the development of better models.

Algorithms↗

Three-state frailty model for age at onset of dementia and death in Swedish twins.

We present a frailty model to estimate the relative importance of genetic and environmental factors on age at onset of dementia in a twin design. We use modern survival methodology to define a model that accounts simultaneously for longitudinal aspects, e.g., left truncation and right censoring in data, and the multivariate nature of twin data. Additionally, we present a novel three-state frailty model, with nondemented, demented, and dead states, describing variation in the onset of disease and mortality simultaneously in one model, while accounting for possible dependence for the two competing events. The frailty structure, i.e., the latent random effects structure, mimics the traditional twin model for continuous variables used in quantitative genetics, and as such describes within-pair dependence. This in turn leads to estimates for intrapair correlations, as well as for additive genetic, and shared and nonshared environmental components of variance. A hierarchical Bayesian model formulation and Gibbs sampling are used to estimate posterior distributions of the parameters. The models are applied to Swedish Twin Registry data on the onset of dementia in elderly twins. Based on the three-state frailty model, we estimate the intrapair correlations for dementia to be 0.87 [90% credible interval: 0.61,0.98] and 0.68[0.18,0.91] for MZ and DZ twins, respectively. Based on our model, we estimate that genetic effects account for about one third, and shared environmental effects for almost a half, of the variation in dementia hazards between individuals. More data, however, are needed to gain precision in these estimates.

Age of Onset↗

Modelling the random effects covariance matrix in longitudinal data.

A common class of models for longitudinal data are random effects (mixed) models. In these models, the random effects covariance matrix is typically assumed constant across subject. However, in many situations this matrix may differ by measured covariates. In this paper, we propose an approach to model the random effects covariance matrix by using a special Cholesky decomposition of the matrix. In particular, we will allow the parameters that result from this decomposition to depend on subject-specific covariates and also explore ways to parsimoniously model these parameters. An advantage of this parameterization is that there is no concern about the positive definiteness of the resulting estimator of the covariance matrix. In addition, the parameters resulting from this decomposition have a sensible interpretation. We propose fully Bayesian modelling for which a simple Gibbs sampler can be implemented to sample from the posterior distribution of the parameters. We illustrate these models on data from depression studies and examine the impact of heterogeneity in the covariance matrix on estimation of both fixed and random effects.

Antidepressive Agents↗

Triage protein fold prediction.

We have constructed, in a completely automated fashion, a new structure template library for threading that represents 358 distinct SCOP folds where each model is mathematically represented as a Hidden Markov model (HMM). Because the large number of models in the library can potentially dilute the prediction measure, a new triage method for fold prediction is employed. In the first step of the triage method, the most probable structural class is predicted using a set of manually constructed, high-level, generalized structural HMMs that represent seven general protein structural classes: all-alpha, all-beta, alpha/beta, alpha+beta, irregular small metal-binding, transmembrane beta-barrel, and transmembrane alpha-helical. In the second step, only those fold models belonging to the determined structural class are selected for the final fold prediction. This triage method gave more predictions as well as more correct predictions compared with a simple prediction method that lacks the initial classification step. Two different schemes of assigning Bayesian model priors are presented and discussed.

Animals↗

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3·8 × 10-10 to 2·4 × 10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7·87 million to $287 000) and 29-fold (from $1·98 million to $69 100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0·01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans↗

Model choice in gene mapping: what and why.

The choice of an appropriate genetic model describing the genetic architecture underlying a character of interest is an inherent part of the gene mapping studies of human and other living organisms. The genetic model specifies the statistical parameters for the number of genes, their positions, and the types and magnitudes of their contributions to the phenotype. There are many considerations involved in model formulation (choice) ranging from the assumptions concerning the data, the role of environment, and the number of oligogenes (or quantitative trait loci) influencing the trait behavior. There are several model selection procedures and criteria under specific sampling designs in the genetic literature. These approaches often have their origin in computer science or in general statistical theory. Our aim here is to give an overview of the most popular statistical criteria and to present principles behind them. Bayesian model averaging is suggested as a robust alternative for such methods.

Animals↗

Modeling human fertility in the presence of measurement error.

The probability of conception in a given menstrual cycle is closely related to the timing of intercourse relative to ovulation. Although commonly used markers of time of ovulation are known to be error prone, most fertility models assume the day of ovulation is measured without error. We develop a mixture model that allows the day to be misspecified. We assume that the measurement errors are i.i.d. across menstrual cycles. Heterogeneity among couples in the per cycle likelihood of conception is accounted for using a beta mixture model. Bayesian estimation is straightforward using Markov chain Monte Carlo techniques. The methods are applied to a prospective study of couples at risk of pregnancy. In the absence of validation data or multiple independent markers of ovulation, the identifiability of the measurement error distribution depends on the assumed model. Thus, the results of studies relating the timing of intercourse to the probability of conception should be interpreted cautiously.

Algorithms↗

Bayesian nonstationary autoregressive models for biomedical signal analysis.

We describe a variational Bayesian algorithm for the estimation of a multivariate autoregressive model with time-varying coefficients that adapt according to a linear dynamical system. The algorithm allows for time and frequency domain characterization of nonstationary multivariate signals and is especially suited to the analysis of event-related data. Results are presented on synthetic data and real electroencephalogram data recorded in event-related desynchronization and photic synchronization scenarios.

Algorithms↗

Bayesian synthesis for quantifying uncertainty in predictions from process models.

The Bayesian synthesis method is reviewed and judged to be useful for determining posterior distributions and interval estimates for inputs and outputs of process-based forest models. The method furnishes posterior distributions of the values of a model's parameters and response variables. The method also provides estimates of correlation among the parameters and output variables. Bayesian synthesis is the only type of uncertainty analysis that affords incorporation of all the information available to the investigator, in addition to the information contained in the model itself.

Journal Article↗

Pharmacokinetics of vancomycin in adult cystic fibrosis patients.

Although the depositions of many antibiotics are altered in cystic fibrosis patients, that of vancomycin has not been studied. To assess vancomycin pharmacokinetics, 10 adult cystic fibrosis patients were given a parenteral dose of vancomycin (15 mg/kg) during the first 72 h of hospitalization for acute bronchopulmonary exacerbation. Blood samples were obtained at 0, 1, 1.25, 1.5, 2, 3, 4, 6, 8, 12, 15, and 24 h. The mean (standard deviation) weight, measured creatinine clearance, and Taussig clinical score were 51 (13) kg, 130 (39) ml/min/1.73 m2, and 64 (13), respectively. Multicompartmental pharmacokinetic parameters were best described by a two-compartment model. The mean (standard deviation) volume of distribution, total body clearance, and terminal elimination rate constant were 0.58 (0.15) liter/kg, 91 (19) ml/min/1.73 m2, and 0.123 (0.05) h-1, respectively. These values were consistent with vancomycin pharmacokinetic parameters obtained in previous studies of healthy adult volunteers. Vancomycin dosages predicted by using a two-compartment Bayesian model were approximately 15 mg/kg every 8 to 12 h. There were poor correlations between clinical score or creatinine clearance and any pharmacokinetic parameter (r values of < 0.32). The coefficient of correlation between urine flow rate and total body clearance was 0.7 (P < 0.05). Adult cystic fibrosis patients exhibit a disposition of vancomycin similar to that exhibited by healthy adults, and thus cystic fibrosis does not alter vancomycin pharmacokinetics.

Adult↗

Commentary: practical advantages of Bayesian analysis of epidemiologic data.

In the past decade, there have been enormous advances in the use of Bayesian methodology for analysis of epidemiologic data, and there are now many practical advantages to the Bayesian approach. Bayesian models can easily accommodate unobserved variables such as an individual's true disease status in the presence of diagnostic error. The use of prior probability distributions represents a powerful mechanism for incorporating information from previous studies and for controlling confounding. Posterior probabilities can be used as easily interpretable alternatives to p values. Recent developments in Markov chain Monte Carlo methodology facilitate the implementation of Bayesian analyses of complex data sets containing missing observations and multidimensional outcomes. Tools are now available that allow epidemiologists to take advantage of this powerful approach to assessment of exposure-disease relations.

Bayes Theorem↗

Background estimation in experimental spectra

A general probabilistic technique for estimating background contributions to measured spectra is presented. A Bayesian model is used to capture the defining characteristics of the problem, namely, that the background is smoother than the signal. The signal is allowed to have positive and/or negative components. The background is represented in terms of a cubic spline basis. A variable degree of smoothness of the background is attained by allowing the number of knots and the knot positions to be adaptively chosen on the basis of the data. The fully Bayesian approach taken provides a natural way to handle knot adaptivity and allows uncertainties in the background to be estimated. Our technique is demonstrated on a particle induced x-ray emission spectrum from a geological sample and an Auger spectrum from iron, which contains signals with both positive and negative components.

Journal Article↗

BTS: a scalable Bayesian Tissue Score for prioritizing GWAS variants and their functional contexts across >1000s of omics datasets.

MOTIVATION: statistics from genome-wide association studies (GWAS) are widely used in fine-mapping and colocalization analyses to identify causal variants and their enrichment in functional contexts, such as affected cell types and genomic features. With the expansion of functional genomic (FG) datasets, which now include hundreds of thousands of tracks across various cell and tissue types, it is critical to establish scalable algorithms integrating thousands of diverse FG annotations with GWAS results. RESULTS: We propose BTS (Bayesian Tissue Score), a novel, highly efficient algorithm uniquely designed for (i) identifying affected cell types and functional elements (context-mapping) and (ii) fine-mapping potentially causal variants in a context-specific manner using large collections of cell type-specific FG annotation tracks. BTS leverages GWAS summary statistics and annotation-specific Bayesian models to analyze genome-wide annotation tracks, including enhancers, open chromatin, and histone marks. We evaluated BTS on GWAS summary statistics for immune and cardiovascular traits, such as Inflammatory Bowel Disease (IBD), Rheumatoid Arthritis (RA), Systemic Lupus Erythematosus (SLE), and Coronary Artery Disease (CAD). Our results demonstrate that BTS is over 100&#xd7; more efficient in estimating functional annotation effects and context-specific variant fine-mapping compared to existing methods. Importantly, this large-scale Bayesian approach prioritizes both known and novel annotations, cell types, genomic regions, and variants and provides valuable biological insights into the functional contexts of these diseases. AVAILABILITY AND IMPLEMENTATION: Docker image is available at https://hub.docker.com/r/wanglab/bts with preinstalled BTS R package (https://bitbucket.org/wanglab-upenn/BTS-R) and BTS GWAS summary statistics analysis pipeline (https://bitbucket.org/wanglab-upenn/bts-pipeline).

Genome-Wide Association Study↗

An illustration of the modelling of cost and efficacy data from a clinical trial.

Health care providers, purchasers and policy makers need to make informed decisions regarding the provision of cost-effective care. When a new health care intervention is to be compared with the current standard, an economic evaluation alongside an evaluation of health benefits provides useful information for the decision making process. We consider the information on cost-effectiveness which arises from an individual clinical trial comparing the two interventions. Recent methods for conducting a cost-effectiveness analysis for a clinical trial have focused on the net benefit parameter. The net benefit parameter, a function of costs and health benefits, is positive if the new intervention is cost-effective compared with the standard. In this paper we describe frequentist and Bayesian approaches to cost-effectiveness analysis which have been suggested in the literature and apply them to data from a clinical trial comparing laparoscopic surgery with open mesh surgery for the repair of inguinal hernias. We extend the Bayesian model to allow the total cost to be divided into a number of different components. The advantages and disadvantages of the different approaches are discussed. In January 2001, NICE issued guidance on the type of surgery to be used for inguinal hernia repair. We discuss our example in the light of this information.

Bayes Theorem↗

Cerebrospinal fluid pharmacokinetics and penetration of continuous infusion topotecan in children with central nervous system tumors.

The purpose of this study was to describe the cerebrospinal fluid (CSF) penetration of topotecan in humans, to generate a pharmacokinetic model to simultaneously describe topotecan lactone and total concentrations in the plasma and CSF, and to characterize the CSF and plasma pharmacokinetics of topotecan administered as a continuous infusion (CI). Plasma and CSF samples were collected from 17 patients receiving 5.5 or 7.5 mg/m2 per day as a 24-h CI (5 patients, 7 courses), or 0.5 to 1.25 mg/m2 per day as a 72-h CI (12 patients, 12 courses). CSF samples were obtained from either a ventricular reservoir (VR) or a lumbar puncture (LP). Topotecan lactone and total (lactone plus hydroxy acid) concentrations were determined by HPLC and fluorescence detection. Using MAP-Bayesian modelling, a three-compartment model was fitted simultaneously to topotecan lactone and total concentrations in the plasma and CSF. The penetration of topotecan into the CSF was determined from the ratio of the CSF to the plasma area under the concentration-time curve. The median CSF ventricular lactone concentrations, obtained prior to the end of infusion (EOI), were 0.86, 1.4, 0.73, 5.3, and 4.6 ng/ml for patients receiving 0.5, 1.0, 1.25, 5.5, and 7.5 mg/m2 per day, respectively. EOI CSF lumbar lactone concentrations measured in three patients were 0.44, 1.1, and 1.7 ng/ml for topotecan doses of 1.0, 5.5, and 7.5 mg/m2 per day, respectively. In two patients receiving 1.25 mg/m2 per day, EOI CSF concentrations were obtained simultaneously from a VR and LP; the lumbar lactone concentrations were 30% and 49% lower than the ventricular concentrations. During a 24-h and a 72-h CI, the median CSF penetration of topotecan lactone was 0.29 (range 0.10 to 0.59) and 0.42 (range 0.11 to 0.86), respectively. A three-compartment model adequately described topotecan lactone and total concentrations in the plasma and CSF. Topotecan was therefore found to significantly penetrate into the CSF in humans. The pharmacokinetic model presented may be useful in the design of clinical studies of topotecan to treat CNS tumors.

Adolescent↗