PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Bayesian modelling”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 397 records · Page 22Linked to original sources

2D observers for human 3D object recognition?

In human object recognition, converging evidence has shown that subjects' performance depends on their familiarity with an object's appearance. The extent of such dependence is a function of the inter-object similarity. The more similar the objects are, the stronger this dependence will be and the more dominant the two-dimensional (2D) image-based information will be. However, the degree to which three-dimensional (3D) model-based information is used remains an area of strong debate. Previously the authors showed that all models with independent 2D templates that allowed 2D rotations in the image plane cannot account for human performance in discriminating novel object views. Here the authors derive an analytic formulation of a Bayesian model that gives rise to the best possible performance under 2D affine transformations and demonstrate that this model cannot account for human performance in 3D object discrimination. Relative to this model, human statistical efficiency is higher for novel views than for learned views, suggesting that human observers have used some 3D structural information.

Computer Simulation↗

Space-time models with time-dependent covariates for the analysis of the temporal lag between socioeconomic factors and lung cancer mortality.

The relationship between socioeconomic factors and mortality for lung cancer is investigated. To identify the proper lag time between socioeconomic factors and lung cancer mortality, a space-time hierarchical Bayesian model with time-dependent covariates is adopted. A real example on lung cancer mortality, males, in Tuscany (Italy) during the period 1971-1999, is provided. Results confirm the presence of an association between mortality for lung cancer and socioeconomic factors with a temporal lag (latency time) of at least 10 years.

Adolescent↗

[Statistical models for spatial analysis in parasitology].

The simplest way to study the spatial pattern of a disease is the geographical representation of its cases (or some indicators of them) over a map. Maps based on raw data are generally "wrong" since they do not take into consideration for sampling errors. Indeed, the observed differences between areas (or points in the map) are not directly interpretable, as they derive from the composition of true, structural differences and of the noise deriving from the sampling process. This problem is well known in human epidemiology, and several solutions have been proposed to filter the signal from the noise. These statistical methods are usually referred to as Disease Mapping. In geographical analysis a first goal is to evaluate the statistical significance of the heterogeneity between areas (or points). If the test indicates rejection of the hypothesis of homogeneity the following task is to study the spatial pattern of the disease. The spatial variability of risk is usually decomposed into two terms: a spatially structured (clustering) and a non spatially structured (heterogeneity) one. The heterogeneity term reflects spatial variability due to intrinsic characteristics of the sampling units (e.g. igienic conditions of farms), while the clustering term models the association due to proximity between sampling units, that usually depends on ecological conditions that vary over the study area and that affect in similar way breedings that are close to each other. Hierarchical bayesian models are the main tool to make inference over the clustering and heterogeneity components. The results are based on the marginal posterior distributions of the parameters of the model, that are approximated by Monte Carlo Markov Chain methods. Different models can be defined depending on the terms that are considered, namely a model with only the clustering term, a model with only the heterogeneity term and a model where both are included. Model selection criteria based on a compromise between degree of complexity and goodness of fit are then needed to discriminate among them, because each specification has a different biological meaning. Our aim is to demonstrate that these techniques can be used to study the geographical distribution of a parasite infection. Our analyses are based on data collected in 142 farms of the province of Latina. In each breeding a fixed number of sheeps has been sampled (20) and checked for the presence of C. daubneyi. We have specified a Binomial model for the proportion of infected animals in each breeding. The heterogeneity component is modelled in a standard way, while we have used different prior specifications for the clustering term to show how they affect the results. When we use the usual specification also for clustering, the two models show a completely different spatial pattern of infection, probably because the intrinsic spatial structure of the clustering term tend to bias our inferences. The selection criterion indicates in this case the heterogeneity model as the "best" one. However, if we modify the prior so that a lower degree of spatial interaction is assumed, the clustering model is less complex and its goodness of fit better and it should be preferred.

Animal Husbandry↗

The analysis of repeated-measures data on schizophrenic reaction times using mixture models.

Reaction times for schizophrenic individuals in a simple visual tracking experiment can be substantially more variable than for non-schizophrenic individuals. Current psychological theory suggests that at least some of this extra variability arises from an attentional lapse that delays some, but not all, of each schizophrenic's reaction times. Based on this theory, we pursue models in which measurements from non-schizophrenics arise from a normal linear model with a separate mean for each individual, whereas measurements from schizophrenics arise from a mixture of (i) a component analogous to the distribution of response times for non-schizophrenics and (ii) a mean-shifted component. We fit four mixture models within this framework, where the distinctions between models arise from assumptions about the variance of the shifted observations and the exchangeability of schizophrenic individuals. Some of these models can be fit by maximum likelihood using the EM algorithm, and all can be fit using the ECM algorithm, where the covariance matrices associated with the parameters are calculated by the SEM and SECM algorithms, respectively. Bayesian model monitoring using posterior predictive checks is invoked to discard models that fail to reproduce certain observed features of the data and to stimulate the development of better models.

Algorithms↗

Three-state frailty model for age at onset of dementia and death in Swedish twins.

We present a frailty model to estimate the relative importance of genetic and environmental factors on age at onset of dementia in a twin design. We use modern survival methodology to define a model that accounts simultaneously for longitudinal aspects, e.g., left truncation and right censoring in data, and the multivariate nature of twin data. Additionally, we present a novel three-state frailty model, with nondemented, demented, and dead states, describing variation in the onset of disease and mortality simultaneously in one model, while accounting for possible dependence for the two competing events. The frailty structure, i.e., the latent random effects structure, mimics the traditional twin model for continuous variables used in quantitative genetics, and as such describes within-pair dependence. This in turn leads to estimates for intrapair correlations, as well as for additive genetic, and shared and nonshared environmental components of variance. A hierarchical Bayesian model formulation and Gibbs sampling are used to estimate posterior distributions of the parameters. The models are applied to Swedish Twin Registry data on the onset of dementia in elderly twins. Based on the three-state frailty model, we estimate the intrapair correlations for dementia to be 0.87 [90% credible interval: 0.61,0.98] and 0.68[0.18,0.91] for MZ and DZ twins, respectively. Based on our model, we estimate that genetic effects account for about one third, and shared environmental effects for almost a half, of the variation in dementia hazards between individuals. More data, however, are needed to gain precision in these estimates.

Age of Onset↗

Modelling the random effects covariance matrix in longitudinal data.

A common class of models for longitudinal data are random effects (mixed) models. In these models, the random effects covariance matrix is typically assumed constant across subject. However, in many situations this matrix may differ by measured covariates. In this paper, we propose an approach to model the random effects covariance matrix by using a special Cholesky decomposition of the matrix. In particular, we will allow the parameters that result from this decomposition to depend on subject-specific covariates and also explore ways to parsimoniously model these parameters. An advantage of this parameterization is that there is no concern about the positive definiteness of the resulting estimator of the covariance matrix. In addition, the parameters resulting from this decomposition have a sensible interpretation. We propose fully Bayesian modelling for which a simple Gibbs sampler can be implemented to sample from the posterior distribution of the parameters. We illustrate these models on data from depression studies and examine the impact of heterogeneity in the covariance matrix on estimation of both fixed and random effects.

Antidepressive Agents↗

Triage protein fold prediction.

We have constructed, in a completely automated fashion, a new structure template library for threading that represents 358 distinct SCOP folds where each model is mathematically represented as a Hidden Markov model (HMM). Because the large number of models in the library can potentially dilute the prediction measure, a new triage method for fold prediction is employed. In the first step of the triage method, the most probable structural class is predicted using a set of manually constructed, high-level, generalized structural HMMs that represent seven general protein structural classes: all-alpha, all-beta, alpha/beta, alpha+beta, irregular small metal-binding, transmembrane beta-barrel, and transmembrane alpha-helical. In the second step, only those fold models belonging to the determined structural class are selected for the final fold prediction. This triage method gave more predictions as well as more correct predictions compared with a simple prediction method that lacks the initial classification step. Two different schemes of assigning Bayesian model priors are presented and discussed.

Animals↗

Bayesian variable selection for gene expression modeling with regulatory motif binding sites in neuroinflammatory events.

Multiple transcription factors (TFs) coordinately control transcriptional regulation of genes in eukaryotes. Although numerous computational methods focus on the identification of individual TF-binding sites (TFBSs), very few consider the interdependence among these sites. In this article, we studied the relationship between TFBSs and microarray gene expression levels using both family-wise and memberspecific motifs, under various combination of regression models with Bayesian variable selection, as well as motif scoring and sharing conditions, in order to account for the coordination complexity of transcription regulation. We proposed a three-step approach to model the relationship. In the first step, we preprocessed microarray data and used p-values and expression ratios to preselect upregulated and downregulated genes. The second step aimed to identify and score individual TFBSs within DNA sequence of each gene. A method based on the degree of similarity and the number of TFBSs was employed to calculate the score of each TFBS in each gene sequence. In the last step, linear regression and probit regression were used to build a predictive model of gene expression outcomes using these TFBSs as predictors. Given a certain number of predictors to be used, a full search of all possible predictor sets is usually combinatorially prohibitive. Therefore, this article considered the Bayesian variable selection for prediction using either of the regression models. The Bayesian variable selection has been applied in the context of gene selection, missing value estimation, and regulatory motif identification. In our modeling, the regressor was approximated as a linear combination of the TFBSs and a Gibbs sampler was employed to find the strongest TFBSs. We applied these regression models with the Bayesian variable selection on spinal cord injury gene expression data set. These TFs demonstrated intricate regulatory roles either as a family or as individual members in neuroinflammatory events. Our analysis can be applied to create plausible hypotheses for combinatorial regulation by TFBSs and avoiding false-positive candidates in the modeling process at the same time. Such a systematic approach provides the possibility to dissect transcription regulation, from a more comprehensive perspective, through which phenotypical events at cellular and tissue levels are moved forward by molecular events at gene transcription and translation levels.

Amino Acid Motifs↗

Analysis of HIV-1 pol sequences using Bayesian Networks: implications for drug resistance.

Human Immunodeficiency Virus-1 (HIV-1) antiviral resistance is a major cause of antiviral therapy failure and compromises future treatment options. As a consequence, resistance testing is the standard of care. Because of the high degree of HIV-1 natural variation and complex interactions, the role of resistance mutations is in many cases insufficiently understood. We applied a probabilistic model, Bayesian networks, to analyze direct influences between protein residues and exposure to treatment in clinical HIV-1 protease sequences from diverse subtypes. We can determine the specific role of many resistance mutations against the protease inhibitor nelfinavir, and determine relationships between resistance mutations and polymorphisms. We can show for example that in addition to the well-known major mutations 90M and 30N for nelfinavir resistance, 88S should not be treated as 88D but instead considered as a major mutation and explain the subtype-dependent prevalence of the 30N resistance pathway.

Amino Acid Sequence↗

Modeling of under-detection of cases in disease surveillance.

PURPOSE: Accurate epidemiological surveillance of leprosy is a matter of international public health concern. It often suffers, however, from potential problems of under-registration of reported cases, particularly in poorer and more socially deprived areas. Such problems also apply in the surveillance of many other communicable or transmissible diseases. We develop a Bayesian model for small-area disease rates that allows for censoring of case detection in suspect districts and can therefore be used to estimate under-reporting of cases in a given study region. METHODS: Such methods are applied to leprosy incidence in a municipality of Pernambuco State in North Eastern Brazil, using a social deprivation indicator as the basis for considering data from certain districts to be censored. The time period we consider was immediately prior to an extension of the coverage and efficacy of the control program and model predictions concerning under reporting can therefore be compared with more reliable data subsequently collected from the same region. RESULTS: The proposed method produces informative estimates of under detection of leprosy cases in the defined study region and these estimates compare well, both in size and in geographical location, with the numbers of cases subsequently detected. CONCLUSIONS: As illustrated by the application discussed in this article, the proposed model provides a general tool that may be used in spatial epidemiological surveillance situations where the available data is suspected to contain significant under-registrations of cases in certain geographical areas.

Bayes Theorem↗

Inferring the sensitivity of wastewater metagenomic sequencing for early detection of viruses: a statistical modelling study.

BACKGROUND: Metagenomic sequencing of wastewater (W-MGS) can in principle detect any known or novel pathogen in a population. We aimed to quantify the sensitivity and cost of W-MGS for viral pathogen detection by jointly analysing W-MGS and epidemiological data for a range of human-infecting viruses. METHODS: In this statistical modelling study, we analysed sequencing data from four studies of untargeted W-MGS to estimate the relative abundance of 11 human-infecting viruses. Corresponding prevalence and incidence estimates were obtained or calculated from academic and public health reports. We combined these estimates using a hierarchical Bayesian model to predict relative abundance at set prevalence or incidence values, allowing comparison across studies and viruses. These predictions were then used to estimate the sequencing depth and concomitant cost required for pathogen detection using W-MGS with or without use of a hybridisation capture enrichment panel. FINDINGS: After controlling for variation in local infection rates, relative abundance varied by orders of magnitude across studies for a given virus. For instance, a local SARS-CoV-2 weekly incidence of 1% corresponded to a predicted SARS-CoV-2 relative abundance ranging from 3·8 × 10-10 to 2·4 × 10-7 across studies, translating to orders-of-magnitude variation in the cost of operating a system able to detect a SARS-CoV-2-like pathogen at a given sensitivity. Use of a respiratory virus enrichment panel in two studies greatly increased predicted relative abundance of SARS-CoV-2, lowering yearly costs by 27-fold (from US$7·87 million to $287 000) and 29-fold (from $1·98 million to $69 100) for a system able to detect a SARS-CoV-2-like pathogen before reaching 0·01% cumulative incidence. INTERPRETATION: The large variation in viral relative abundance after controlling for epidemiological factors indicates that other sources of inter-study variation, such as differences in sewershed hydrology and laboratory protocols, have a substantial impact on the sensitivity and cost of W-MGS. Well chosen hybridisation capture panels can greatly increase sensitivity and reduce cost for viruses in the panel, but might reduce sensitivity to unknown or unexpected pathogens. FUNDING: The Wellcome Trust, Open Philanthropy, and Musk Foundation.

Humans↗

Model choice in gene mapping: what and why.

The choice of an appropriate genetic model describing the genetic architecture underlying a character of interest is an inherent part of the gene mapping studies of human and other living organisms. The genetic model specifies the statistical parameters for the number of genes, their positions, and the types and magnitudes of their contributions to the phenotype. There are many considerations involved in model formulation (choice) ranging from the assumptions concerning the data, the role of environment, and the number of oligogenes (or quantitative trait loci) influencing the trait behavior. There are several model selection procedures and criteria under specific sampling designs in the genetic literature. These approaches often have their origin in computer science or in general statistical theory. Our aim here is to give an overview of the most popular statistical criteria and to present principles behind them. Bayesian model averaging is suggested as a robust alternative for such methods.

Animals↗

Modeling human fertility in the presence of measurement error.

The probability of conception in a given menstrual cycle is closely related to the timing of intercourse relative to ovulation. Although commonly used markers of time of ovulation are known to be error prone, most fertility models assume the day of ovulation is measured without error. We develop a mixture model that allows the day to be misspecified. We assume that the measurement errors are i.i.d. across menstrual cycles. Heterogeneity among couples in the per cycle likelihood of conception is accounted for using a beta mixture model. Bayesian estimation is straightforward using Markov chain Monte Carlo techniques. The methods are applied to a prospective study of couples at risk of pregnancy. In the absence of validation data or multiple independent markers of ovulation, the identifiability of the measurement error distribution depends on the assumed model. Thus, the results of studies relating the timing of intercourse to the probability of conception should be interpreted cautiously.

Algorithms↗

Bayesian nonstationary autoregressive models for biomedical signal analysis.

We describe a variational Bayesian algorithm for the estimation of a multivariate autoregressive model with time-varying coefficients that adapt according to a linear dynamical system. The algorithm allows for time and frequency domain characterization of nonstationary multivariate signals and is especially suited to the analysis of event-related data. Results are presented on synthetic data and real electroencephalogram data recorded in event-related desynchronization and photic synchronization scenarios.

Algorithms↗

Analyzing gene expression time-courses.

Measuring gene expression over time can provide important insights into basic cellular processes. Identifying groups of genes with similar expression time-courses is a crucial first step in the analysis. As biologically relevant groups frequently overlap, due to genes having several distinct roles in those cellular processes, this is a difficult problem for classical clustering methods. We use a mixture model to circumvent this principal problem, with hidden Markov models (HMMs) as effective and flexible components. We show that the ensuing estimation problem can be addressed with additional labeled data-partially supervised learning of mixtures-through a modification of the Expectation-Maximization (EM) algorithm. Good starting points for the mixture estimation are obtained through a modification to Bayesian model merging, which allows us to learn a collection of initial HMMs. We infer groups from mixtures with a simple information-theoretic decoding heuristic, which quantifies the level of ambiguity in group assignment. The effectiveness is shown with high-quality annotation data. As the HMMs we propose capture asynchronous behavior by design, the groups we find are also asynchronous. Synchronous subgroups are obtained from a novel algorithm based on Viterbi paths. We show the suitability of our HMM mixture approach on biological and simulated data and through the favorable comparison with previous approaches. A software implementing the method is freely available under the GPL from http://ghmm.org/gql.

Algorithms↗

Differentiation among populations with migration, mutation, and drift: implications for genetic inference.

Populations may become differentiated from one another as a result of genetic drift. The amounts and patterns of differentiation at neutral loci are determined by local population sizes, migration rates among populations, and mutation rates. We provide exact analytical expressions for the mean, variance, and covariance of a stochastic model for hierarchically structured populations subject to migration, mutation, and drift. In addition to the expected correlation in allele frequencies among populations in the same geographic region, we demonstrate that there is a substantial correlation in allele frequencies among regions at the top level of the hierarchy. We propose a hierarchical Bayesian model for inference of Wright's F-statistics in a two-level hierarchy in which we estimate the among-region correlation in allele frequencies by substituting replication across loci for replication across time. We illustrate the approach through an analysis of human microsatellite data, and we show that approaches ignoring the among-region correlation in allele frequencies underestimate the amount of genetic differentiation among major geographic population groups by approximately 30%. Finally, we discuss the implications of these results for the use and interpretation of F-statistics in evolutionary studies.

Animal Migration↗

Bayesian synthesis for quantifying uncertainty in predictions from process models.

The Bayesian synthesis method is reviewed and judged to be useful for determining posterior distributions and interval estimates for inputs and outputs of process-based forest models. The method furnishes posterior distributions of the values of a model's parameters and response variables. The method also provides estimates of correlation among the parameters and output variables. Bayesian synthesis is the only type of uncertainty analysis that affords incorporation of all the information available to the investigator, in addition to the information contained in the model itself.

Journal Article↗

Pharmacokinetics of vancomycin in adult cystic fibrosis patients.

Although the depositions of many antibiotics are altered in cystic fibrosis patients, that of vancomycin has not been studied. To assess vancomycin pharmacokinetics, 10 adult cystic fibrosis patients were given a parenteral dose of vancomycin (15 mg/kg) during the first 72 h of hospitalization for acute bronchopulmonary exacerbation. Blood samples were obtained at 0, 1, 1.25, 1.5, 2, 3, 4, 6, 8, 12, 15, and 24 h. The mean (standard deviation) weight, measured creatinine clearance, and Taussig clinical score were 51 (13) kg, 130 (39) ml/min/1.73 m2, and 64 (13), respectively. Multicompartmental pharmacokinetic parameters were best described by a two-compartment model. The mean (standard deviation) volume of distribution, total body clearance, and terminal elimination rate constant were 0.58 (0.15) liter/kg, 91 (19) ml/min/1.73 m2, and 0.123 (0.05) h-1, respectively. These values were consistent with vancomycin pharmacokinetic parameters obtained in previous studies of healthy adult volunteers. Vancomycin dosages predicted by using a two-compartment Bayesian model were approximately 15 mg/kg every 8 to 12 h. There were poor correlations between clinical score or creatinine clearance and any pharmacokinetic parameter (r values of < 0.32). The coefficient of correlation between urine flow rate and total body clearance was 0.7 (P < 0.05). Adult cystic fibrosis patients exhibit a disposition of vancomycin similar to that exhibited by healthy adults, and thus cystic fibrosis does not alter vancomycin pharmacokinetics.

Adult↗