PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Markov Chain”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Searching for convergence in phylogenetic Markov chain Monte Carlo.

Markov chain Monte Carlo (MCMC) is a methodology that is gaining widespread use in the phylogenetics community and is central to phylogenetic software packages such as MrBayes. An important issue for users of MCMC methods is how to select appropriate values for adjustable parameters such as the length of the Markov chain or chains, the sampling density, the proposal mechanism, and, if Metropolis-coupled MCMC is being used, the number of heated chains and their temperatures. Although some parameter settings have been examined in detail in the literature, others are frequently chosen with more regard to computational time or personal experience with other data sets. Such choices may lead to inadequate sampling of tree space or an inefficient use of computational resources. We performed a detailed study of convergence and mixing for 70 randomly selected, putatively orthologous protein sets with different sizes and taxonomic compositions. Replicated runs from multiple random starting points permit a more rigorous assessment of convergence, and we developed two novel statistics, delta and epsilon, for this purpose. Although likelihood values invariably stabilized quickly, adequate sampling of the posterior distribution of tree topologies took considerably longer. Our results suggest that multimodality is common for data sets with 30 or more taxa and that this results in slow convergence and mixing. However, we also found that the pragmatic approach of combining data from several short, replicated runs into a "metachain" to estimate bipartition posterior probabilities provided good approximations, and that such estimates were no worse in approximating a reference posterior distribution than those obtained using a single long run of the same length as the metachain. Precision appears to be best when heated Markov chains have low temperatures, whereas chains with high temperatures appear to sample trees with high posterior probabilities only rarely.

Bayes Theorem↗

LD-SPatt: large deviations statistics for patterns on Markov chains.

Statistics on Markov chains are widely used for the study of patterns in biological sequences. Statistics on these models can be done through several approaches. Central limit theorem (CLT) producing Gaussian approximations are one of the most popular ones. Unfortunately, in order to find a pattern of interest, these methods have to deal with tail distribution events where CLT is especially bad. In this paper, we propose a new approach based on the large deviations theory to assess pattern statistics. We first recall theoretical results for empiric mean (level 1) as well as empiric distribution (level 2) large deviations on Markov chains. Then, we present the applications of these results focusing on numerical issues. LD-SPatt is the name of GPL software implementing these algorithms. We compare this approach to several existing ones in terms of complexity and reliability and show that the large deviations are more reliable than the Gaussian approximations in absolute values as well as in terms of ranking and are at least as reliable as compound Poisson approximations. We then finally discuss some further possible improvements and applications of this new method.

Computational Biology↗

A bayesian approach to detect quantitative trait loci using Markov chain Monte Carlo.

Markov chain Monte Carlo (MCMC) techniques are applied to simultaneously identify multiple quantitative trait loci (QTL) and the magnitude of their effects. Using a Bayesian approach a multi-locus model is fit to quantitative trait and molecular marker data, instead of fitting one locus at a time. The phenotypic trait is modeled as a linear function of the additive and dominance effects of the unknown QTL genotypes. Inference summaries for the locations of the QTL and their effects are derived from the corresponding marginal posterior densities obtained by integrating the likelihood, rather than by optimizing the joint likelihood surface. This is done using MCMC by treating the unknown QTL, genotypes, and any missing marker genotypes, as augmented data and then by including these unknowns in the Markov chain cycle alone with the unknown parameters. Parameter estimates are obtained as means of the corresponding marginal posterior densities. High posterior density regions of the marginal densities are obtained as confidence regions. We examine flowering time data from double haploid progeny of Brassica napus to illustrate the proposed method.

Bayes Theorem↗

Transition records of stationary Markov chains.

In any Markov chain with finite state space the distribution of transition records always belongs to the exponential family. This observation is used to prove a fluctuation theorem, and to show that the dynamical entropy of a stationary Markov chain is linear in the number of steps. Three applications are discussed. A known result about entropy production is reproduced. A thermodynamic relation is derived for equilibrium systems with Metropolis dynamics. Finally, a link is made with recent results concerning a one-dimensional polymer model.

Journal Article↗

Covariate adjustment of event histories estimated from Markov chains: the additive approach.

Markov chain models are frequently used for studying event histories that include transitions between several states. An empirical transition matrix for nonhomogeneous Markov chains has previously been developed, including a detailed statistical theory based on counting processes and martingales. In this article, we show how to estimate transition probabilities dependent on covariates. This technique may, e.g., be used for making estimates of individual prognosis in epidemiological or clinical studies. The covariates are included through nonparametric additive models on the transition intensities of the Markov chain. The additive model allows for estimation of covariate-dependent transition intensities, and again a detailed theory exists based on counting processes. The martingale setting now allows for a very natural combination of the empirical transition matrix and the additive model, resulting in estimates that can be expressed as stochastic integrals, and hence their properties are easily evaluated. Two medical examples will be given. In the first example, we study how the lung cancer mortality of uranium miners depends on smoking and radon exposure. In the second example, we study how the probability of being in response depends on patient group and prophylactic treatment for leukemia patients who have had a bone marrow transplantation. A program in R and S-PLUS that can carry out the analyses described here has been developed and is freely available on the Internet.

Analysis of Variance↗

Characterization of endothelial cell locomotion using a Markov chain model.

A Markov chain model was developed to characterize the two-dimensional locomotion of bovine pulmonary artery endothelial (BPAE) cells cultured with or without basic fibroblast growth factor (bFGF). This model provides a detailed description of the migration process by computing the following locomotory parameters: (i) the speed of cell locomotion; (ii) the expected duration of cell movement in any given direction; (iii) the probability distribution of turn angles that will decide the next direction of cell movement; (iv) the frequency of cell stops; and (v) the duration of cell stops. Eight directional states and a stationary state were used in our Markov analysis. From cell trajectory data, the transition probabilities among the various states and the waiting times for the directional and the stationary states were computed. The steady-state probabilities were also calculated to obtain the ultimate direction of cell motion and, thus, determine whether cell motion was random. Our results showed how the addition of bFGF enhanced the locomotory capability of BPAE cells. Cells cultured with 30 ng/mL bFGF had lower probability of moving to the stationary state than those cultured without bFGF. In addition, cells cultured with 30 ng/mL bFGF remained in the stationary state for shorter periods of time than cells cultured without bFGF. In both these cases, however, the transition probabilities from the stationary state to any directional state were uniformly distributed and were not affected by the presence of bFGF.

Algorithms↗

Simulation of human hypnograms using a Markov chain model.

A Markov chain model has been proposed as a mechanism that generates human sleep stages. A method for estimating the parameters of the model, i.e., the transition probabilities (rates) between sleep stages, has been introduced and applied to 95 hypnograms taken from 23 subjects. The rates characterize interindividual differences and nightly variations of the sleep mechanism, related to sleep-onset behavior, to the decreasing amount of slow wave sleep in the course of the night, and to the REM-NREM periodicity. The model simulates both probabilistic and the above-mentioned predictable dynamics of sleep, but only if these time-varying, individual rates are applied.

Humans↗

Bayesian restoration of a hidden Markov chain with applications to DNA sequencing.

Hidden Markov models (HMMs) are a class of stochastic models that have proven to be powerful tools for the analysis of molecular sequence data. A hidden Markov model can be viewed as a black box that generates sequences of observations. The unobservable internal state of the box is stochastic and is determined by a finite state Markov chain. The observable output is stochastic with distribution determined by the state of the hidden Markov chain. We present a Bayesian solution to the problem of restoring the sequence of states visited by the hidden Markov chain from a given sequence of observed outputs. Our approach is based on a Monte Carlo Markov chain algorithm that allows us to draw samples from the full posterior distribution of the hidden Markov chain paths. The problem of estimating the probability of individual paths and the associated Monte Carlo error of these estimates is addressed. The method is illustrated by considering a problem of DNA sequence multiple alignment. The special structure for the hidden Markov model used in the sequence alignment problem is considered in detail. In conclusion, we discuss certain interesting aspects of biological sequence alignments that become accessible through the Bayesian approach to HMM restoration.

Algorithms↗

Finding noncommunicating sets for Markov chain Monte Carlo estimations on pedigrees.

Markov chain Monte Carlo (MCMC) has recently gained use as a method of estimating required probability and likelihood functions in pedigree analysis, when exact computation is impractical. However, when a multiallelic locus is involved, irreducibility of the constructed Markov chain, an essential requirement of the MCMC method, may fail. Solutions proposed by several researchers, which do not identify all the noncommunicating sets of genotypic configurations, are inefficient with highly polymorphic loci. This is a particularly serious problem in linkage analysis, because highly polymorphic markers are much more informative and thus are preferred. In the present paper, we describe an algorithm that finds all the noncommunicating classes of genotypic configurations on any pedigree. This leads to a more efficient method of defining an irreducible Markov chain. Examples, including a pedigree from a genetic study of familial Alzheimer disease, are used to illustrate how the algorithm works and how penetrances are modified for specific individuals to ensure irreducibility.

Algorithms↗

Distinguishing migration from isolation: a Markov chain Monte Carlo approach.

A Markov chain Monte Carlo method for estimating the relative effects of migration and isolation on genetic diversity in a pair of populations from DNA sequence data is developed and tested using simulations. The two populations are assumed to be descended from a panmictic ancestral population at some time in the past and may (or may not) after that be connected by migration. The use of a Markov chain Monte Carlo method allows the joint estimation of multiple demographic parameters in either a Bayesian or a likelihood framework. The parameters estimated include the migration rate for each population, the time since the two populations diverged from a common ancestral population, and the relative size of each of the two current populations and of the common ancestral population. The results show that even a single nonrecombining genetic locus can provide substantial power to test the hypothesis of no ongoing migration and/or to test models of symmetric migration between the two populations. The use of the method is illustrated in an application to mitochondrial DNA sequence data from a fish species: the threespine stickleback (Gasterosteus aculeatus).

Animals↗

Temporal relation between the ADC and DC potential responses to transient focal ischemia in the rat: a Markov chain Monte Carlo simulation analysis.

Markov chain Monte Carlo simulation was used in a reanalysis of the longitudinal data obtained by Harris et al. (J Cereb Blood Flow Metab 20:28-36) in a study of the direct current (DC) potential and apparent diffusion coefficient (ADC) responses to focal ischemia. The main purpose was to provide a formal analysis of the temporal relationship between the ADC and DC responses, to explore the possible involvement of a common latent (driving) process. A Bayesian nonlinear hierarchical random coefficients model was adopted. DC and ADC transition parameter posterior probability distributions were generated using three parallel Markov chains created using the Metropolis algorithm. Particular attention was paid to the within-subject differences between the DC and ADC time course characteristics. The results show that the DC response is biphasic, whereas the ADC exhibits monophasic behavior, and that the two DC components are each distinguishable from the ADC response in their time dependencies. The DC and ADC changes are not, therefore, driven by a common latent process. This work demonstrates a general analytical approach to the multivariate, longitudinal data-processing problem that commonly arises in stroke and other biomedical research.

Animals↗

Achieving irreducibility of the Markov chain Monte Carlo method applied to pedigree data.

Markov chain Monte Carlo (MCMC) methods have been explored by various researchers as an alternative to exact probability computation in statistical genetics. The objective is to simulate a Markov chain with the desired equilibrium distribution. If the transition kernel is aperiodic and irreducible, then convergence to the equilibrium distribution is guaranteed; realizations of the Markov chain can thus be used to estimate desired probabilities. Aperiodicity is easily satisfied, but, although it has been shown that irreducibility is satisfied for a diallelic locus, reducibility is a potential problem for a multiallelic locus. This is a particularly serious problem in linkage analysis, because multiallelic markers are much more informative than diallelic markers and thus highly preferred. In this paper, the authors propose a new algorithm to achieve irreducibility of the Markov chain of interest by introducing an irreducible auxiliary chain. The irreducibility of the auxiliary chain is obtained by assigning positive probabilities to a small subset of the genotypic configurations inconsistent with the data, to bridge the gap between the irreducible sets.

Algorithms↗

Two-locus modeling of asthma in a Hutterite pedigree via Markov chain Monte Carlo.

Bayesian Markov chain Monte Carlo (MCMC) segregation analysis for asthma was performed on the whole 1,544-member Hutterite pedigree. Heterogeneous and epistatic two-locus models and complex one-locus models were investigated, with trait loci postulated to be linked to markers in regions previously found to be possibly linked to asthma or atopy. The epistatic two-locus dominant-dominant model provided the best estimates, among the models investigated, in terms of prediction of population prevalence and relative risk for sibs of the affected.

Asthma↗

Sibship reconstruction in hierarchical population structures using Markov chain Monte Carlo techniques.

Markov chain Monte Carlo procedures allow the reconstruction of full-sibships using data from genetic marker loci only. In this study, these techniques are extended to allow the reconstruction of nested full- within half-sib families, and to present an efficient method for calculating the likelihood of the observed marker data in a nested family. Simulation is used to examine the properties of the reconstructed sibships, and of estimates of heritability and common environmental variance of quantitative traits obtained from those populations. Accuracy of reconstruction increases with increasing marker information and with increasing size of the nested full-sibships, but decreases with increasing population size. Estimates of variance component are biased, with the direction and magnitude of bias being dependent upon the underlying errors made during pedigree reconstruction.

Analysis of Variance↗

Estimation of the transition matrix of a discrete-time Markov chain.

Discrete-time Markov chains have been successfully used to investigate treatment programs and health care protocols for chronic diseases. In these situations, the transition matrix, which describes the natural progression of the disease, is often estimated from a cohort observed at common intervals. Estimation of the matrix, however, is often complicated by the complex relationship among transition probabilities. This paper summarizes methods to obtain the maximum likelihood estimate of the transition matrix when the cycle length of the model coincides with the observation interval, the cycle length does not coincide with the observation interval, and when the observation intervals are unequal in length. In addition, the bootstrap is discussed as a method to assess the uncertainty of the maximum likelihood estimate and to construct confidence intervals for functions of the transition matrix such as expected survival.

Algorithms↗

Searching for alcoholism susceptibility genes using Markov chain Monte Carlo methods.

Markov chain Monte Carlo (MCMC) methods offer a rapid parametric approach that can test for linkage throughout the entire genome. It has an advantage similar to nonparametric methods in that the model does not have to be completely specified a priori. However, unlike nonparametric methods, there are no limitations on pedigree size and MCMC methods can also handle relatively complex pedigree structures. In addition MCMC methods can be used to carry segregation analysis in order to answer questions on the genetic components of a disease phenotype. Segregation analysis gave evidence for between two and eight alcoholism susceptibility loci, each having a modest effect on the phenotype. MCMC methods were used to map alcoholism loci using the phenotypes ALDX1 (DSM-III-R and Feighner criteria) and ALDX2 (World Health Organization diagnosis ICD-10 criteria). There was mild evidence for quantitative trait loci on chromosomes 2, 10, and 11.

Adolescent↗

[Statistical properties of Markov chain in land use and landscape study].

Markov chain has been widely applied in the study of land use and landscape changes, but its statistical properties were less tested. Based on the land use change data monitored in Beijing, and with Pearson chi-squared goodness-of-fit test, this paper examined the time stability and time independence of Markov chain of land use change. The results indicated that the hypothesis of time stationary and Markov property is not tenable, which meant that the land use change in Beijing was an un-stationary and highly order Markov chain. Pearson chi-squared test was not as restricted a the likelihood ratio test, its transitional probabilities being allowed to be greater than zero, and thus, could be more useful in testing the assumption of homogeneity and independence.

China↗

Exchangeability in multivariate Markov chain models.

Time-homogeneous Markov chain models with state space [0, 1]k are useful in analysis of binary follow-up data on k individuals that interact. The number of parameters increases exponentially with k so more restrictive models are imperative for statistical inference. The hypothesis that the matrix of transition probabilities is invariant under permutation of individuals is discussed. It is shown that if individuals are exchangeable, then the process counting the number of individuals occupying a given state is a Markov chain. This reduction of data is sufficient if either at most a single individual may change state between two consecutive time points or if a state is absorbing. Similar results are obtained for exchangeability within two subgroups. Inference in the multivariate process reduces to a univariate problem if individuals are independent given the group's previous response. It is shown how conditional independence could be tested assuming exchangeability. The different hypotheses re examined in an analysis of the occurrence of bacteria in milk samples of Danish dairy cattle.

Animals↗