PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 523 records · Page 29Linked to original sources

A mover-stayer model for longitudinal marker data.

Studies of chronic disease often focus on estimating prevalence and incidence in which the presence of active disease is based on dichotomizing a continuous marker variable measured with error. Examples include hypertension, asthma, and depression, where active disease is defined by setting a threshold on a continuous measure of blood pressure, respiratory function, and mood, respectively. This paper proposes a model for inference about prevalence and incidence when active disease is determined by dichotomizing a continuous marker variable in a population-based study. In this formulation, it is postulated that there are three groups of people, those that are not susceptible to the disease, those who are always in the disease state, and those who have the potential to transition between the disease and the disease-free states over time. The model is used to estimate the prevalence and incidence of the disease in the population while accounting for measurement error in the marker. An EM algorithm is used for parameter estimation and the methodology is illustrated on Framingham heart study hypertension data. A simulation study is conducted in order to demonstrate the importance of accounting for measurement error in estimating prevalence and incidence for this example.

Adult↗

Estimation in a Cox proportional hazards cure model.

Some failure time data come from a population that consists of some subjects who are susceptible to and others who are nonsusceptible to the event of interest. The data typically have heavy censoring at the end of the follow-up period, and a standard survival analysis would not always be appropriate. In such situations where there is good scientific or empirical evidence of a nonsusceptible population, the mixture or cure model can be used (Farewell, 1982, Biometrics 38, 1041-1046). It assumes a binary distribution to model the incidence probability and a parametric failure time distribution to model the latency. Kuk and Chen (1992, Biometrika 79, 531-541) extended the model by using Cox's proportional hazards regression for the latency. We develop maximum likelihood techniques for the joint estimation of the incidence and latency regression parameters in this model using the nonparametric form of the likelihood and an EM algorithm. A zero-tail constraint is used to reduce the near nonidentifiability of the problem. The inverse of the observed information matrix is used to compute the standard errors. A simulation study shows that the methods are competitive to the parametric methods under ideal conditions and are generally better when censoring from loss to follow-up is heavy. The methods are applied to a data set of tonsil cancer patients treated with radiation therapy.

Algorithms↗

A nonparametric mixture model for cure rate estimation.

Nonparametric methods have attracted less attention than their parametric counterparts for cure rate analysis. In this paper, we study a general nonparametric mixture model. The proportional hazards assumption is employed in modeling the effect of covariates on the failure time of patients who are not cured. The EM algorithm, the marginal likelihood approach, and multiple imputations are employed to estimate parameters of interest in the model. This model extends models and improves estimation methods proposed by other researchers. It also extends Cox's proportional hazards regression model by allowing a proportion of event-free patients and investigating covariate effects on that proportion. The model and its estimation method are investigated by simulations. An application to breast cancer data, including comparisons with previous analyses using a parametric model and an existing nonparametric model by other researchers, confirms the conclusions from the parametric model but not those from the existing nonparametric model.

Algorithms↗

Risk assessment for quantitative responses using a mixture model.

A problem that frequently occurs in biological experiments with laboratory animals is that some subjects are less susceptible to the treatment than others. A mixture model has traditionally been proposed to describe the distribution of responses in treatment groups for such experiments. Using a mixture dose-response model, we derive an upper confidence limit on additional risk, defined as the excess risk over the background risk due to an added dose. Our focus will be on experiments with continuous responses for which risk is the probability of an adverse effect defined as an event that is extremely rare in controls. The asymptotic distribution of the likelihood ratio statistic is used to obtain the upper confidence limit on additional risk. The method can also be used to derive a benchmark dose corresponding to a specified level of increased risk. The EM algorithm is utilized to find the maximum likelihood estimates of model parameters and an extension of the algorithm is proposed to derive the estimates when the model is subject to a specified level of added risk. An example is used to demonstrate the results, and it is shown that by using the mixture model a more accurate measure of added risk is obtained.

Air Pollutants, Occupational↗

A transitional model for longitudinal binary data subject to nonignorable missing data.

Binary longitudinal data are often collected in clinical trials when interest is on assessing the effect of a treatment over time. Our application is a recent study of opiate addiction that examined the effect of a new treatment on repeated urine tests to assess opiate use over an extended follow-up. Drug addiction is episodic, and a new treatment may affect various features of the opiate-use process such as the proportion of positive urine tests over follow-up and the time to the first occurrence of a positive test. Complications in this trial were the large amounts of dropout and intermittent missing data and the large number of observations on each subject. We develop a transitional model for longitudinal binary data subject to nonignorable missing data and propose an EM algorithm for parameter estimation. We use the transitional model to derive summary measures of the opiate-use process that can be compared across treatment groups to assess treatment effect. Through analyses and simulations, we show the importance of properly accounting for the missing data mechanism when assessing the treatment effect in our example.

Algorithms↗

Latent variable models for longitudinal data with multiple continuous outcomes.

Multiple outcomes are often used to properly characterize an effect of interest. This paper proposes a latent variable model for the situation where repeated measures over time are obtained on each outcome. These outcomes are assumed to measure an underlying quantity of main interest from different perspectives. We relate the observed outcomes using regression models to a latent variable, which is then modeled as a function of covariates by a separate regression model. Random effects are used to model the correlation due to repeated measures of the observed outcomes and the latent variable. An EM algorithm is developed to obtain maximum likelihood estimates of model parameters. Unit-specific predictions of the latent variables are also calculated. This method is illustrated using data from a national panel study on changes in methadone treatment practices.

Longitudinal Studies↗

Semiparametric regression analysis of interval-censored data.

We propose a semiparametric approach to the proportional hazards regression analysis of interval-censored data. An EM algorithm based on an approximate likelihood leads to an M-step that involves maximizing a standard Cox partial likelihood to estimate regression coefficients and then using the Breslow estimator for the unknown baseline hazards. The E-step takes a particularly simple form because all incomplete data appear as linear terms in the complete-data log likelihood. The algorithm of Turnbull (1976, Journal of the Royal Statistical Society, Series B 38, 290-295) is used to determine times at which the hazard can take positive mass. We found multiple imputation to yield an easily computed variance estimate that appears to be more reliable than asymptotic methods with small to moderately sized data sets. In the right-censored survival setting, the approach reduces to the standard Cox proportional hazards analysis, while the algorithm reduces to the one suggested by Clayton and Cuzick (1985, Applied Statistics 34, 148-156). The method is illustrated on data from the breast cancer cosmetics trial, previously analyzed by Finkelstein (1986, Biometrics 42, 845-854) and several subsequent authors.

Algorithms↗

A Bayesian approach to finite mixture models in bioassay via data augmentation and Gibbs sampling and its application to-insecticide resistance.

After continued treatment with an insecticide, within the population of the susceptible insects, resistant strains will occur. It is important to know whether there are any resistant strains, what the proportions are, and what the median lethal doses are for the insecticide. Lwin and Martin (1989, Biometrics 45, 721-732) propose a probit mixture model and use the EM algorithm to obtain the maximum likelihood estimates for the parameters. This approach has difficulties in estimating the confidence intervals and in testing the number of components. We propose a Bayesian approach to obtaining the credible intervals for the location and scale of the tolerances in each component and for the mixture proportions by using data augmentation and Gibbs sampler. We use Bayes factor for model selection and determining the number of components. We illustrate the method with data published in Lwin and Martin (1989).

Algorithms↗

ROC curve estimation when covariates affect the verification process.

A receiver operating characteristic (ROC) curve is commonly used to measure the accuracy of a medical test. It is a plot of the true positive fraction (sensitivity) against the false positive fraction (1-specificity) for increasingly stringent positivity criterion. Bias can occur in estimation of an ROC curve if only some of the tested patients are selected for disease verification and if analysis is restricted only to the verified cases. This bias is known as verification bias. In this paper, we address the problem of correcting for verification bias in estimation of an ROC curve when the verification process and efficacy of the diagnostic test depend on covariates. Our method applies the EM algorithm to ordinal regression models to derive ML estimates for ROC curves as a function of covariates, adjusted for covariates affecting the likelihood of being verified. Asymptotic variance estimates are obtained using the observed information matrix of the observed data. These estimates are derived under the missing-at-random assumption, which means that selection for disease verification depends only on the observed data, i.e., the test result and the observed covariates. We also address the issues of model selection and model checking. Finally, we illustrate the proposed method on data from a two-phase study of dementia disorders, where selection for verification depends on the screening test result and age.

Aged↗

Maximum likelihood analysis of logistic regression models with incomplete covariate data and auxiliary information.

This article presents a new method for maximum likelihood estimation of logistic regression models with incomplete covariate data where auxiliary information is available. This auxiliary information is extraneous to the regression model of interest but predictive of the covariate with missing data. Ibrahim (1990, Journal of the American Statistical Association 85, 765-769) provides a general method for estimating generalized linear regression models with missing covariates using the EM algorithm that is easily implemented when there is no auxiliary data. Vach (1997, Statistics in Medicine 16, 57-72) describes how the method can be extended when the outcome and auxiliary data are conditionally independent given the covariates in the model. The method allows the incorporation of auxiliary data without making the conditional independence assumption. We suggest tests of conditional independence and compare the performance of several estimators in an example concerning mental health service utilization in children. Using an artificial dataset, we compare the performance of several estimators when auxiliary data are available.

Analysis of Variance↗

Semiparametric maximum likelihood for measurement error model regression.

This paper presents an EM algorithm for semiparametric likelihood analysis of linear, generalized linear, and nonlinear regression models with measurement errors in explanatory variables. A structural model is used in which probability distributions are specified for (a) the response and (b) the measurement error. A distribution is also assumed for the true explanatory variable but is left unspecified and is estimated by nonparametric maximum likelihood. For various types of extra information about the measurement error distribution, the proposed algorithm makes use of available routines that would be appropriate for likelihood analysis of (a) and (b) if the true x were available. Simulations suggest that the semiparametric maximum likelihood estimator retains a high degree of efficiency relative to the structural maximum likelihood estimator based on correct distributional assumptions and can outperform maximum likelihood based on an incorrect distributional assumption. The approach is illustrated on three examples with a variety of structures and types of extra information about the measurement error distribution.

Algorithms↗

Nonparametric mixed effects models for unequally sampled noisy curves.

We propose a method of analyzing collections of related curves in which the individual curves are modeled as spline functions with random coefficients. The method is applicable when the individual curves are sampled at variable and irregularly spaced points. This produces a low-rank, low-frequency approximation to the covariance structure, which can be estimated naturally by the EM algorithm. Smooth curves for individual trajectories are constructed as best linear unbiased predictor (BLUP) estimates, combining data from that individual and the entire collection. This framework leads naturally to methods for examining the effects of covariates on the shapes of the curves. We use model selection techniques--Akaike information criterion (AIC), Bayesian information criterion (BIC), and cross-validation--to select the number of breakpoints for the spline approximation. We believe that the methodology we propose provides a simple, flexible, and computationally efficient means of functional data analysis.

Algorithms↗

Estimating the frequency distribution of crossovers during meiosis from recombination data.

Estimation of tetrad crossover frequency distributions from genetic recombination data is a classic problem dating back to Weinstein (1936, Genetics 21, 155-199). But a number of important issues, such as how to specify the maximum number of crossovers, how to construct confidence intervals for crossover probabilities, and how to obtain correct p-values for hypothesis tests, have never been adequately addressed. In this article, we obtain some properties of the maximum likelihood estimate (MLE) for crossover probabilities that imply guidelines for choosing the maximum number of crossovers. We give these results for both normal meiosis and meiosis with nondisjunction. We also develop an accelerated EM algorithm to find the MLE more efficiently. We propose bootstrap-based methods to find confidence intervals and p-values and conduct simulation studies to check the validity of the bootstrap approach.

Algorithms↗

A maximum likelihood-based method for mining major genes affecting a quantitative character.

In this article, we present a maximum likelihood-based analytical approach for detecting a major gene of large effect on a quantitative trait in a progeny population derived from a mating design. Our analysis is based on a mixed genetic model specifying both major gene and background polygenic inheritance. The likelihood of the data is formulated by combining the information about population behaviors of the major gene during hybridization and its phenotypic distribution densities. The EM algorithm is implemented to obtain maximum likelihood estimates for population and quantitative genetic parameters of the major locus. This approach is applied to detect an overdominant gene governing stem volume growth in a factorial mating design of aspen trees. It is suggested that further molecular genetic research toward mapping single genes affecting aspen growth and production based on the same experimental data has a high probability of success.

Algorithms↗

Maximum likelihood estimation of two-level latent variable models with mixed continuous and polytomous data.

Two-level data with hierarchical structure and mixed continuous and polytomous data are very common in biomedical research. In this article, we propose a maximum likelihood approach for analyzing a latent variable model with these data. The maximum likelihood estimates are obtained by a Monte Carlo EM algorithm that involves the Gibbs sampler for approximating the E-step and the M-step and the bridge sampling for monitoring the convergence. The approach is illustrated by a two-level data set concerning the development and preliminary findings from an AIDS preventative intervention for Filipina commercial sex workers where the relationship between some latent quantities is investigated.

Acquired Immunodeficiency Syndrome↗

Frailty models with missing covariates.

We present a method for estimating the parameters in random effects models for survival data when covariates are subject to missingness. Our method is more general than the usual frailty model as it accommodates a wide range of distributions for the random effects, which are included as an offset in the linear predictor in a manner analogous to that used in generalized linear mixed models. We propose using a Monte Carlo EM algorithm along with the Gibbs sampler to obtain parameter estimates. This method is useful in reducing the bias that may be incurred using complete-case methods in this setting. The methodology is applied to data from Eastern Cooperative Oncology Group melanoma clinical trials in which observations were believed to be clustered and several tumor characteristics were not always observed.

Clinical Trials, Phase III as Topic↗

Methods for analyzing the spatial distribution of chiasmata during meiosis based on recombination data.

Using genetic recombination data to make inferences about chiasmata on the tetrad during meiosis is a classic problem dating back to Weinstein's paper in 1936 (Genetics 21, 155-199). In the last few years, Weinstein's methods have been revived and applied to new problems, but a number of important statistical issues remain unresolved. Recently, we developed improved statistical methods for studying the frequency distribution of the number of chiasmata (Yu and Feingold, 2001, Biometrics 57, 427-434). In the current article, we develop methods for the complementary issue of studying the spatial distribution of chiasmata. Somewhat different statistical approaches are needed for the spatial problem than for the frequency problem because different scientific questions are of interest. We explore the properties of the maximum likelihood estimate (MLE) for chiasma spatial distributions and propose improvements to the estimation procedures. We develop a class of statistical tests for comparing chiasma patterns in tetrads that have undergone normal meiosis and tetrads that have had a nondisjunction event. Finally, we propose an EM algorithm to find the MLE when the observed data is ambiguous, as is often the case in human datasets. We apply our improved methods to reanalyze a dataset from the literature studying the association between crossover location and meiotic nondisjunction of chromosome 21.

Algorithms↗

A frailty model for informative censoring.

To account for the correlation between failure and censoring, we propose a new frailty model for clustered data. In this model, the risk to be censored is affected by the risk of failure. This model allows flexibility in the direction and degree of dependence between failure and censoring. It includes the traditional frailty model as a special case. It allows censoring by some causes to be analyzed as informative while treating censoring by other causes as noninformative. It can also analyze data for competing risks. To fit the model, the EM algorithm is used with Markov chain Monte Carlo simulations in the E-steps. Simulation studies and analysis of data for kidney disease patients are provided. Consequences of incorrectly assuming noninformative censoring are investigated.

Algorithms↗