PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 541 records · Page 30Linked to original sources

Small-sample inference for incomplete longitudinal data with truncation and censoring in tumor xenograft models.

In cancer drug development, demonstrating activity in xenograft models, where mice are grafted with human cancer cells, is an important step in bringing a promising compound to humans. A key outcome variable is the tumor volume measured in a given period of time for groups of mice given different doses of a single or combination anticancer regimen. However, a mouse may die before the end of a study or may be sacrificed when its tumor volume quadruples, and its tumor may be suppressed for some time and then grow back. Thus, incomplete repeated measurements arise. The incompleteness or missingness is also caused by drastic tumor shrinkage (<0.01 cm3) or random truncation. Because of the small sample sizes in these models, asymptotic inferences are usually not appropriate. We propose two parametric test procedures based on the EM algorithm and the Bayesian method to compare treatment effects among different groups while accounting for informative censoring. A real xenograft study on a new antitumor agent, temozolomide, combined with irinotecan is analyzed using the proposed methods.

Algorithms↗

A semiparametric likelihood approach to joint modeling of longitudinal and time-to-event data.

Joint models for a time-to-event (e.g., survival) and a longitudinal response have generated considerable recent interest. The longitudinal data are assumed to follow a mixed effects model, and a proportional hazards model depending on the longitudinal random effects and other covariates is assumed for the survival endpoint. Interest may focus on inference on the longitudinal data process, which is informatively censored, or on the hazard relationship. Several methods for fitting such models have been proposed, most requiring a parametric distributional assumption (normality) on the random effects. A natural concern is sensitivity to violation of this assumption; moreover, a restrictive distributional assumption may obscure key features in the data. We investigate these issues through our proposal of a likelihood-based approach that requires only the assumption that the random effects have a smooth density. Implementation via the EM algorithm is described, and performance and the benefits for uncovering noteworthy features are illustrated by application to data from an HIV clinical trial and by simulation.

Anti-HIV Agents↗

A Bayesian hierarchical model for categorical data with nonignorable nonresponse.

Log-linear models have been shown to be useful for smoothing contingency tables when categorical outcomes are subject to nonignorable nonresponse. A log-linear model can be fit to an augmented data table that includes an indicator variable designating whether subjects are respondents or nonrespondents. Maximum likelihood estimates calculated from the augmented data table are known to suffer from instability due to boundary solutions. Park and Brown (1994, Journal of the American Statistical Association 89, 44-52) and Park (1998, Biometrics 54, 1579-1590) developed empirical Bayes models that tend to smooth estimates away from the boundary. In those approaches, estimates for nonrespondents were calculated using an EM algorithm by maximizing a posterior distribution. As an extension of their earlier work, we develop a Bayesian hierarchical model that incorporates a log-linear model in the prior specification. In addition, due to uncertainty in the variable selection process associated with just one log-linear model, we simultaneously consider a finite number of models using a stochastic search variable selection (SSVS) procedure due to George and McCulloch (1997, Statistica Sinica 7, 339-373). The integration of the SSVS procedure into a Markov chain Monte Carlo (MCMC) sampler is straightforward, and leads to estimates of cell frequencies for the nonrespondents that are averages resulting from several log-linear models. The methods are demonstrated with a data example involving serum creatinine levels of patients who survived renal transplants. A simulation study is conducted to investigate properties of the model.

Algorithms↗

Modeling the dependence between number of trials and success probability in beta-binomial-Poisson mixture distributions.

Beta-binomial models are widely used for overdispersed binomial data, with the binomial success probability modeled as following a beta distribution. The number of binary trials in each binomial is assumed to be nonrandom and unrelated to the success probability. In many behavioral studies, however, binomial observations demonstrate more complex structures. In this article, a general beta-binomial-Poisson mixture model is developed, to allow for a relation between the number of trials and the success probability for overdispersed binomial data. An EM algorithm is implemented to compute both the maximum likelihood estimates of the model parameters and the corresponding standard errors. For illustration, the methodology is applied to study the feeding behavior of green-backed herons in two southeastern Missouri streams.

Algorithms↗

Estimation of competing risks with general missing pattern in failure types.

In competing risks data, missing failure types (causes) is a very common phenomenon. In this work, we consider a general missing pattern in which, if a failure type is not observed, one observes a set of possible types containing the true type, along with the failure time. We first consider maximum likelihood estimation with missing-at-random assumption via the expectation maximization (EM) algorithm. We then propose a Nelson-Aalen type estimator for situations when certain information on the conditional probability of the true type given a set of possible failure types is available from the experimentalists. This is based on a least-squares type method using the relationships between hazards for different types and hazards for different combinations of missing types. We conduct a simulation study to investigate the performance of this method, which indicates that bias may be small, even for high proportion of missing data, for sufficiently large number of observations. The estimates are somewhat sensitive to misspecification of the conditional probabilities of the true types when the missing proportion is high. We also consider an example from an animal experiment to illustrate our methodology.

Administration, Oral↗

Maximum likelihood methods for nonignorable missing responses and covariates in random effects models.

This article analyzes quality of life (QOL) data from an Eastern Cooperative Oncology Group (ECOG) melanoma trial that compared treatment with ganglioside vaccination to treatment with high-dose interferon. The analysis of this data set is challenging due to several difficulties, namely, nonignorable missing longitudinal responses and baseline covariates. Hence, we propose a selection model for estimating parameters in the normal random effects model with nonignorable missing responses and covariates. Parameters are estimated via maximum likelihood using the Gibbs sampler and a Monte Carlo expectation maximization (EM) algorithm. Standard errors are calculated using the bootstrap. The method allows for nonmonotone patterns of missing data in both the response variable and the covariates. We model the missing data mechanism and the missing covariate distribution via a sequence of one-dimensional conditional distributions, allowing the missing covariates to be either categorical or continuous, as well as time-varying. We apply the proposed approach to the ECOG quality-of-life data and conduct a small simulation study evaluating the performance of the maximum likelihood estimates. Our results indicate that a patient treated with the vaccine has a higher QOL score on average at a given time point than a patient treated with high-dose interferon.

Analysis of Variance↗

Joint analysis of time-to-event and multiple binary indicators of latent classes.

Multiple categorical variables are commonly used in medical and epidemiological research to measure specific aspects of human health and functioning. To analyze such data, models have been developed considering these categorical variables as imperfect indicators of an individual's "true" status of health or functioning. In this article, the latent class regression model is used to model the relationship between covariates, a latent class variable (the unobserved status of health or functioning), and the observed indicators (e.g., variables from a questionnaire). The Cox model is extended to encompass a latent class variable as predictor of time-to-event, while using information about latent class membership available from multiple categorical indicators. The expectation-maximization (EM) algorithm is employed to obtain maximum likelihood estimates, and standard errors are calculated based on the profile likelihood, treating the nonparametric baseline hazard as a nuisance parameter. A sampling-based method for model checking is proposed. It allows for graphical investigation of the assumption of proportional hazards across latent classes. It may also be used for checking other model assumptions, such as no additional effect of the observed indicators given latent class. The usefulness of the model framework and the proposed techniques are illustrated in an analysis of data from the Women's Health and Aging Study concerning the effect of severe mobility disability on time-to-death for elderly women.

Aged↗

An extended general location model for causal inferences from data subject to noncompliance and missing values.

Noncompliance is a common problem in experiments involving randomized assignment of treatments, and standard analyses based on intention-to-treat or treatment received have limitations. An attractive alternative is to estimate the Complier-Average Causal Effect (CACE), which is the average treatment effect for the subpopulation of subjects who would comply under either treatment (Angrist, Imbens, and Rubin, 1996, Journal of American Statistical Association 91, 444-472). We propose an extended general location model to estimate the CACE from data with noncompliance and missing data in the outcome and in baseline covariates. Models for both continuous and categorical outcomes and ignorable and latent ignorable (Frangakis and Rubin, 1999, Biometrika 86, 365-379) missing-data mechanisms are developed. Inferences for the models are based on the EM algorithm and Bayesian MCMC methods. We present results from simulations that investigate sensitivity to model assumptions and the influence of missing-data mechanism. We also apply the method to the data from a job search intervention for unemployed workers.

Algorithms↗

Maximum likelihood analysis of a general latent variable model with hierarchically mixed data.

A general two-level latent variable model is developed to provide a comprehensive framework for model comparison of various submodels. Nonlinear relationships among the latent variables in the structural equations at both levels, as well as the effects of fixed covariates in the measurement and structural equations at both levels, can be analyzed within the framework. Moreover, the methodology can be applied to hierarchically mixed continuous, dichotomous, and polytomous data. A Monte Carlo EM algorithm is implemented to produce the maximum likelihood estimate. The E-step is completed by approximating the conditional expectations through observations that are simulated by Markov chain Monte Carlo methods, while the M-step is completed by conditional maximization. A procedure is proposed for computing the complicated observed-data log likelihood and the BIC for model comparison. The methods are illustrated by using a real data set.

Acquired Immunodeficiency Syndrome↗

Shared frailty models for recurrent events and a terminal event.

There has been an increasing interest in the analysis of recurrent event data (Cook and Lawless, 2002, Statistical Methods in Medical Research 11, 141-166). In many situations, a terminating event such as death can happen during the follow-up period to preclude further occurrence of the recurrent events. Furthermore, the death time may be dependent on the recurrent event history. In this article we consider frailty proportional hazards models for the recurrent and terminal event processes. The dependence is modeled by conditioning on a shared frailty that is included in both hazard functions. Covariate effects can be taken into account in the model as well. Maximum likelihood estimation and inference are carried out through a Monte Carlo EM algorithm with Metropolis-Hastings sampler in the E-step. An analysis of hospitalization and death data for waitlisted dialysis patients is presented to illustrate the proposed methods. Methods to check the validity of the proposed model are also demonstrated. This model avoids the difficulties encountered in alternative approaches which attempt to specify a dependent joint distribution with marginal proportional hazards and yields an estimate of the degree of dependence.

Algorithms↗

Genetic susceptibility to prostate, breast, and colorectal cancer among Nordic twins.

To investigate the role of genetics in the development of cancer, we developed a new approach to analyze data on prostate, breast, and colorectal cancer from the Swedish, Danish, and Finnish twin registries on monozygotic (MZ) and same-sex dizygotic (DZ) twins. In the spirit of a sensitivity analysis, we modeled genetic inheritance as either an autosomal recessive or dominant cancer susceptibility (CS) genotype that involves either a single gene, many genes with equal allele frequencies, or three genes with a ninefold range of allele frequencies. We also modeled the joint probability of cancer incidence among five age categories, conditional on the presence or absence of the CS genotype. The main assumptions are: (1) The joint distribution of unobserved environmental effects in a twin pair conditional on the presence or absence of the CS genotype is the same for MZ and DZ twins, (2) the probability of cancer conditional on the presence or absence of the CS genotype and the unobserved environmental effects (i.e., the gene-environment interaction) is the same for MZ and DZ twins, and (3) the probability of cancer is independent between twins with the CS genotype. Estimation was maximum likelihood via a search over allele frequency and two levels of EM algorithms. Models had acceptable or good fits. Variability was estimated using a bootstrap approach, but only 50 replications were feasible. The 94th percentile of bootstrap replications for the estimated fraction of cancers with the CS genotype ranged, over the various genetic models, from 0.16 to 0.45 for prostate cancer, 0.12 to 0.30 for breast cancer, and 0.08 to 0.27 for colorectal cancer. We conclude that genetic susceptibility makes only a small to moderate contribution to the incidence of prostate, breast, and colorectal cancer.

Algorithms↗

Mixture model equations for marker-assisted genetic evaluation.

Marker-assisted genetic evaluation needs to infer genotypes at quantitative trait loci (QTL) based on the information of linked markers. As the inference usually provides the probability distribution of QTL genotypes rather than a specific genotype, marker-assisted genetic evaluation is characterized by the mixture model because of the uncertainty of QTL genotypes. It is, therefore, necessary to develop a statistical procedure useful for mixture model analyses. In this study, a set of mixture model equations was derived based on the normal mixture model and the EM algorithm for evaluating linear models with uncertain independent variables. The derived equations can be seen as an extension of Henderson's mixed model equations to mixture models and provide a general framework to deal with the issues of uncertain incidence matrices in linear models. The mixture model equations were applied to marker-assisted genetic evaluation with different parameterizations of QTL effects. A sire-QTL-effect model and a founder-QTL-effect model were used to illustrate the application of the mixture model equations. The potential advantages of the mixture model equations for marker-assisted genetic evaluation were discussed. The mixed-effect mixture model equations are flexible in modelling QTL effects and show desirable properties in estimating QTL effects, compared with Henderson's mixed model equations.

Animals↗

Peroxisome proliferator-activated receptor-gamma co-activator-1alpha (PGC-1alpha) gene polymorphisms and their relationship to Type 2 diabetes in Asian Indians.

AIMS: The objective of the present investigation was to examine the relationship of three polymorphisms, Thr394Thr, Gly482Ser and +A2962G, of the peroxisome proliferator activated receptor-gamma co-activator-1 alpha (PGC-1alpha) gene with Type 2 diabetes in Asian Indians. METHODS: The study group comprised 515 Type 2 diabetic and 882 normal glucose tolerant subjects chosen from the Chennai Urban Rural Epidemiology Study, an ongoing population-based study in southern India. The three polymorphisms were genotyped using polymerase chain reaction-restriction fragment length polymorphism (PCR-RFLP). Haplotype frequencies were estimated using an expectation-maximization (EM) algorithm. Linkage disequilibrium was estimated from the estimates of haplotypic frequencies. RESULTS: The three polymorphisms studied were not in linkage disequilibrium. With respect to the Thr394Thr polymorphism, 20% of the Type 2 diabetic patients (103/515) had the GA genotype compared with 12% of the normal glucose tolerance (NGT) subjects (108/882) (P = 0.0004). The frequency of the A allele was also higher in Type 2 diabetic subjects (0.11) compared with NGT subjects (0.07) (P = 0.002). Regression analysis revealed the odds ratio for Type 2 diabetes for the susceptible genotype (XA) to be 1.683 (95% confidence intervals: 1.264-2.241, P = 0.0004). Age adjusted glycated haemoglobin (P = 0.003), serum cholesterol (P = 0.001) and low-density lipoprotein (LDL) cholesterol (P = 0.001) levels and systolic blood pressure (P = 0.001) were higher in the NGT subjects with the XA genotype compared with GG genotype. There were no differences in genotype or allelic distribution between the Type 2 diabetic and NGT subjects with respect to the Gly482Ser and +A2962G polymorphisms. CONCLUSIONS: The A allele of Thr394Thr (G --> A) polymorphism of the PGC-1 gene is associated with Type 2 diabetes in Asian Indian subjects and the XA genotype confers 1.6 times higher risk for Type 2 diabetes compared with the GG genotype in this population.

Adult↗

Case-parent triads: estimating single- and double-dose effects of fetal and maternal disease gene haplotypes.

Case-parent triad data are considered a robust basis for studying association between variants of a gene and a disease. Methods evaluating statistical significance of association, like the TDT-test and its extensions, are frequently used. When there are prior hypotheses of a causal effect of the gene under study, however, methods measuring penetrance of alleles or haplotypes as relative risks will be more informative. Log-linear models have been proposed as a flexible tool for such relative risk estimation. We demonstrate an extension of the log-linear model to a natural framework for also estimating effects of multiple alleles or haplotypes, incorporating both single- and double-dose effects. The model also incorporates effects of single- and double-dose maternal haplotypes on a fetus during pregnancy. Unknown phase of haplotypes as well as missing parents are accounted for by the EM algorithm. A number of numerical improvements to maximum likelihood estimation are also implemented to facilitate a larger number of haplotypes. Software for these analyses, HAPLIN, is publicly available through our web site. As an illustration we have re-analyzed data on the MSX1 homeobox-gene on chromosome 4 to show how haplotypes may influence the risk of oral clefts.

Cleft Lip↗

Missing covariates in longitudinal data with informative dropouts: bias analysis and inference.

We consider estimation in generalized linear mixed models (GLMM) for longitudinal data with informative dropouts. At the time a unit drops out, time-varying covariates are often unobserved in addition to the missing outcome. However, existing informative dropout models typically require covariates to be completely observed. This assumption is not realistic in the presence of time-varying covariates. In this article, we first study the asymptotic bias that would result from applying existing methods, where missing time-varying covariates are handled using naive approaches, which include: (1) using only baseline values; (2) carrying forward the last observation; and (3) assuming the missing data are ignorable. Our asymptotic bias analysis shows that these naive approaches yield inconsistent estimators of model parameters. We next propose a selection/transition model that allows covariates to be missing in addition to the outcome variable at the time of dropout. The EM algorithm is used for inference in the proposed model. Data from a longitudinal study of human immunodeficiency virus (HIV)-infected women are used to illustrate the methodology.

Algorithms↗

A hybrid model for nonignorable dropout in longitudinal binary responses.

This article presents a likelihood-based method for handling nonignorable dropout in longitudinal studies with binary responses. The methodology developed is appropriate when the target of inference is the marginal distribution of the response at each occasion and its dependence on covariates. A "hybrid" model is formulated, which is designed to retain advantageous features of the selection and pattern-mixture model approaches. This formulation accommodates a variety of assumed forms of nonignorable dropout, while maintaining transparency of the constraints required for identifying the overall model. Once appropriate identifying constraints have been imposed, likelihood-based estimation is conducted via the EM algorithm. The article concludes by applying the approach to data from a randomized clinical trial comparing two doses of a contraceptive.

Algorithms↗

Semiparametric models for missing covariate and response data in regression models.

We consider a class of semiparametric models for the covariate distribution and missing data mechanism for missing covariate and/or response data for general classes of regression models including generalized linear models and generalized linear mixed models. Ignorable and nonignorable missing covariate and/or response data are considered. The proposed semiparametric model can be viewed as a sensitivity analysis for model misspecification of the missing covariate distribution and/or missing data mechanism. The semiparametric model consists of a generalized additive model (GAM) for the covariate distribution and/or missing data mechanism. Penalized regression splines are used to express the GAMs as a generalized linear mixed effects model, in which the variance of the corresponding random effects provides an intuitive index for choosing between the semiparametric and parametric model. Maximum likelihood estimates are then obtained via the EM algorithm. Simulations are given to demonstrate the methodology, and a real data set from a melanoma cancer clinical trial is analyzed using the proposed methods.

Algorithms↗

Structural inference in transition measurement error models for longitudinal data.

We propose a new class of models, transition measurement error models, to study the effects of covariates and the past responses on the current response in longitudinal studies when one of the covariates is measured with error. We show that the response variable conditional on the error-prone covariate follows a complex transition mixed effects model. The naive model obtained by ignoring the measurement error correctly specifies the transition part of the model, but misspecifies the covariate effect structure and ignores the random effects. We next study the asymptotic bias in naive estimator obtained by ignoring the measurement error for both continuous and discrete outcomes. We show that the naive estimator of the regression coefficient of the error-prone covariate is attenuated, while the naive estimators of the regression coefficients of the past responses are generally inflated. We then develop a structural modeling approach for parameter estimation using the maximum likelihood estimation method. In view of the multidimensional integration required by full maximum likelihood estimation, an EM algorithm is developed to calculate maximum likelihood estimators, in which Monte Carlo simulations are used to evaluate the conditional expectations in the E-step. We evaluate the performance of the proposed method through a simulation study and apply it to a longitudinal social support study for elderly women with heart disease. An additional simulation study shows that the Bayesian information criterion (BIC) performs well in choosing the correct transition orders of the models.

Aged↗