PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Adjusting for age-related competing mortality in long-term cancer clinical trials.

Mortality related to causes other than the treated disease may have a significant impact on overall survival in long-term clinical trials. We present a model that adjusts for age-related competing mortality when cause of death is missing or only partially available. Through use of a piecewise exponential survival model, we extend relative survival methods to continuous follow-up data, allowing the competing mortality to differ from that of the general population by a scale parameter. An EM algorithm provides a simple way to compute the maximum likelihood estimators (MLEs) and to test hypotheses using widely available software. We compare the bias and relative efficiency of this model to a piecewise exponential Cox model for overall survival. Theoretical results are confirmed by simulations and illustrated with data from a clinical trial in colorectal cancer. This example also shows how age-related and disease-related mortality can be confounded in an analysis of overall survival. We conclude with a discussion of the advantages and disadvantages of the model.

Adolescent↗

A method of non-parametric back-projection and its application to AIDS data.

The method of back-projection has been used to estimate the unobserved past incidence of infection with the human immunodeficiency virus (HIV) and to obtain projections of future AIDS incidence. Here a new approach to back-projection, which avoids parametric assumptions about the form of the HIV infection intensity, is described. This approach gives the data greater opportunity to determine the shape of the estimated intensity function. The method is based on a modification of an EM algorithm for maximum likelihood estimation that incorporates smoothing of the estimated parameters. It is easy to implement on a computer because the computations are based on explicit formulae. The method is illustrated with applications to AIDS data from Australia, U.S.A. and Japanese haemophiliacs.

Acquired Immunodeficiency Syndrome↗

Computational aspects of analysing random effects/longitudinal models.

Random effects and longitudinal models are becoming increasingly popular in the analysis of many types of data, including medical and biopharmaceutical, because of their richness and flexibility. They can be, however, difficult to fit using traditional statistical tools. Fortunately, there now exists a burgeoning collection of newer computational methods that can be applied to draw inferences with such models. This review attempts to provide an introduction to some of these techniques by describing them as extensions of the EM algorithm, currently a standard tool for the analysis of longitudinal and random effects models. For clarity of exposition, the extensions are classified into three types: large-sample iterative; large-sample simulation, and small-sample simulation.

Longitudinal Studies↗

Methods for the analysis of informatively censored longitudinal data.

This paper describes the problem of informative censoring in longitudinal studies where the primary outcome is rate of change in a continuous variable. Standard approaches based on the linear random effects model are valid only when the data are missing in a non-ignorable fashion. Informative censoring, which is a special type of non-ignorably missing data, occurs when the probability of early termination is related to an individual subject's true rate of change. When present, informative censoring causes bias in standard likelihood-based analyses, as well as in weighted averages of individual least-squares slopes. This paper reviews several methods proposed by others for analysis of informatively censored longitudinal data, and outlines a new approach based on a log-normal survival model. Maximum likelihood estimates may be obtained via the EM algorithm. Advantages of this approach are that it allows general unbalanced data caused by staggered entry and unequally-timed visits, it utilizes all available data, including data from patients with only a single measurement, and it provides a unified method for estimating all model parameters. Issues related to study design when informative censoring may occur are also discussed.

Linear Models↗

A latent class model for repeated measurements experiments.

Standard models for the analysis of repeated measurements assume a common response profile for all experimental units within a treatment group. However, in many applications this under-represents the nature of the response. There may be several distinct modes of response within a group (for example, responders versus non-responders to a given treatment), or there may be a set of distinct response profiles which are common to all the treatment groups. In these situations the effect of treatment can be characterized both by the shape of the fitted profiles and by estimating the proportion of cases who exhibit each particular response profile. This paper describes how such experiments may be analysed through the introduction of a latent variable into the standard model. Maximum likelihood estimation is straight-forward using the EM algorithm. Model choice requires some care, but good-fitting models can be identified via inspection of residuals and the use of empirical semi-variogram plots. Once the number of distinct profiles has been determined, treatment effects can be investigated using likelihood-ratio statistics. The approach is illustrated with a re-analysis of a dataset first described by Grizzle and Allen.

Animals↗

Studying the relationship between change and initial value in longitudinal studies.

Blomqvist's problem of studying the relationship between change and initial value in a linear growth curve setting is reformulated from a random effects model perspective. First, a maximum likelihood estimate of the between-individual covariance matrix for a simple linear regression model with stochastic parameters is obtained via an EM algorithm as discussed by Laird and Ware. Second, the regression coefficient of the individual-specific slopes on the individual-specific intercepts is estimated as a ratio of elements of the between-individual covariance matrix as discussed by Zucker et al. Then a Fieller's type confidence interval for this ratio is proposed. Discussion is facilitated by recognizing the Laird-Ware model as a special case of a more general model discussed by Hocking.

Algorithms↗

Using time of first positive HIV test and other auxiliary data in back-projection of AIDS incidence.

Estimation of HIV incidence by the method of back-projection typically uses data on the time of diagnosis of AIDS cases, together with known information about the incubation distribution of AIDS. This paper discusses back-projection using auxiliary data on AIDS cases, particularly the time of first positive HIV test. We discuss the possibility that certain types of auxiliary data, including time of first positive test, can be useful in back-projection because they provide extra information about the incubation period of AIDS cases. Under a back-projection model, theoretical efficiency calculations are given comparing back-projection with and without the time of first positive HIV test of AIDS cases. These calculations suggest that such data have the potential to significantly improve HIV incidence estimates, particularly in the recent past. Smoothed non-parametric estimates of both HIV incidence and time-dependent testing rates are described. These can be obtained using the EM algorithm, in conjunction with a smoothing step or a penalized likelihood. The benefit of these methods in practice needs to be assessed as such data become available.

Acquired Immunodeficiency Syndrome↗

The analysis of repeated-measures data on schizophrenic reaction times using mixture models.

Reaction times for schizophrenic individuals in a simple visual tracking experiment can be substantially more variable than for non-schizophrenic individuals. Current psychological theory suggests that at least some of this extra variability arises from an attentional lapse that delays some, but not all, of each schizophrenic's reaction times. Based on this theory, we pursue models in which measurements from non-schizophrenics arise from a normal linear model with a separate mean for each individual, whereas measurements from schizophrenics arise from a mixture of (i) a component analogous to the distribution of response times for non-schizophrenics and (ii) a mean-shifted component. We fit four mixture models within this framework, where the distinctions between models arise from assumptions about the variance of the shifted observations and the exchangeability of schizophrenic individuals. Some of these models can be fit by maximum likelihood using the EM algorithm, and all can be fit using the ECM algorithm, where the covariance matrices associated with the parameters are calculated by the SEM and SECM algorithms, respectively. Bayesian model monitoring using posterior predictive checks is invoked to discard models that fail to reproduce certain observed features of the data and to stimulate the development of better models.

Algorithms↗

Re-evaluating design specifications of longitudinal clinical trials without unblinding when the key response is rate of change.

The design of clinical trials often requires knowledge of quantities such as between- and within subject variances about which only imprecise information exists. To provide assurance that the study has the desired power to detect a minimum clinically meaningful difference between treatment groups. Gould, Gould and Shih, and Shih have recommended obtaining relevant information from the trial at an interim stage without unblinding. Wittes and Brittain provided a similar recommendation, but viewed the portion up to the interim stage as an (internal) pilot study and required unblinding. This paper considers the problem of re-evaluating the design specifications in longitudinal clinical trials when the key response is the rate of change (slope). The proposed method aims to re-evaluate the sample size and study duration in a way that maintains the trial's blinding, using an EM algorithm. Simulation results show that the effect on type I error rate in negligible, but the potential gain in power can be substantial. The procedure is simple to use in practice, as it does not unblind patients' treatment identifications, and, since it does not unveil the relative efficacy of treatments, it fulfils the requirement of a valid 'administrative' (interim) analysis.

Algorithms↗

Assessing human fertility using several markers of ovulation.

In modelling human fertility one ideally accounts for timing of intercourse relative to ovulation. Measurement error in identifying the day of ovulation can bias estimates of fecundability parameters and attenuate estimates of covariate effects. In the absence of a single perfect marker of ovulation, several error prone markers are sometimes obtained. In this paper we propose a semi-parametric mixture model that uses multiple independent markers of ovulation to account for measurement error. The model assigns each method of assessing ovulation a distinct non-parametric error distribution, and corrects bias in estimates of day-specific fecundability. We use a Monte Carlo EM algorithm for joint estimation of (i) the error distribution for the markers, (ii) the error-corrected fertility parameters, and (iii) the couple-specific random effects. We apply the methods to data from a North Carolina fertility study to assess the magnitude of error in measures of ovulation based on urinary luteinizing hormone and metabolites of ovarian hormones, and estimate the corrected day-specific probabilities of clinical pregnancy. Published in 2001 by John Wiley & Sons, Ltd.

Algorithms↗

Analysis of change in the presence of informative censoring: application to a longitudinal clinical trial of progressive renal disease.

The rate of change in a continuous variable, measured serially over time, is often used as an outcome in longitudinal studies or clinical trials. When patients terminate the study before the scheduled end of the study, there is a potential for bias in estimation of rate of change using standard methods which ignore the missing data mechanism. These methods include the use of unweighted generalized estimating equations methods and likelihood-based methods assuming an ignorable missing data mechanism. We present a model for analysis of informatively censored data, based on an extension of the two-stage linear random effects model, where each subject's random intercept and slope are allowed to be associated with an underlying time to event. The joint distribution of the continuous responses and the time-to-event variable are then estimated via maximum likelihood using the EM algorithm, and using the bootstrap to calculate standard errors. We illustrate this methodology and compare it to simpler approaches and usual maximum likelihood using data from a multi-centre study of the effects of diet and blood pressure control on progression of renal disease, the Modification of Diet in Renal Disease (MDRD) Study. Sensitivity analyses and simulations are used to evaluate the performance of this methodology in the context of the MDRD data, under various scenarios where the drop-out mechanism is ignorable as well as non-ignorable.

Algorithms↗

Proportional hazards model for interval-censored failure times and time-dependent covariates: application to hazard of HIV infection of injecting drug users in prison.

Interval-censored survival data are data in which the failure times are not known precisely, but are known to lie within an interval. Such data can be analysed using a proportional hazards model with piecewise-exponential baseline hazard, a model which can be fitted by an EM algorithm easily programmed in standard statistical software. In this paper we extend the model to allow for time-dependent covariates and left-truncation, and demonstrate its use by assessing the effect of imprisonment on hazard of HIV infection in a cohort of injecting drug users from Edinburgh. No conclusive effect of incarceration on hazard of HIV infection was found, but there was a suggestion that imprisonment might have been a significant relative risk factor for infection in the later period, when risk behaviour among drug users in the community was reduced.

Adolescent↗

Modelling variation of lower leg length growth in early life.

We consider the estimation of sources of variation for panel data with repeated measurements. With no repeated measurements and known measurement error, models for variation decomposition have been proposed when there are one or more types of measurements. Estimation was performed using the EM algorithm accompanied by model augmentation that demands more computational efforts. In this article we extend previous variation models and modify the estimation methods in order to estimate various variation components after eliminating the unknown effects of measurement error. Specifically, methods that dispense with model augmentation and estimation of time-dependent covariates are considered. A set of lower leg length data from Chinese infants is analysed by using the proposed model. Interestingly, our results are consistent with the well-accepted three-phase (infancy-childhood-puberty) growth transition proposition for human growth. Moreover, gender effect is found to be time-varying.

Anthropometry↗

Complete imputation of missing repeated categorical data: one-sample applications.

Longitudinal studies with repeated measures are often subject to non-response. Methods currently employed to alleviate the difficulties caused by missing data are typically unsatisfactory, especially when the cause of the missingness is related to the outcomes. We present an approach for incomplete categorical data in the repeated measures setting that allows missing data to depend on other observed outcomes for a study subject. The proposed methodology also allows a broader examination of study findings through interpretation of results in the framework of the set of all possible test statistics that might have been observed had no data been missing. The proposed approach consists of the following general steps. First, we generate all possible sets of missing values and form a set of possible complete data sets. We then weight each data set according to clearly defined assumptions and apply an appropriate statistical test procedure to each data set, combining the results to give an overall indication of significance. We make use of the EM algorithm and a Bayesian prior in this approach. While not restricted to the one-sample case, the proposed methodology is illustrated for one-sample data and compared to the common complete-case and available-case analysis methods.

Algorithms↗

Maximum likelihood estimation of a survival function with a change point for truncated and interval-censored data.

This paper considers estimation of a survival function when there exists a change point and the survival time of interest is defined as elapsed time between two related events. Furthermore, there exists censoring on observations on the occurrences of both events and truncation on observations on the occurrence of the second event and thus the survival time of interest. To obtain the maximum likelihood estimator of a survival function, an EM algorithm is developed when the survival function is completely unknown before the change point and known up to a vector of unknown parameters after the change point. The idea is a generalization of that discussed in Moeschberger and Klein. Simulations and an example are used to evaluate and illustrate the algorithm.

Acquired Immunodeficiency Syndrome↗

A local likelihood proportional hazards model for interval censored data.

We discuss the use of local likelihood methods to fit proportional hazards regression models to right and interval censored data. The assumed model allows for an arbitrary, smoothed baseline hazard on which a vector of covariates operates in a proportional manner, and thus produces an interpretable baseline hazard function along with estimates of global covariate effects. For estimation, we extend the modified EM algorithm suggested by Betensky, Lindsey, Ryan and Wand. We illustrate the method with data on times to deterioration of breast cosmeses and HIV-1 infection rates among haemophiliacs.

Algorithms↗

Sample Size Requirements of a Mixture Analysis Method with Applications in Systematic Biology.

The available information on sample size requirements of mixture analysis methods is insufficient to permit a precise evaluation of the potential problems facing practical applications of mixture analysis. We use results from Monte Carlo simulation to assess the sample size requirements of a simple mixture analysis method under conditions relevant to biological applications of mixture analysis. The mixture model used includes two univariate normal components with equal variances but assumes that the researcher is ignorant as to the equality of the variances. The method used relies on the EM algorithm to compute the maximum likelihood estimates of the mixture parameters, and the likelihood ratio test to assess the number of components in the mixtures. Our results suggest that sample sizes close to 500 or 1000 data may be required to adequately solve mixtures commonly found in biology. Sample sizes of 500 or 1000 are difficult to achieve. However, use of this MA method may be a reasonable option when the researcher deals with problems which are intractable by other means. Copyright 1999 Academic Press.

Journal Article↗

A general statistical analysis for fMRI data.

We propose a method for the statistical analysis of fMRI data that seeks a compromise between efficiency, generality, validity, simplicity, and execution speed. The main differences between this analysis and previous ones are: a simple bias reduction and regularization for voxel-wise autoregressive model parameters; the combination of effects and their estimated standard deviations across different runs/sessions/subjects via a hierarchical random effects analysis using the EM algorithm; overcoming the problem of a small number of runs/session/subjects using a regularized variance ratio to increase the degrees of freedom.

Algorithms↗