PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Estimation of an errors-in-variables regression model when the variances of the measurement errors vary between the observations.

It is common in the analysis of aggregate data in epidemiology that the variances of the aggregate observations are available. The analysis of such data leads to a measurement error situation, where the known variances of the measurement errors vary between the observations. Assuming multivariate normal distribution for the 'true' observations and normal distributions for the measurement errors, we derive a simple EM algorithm for obtaining maximum likelihood estimates of the parameters of the multivariate normal distributions. The results also facilitate the estimation of regression parameters between the variables as well as the 'true' values of the observations. The approach is applied to re-estimate recent results of the WHO MONICA Project on cardiovascular disease and its risk factors, where the original estimation of the regression coefficients did not adjust for the regression attenuation caused by the measurement errors.

Algorithms↗

Detecting and eliminating erroneous gestational ages: a normal mixture model.

In perinatal research and clinical practice, gestational age is a crucial variable for measuring foetal 'growth' (birth weight for gestational age) and for estimating the risk of mortality and morbidity, yet reported gestational age values are affected by random and systematic errors due to the absence of a gold standard measure. Previous investigators have used birth weight (which is measured with greater validity and precision than is gestational age) to correct such errors, but existing methods are inadequate due to unreasonable assumptions about the distributions of birth weight and gestational age. We propose a new method for identifying and correcting implausible observations using the expectation-maximization (EM) algorithm. Using population-based data from U.S. birth certificates, we compare the resulting gestational ages, birth weight distributions at each gestational age, and gestational age-specific infant mortality based on the new method with those on the same population produced by previous published correction methods. The new method gives the best birth weight distributions for gestational age and the most realistic gestational-age-specific mortality rates, while each of the other methods has at least one significant flaw.

Algorithms↗

Maximum likelihood estimation in the joint analysis of time-to-event and multiple longitudinal variables.

Joint modelling of longitudinal and survival data has received much attention in recent years. Most have concentrated on a single longitudinal variable. This paper considers joint modelling in the presence of multiple longitudinal variables. We explore direct association of time-to-event and multiple longitudinal processes through a frailty model and use a mixed effects model for each of the longitudinal variables. Correlations among the longitudinal variables are induced through correlated random effects. We allow effects of categorical and continuous covariates on both longitudinal and time-to-event responses and explore interactions between the longitudinal variables and other covariates on time-to-event. Estimates of the parameters are obtained by maximizing the joint likelihood for the longitudinal variable processes and the event process. We use a one-step-late EM algorithm to handle the direct dependence of the event process on the modelled longitudinal variables along with the presence of other fixed covariates in both processes. We argue that such a joint analysis with multiple longitudinal variables is advantageous to one with only a single longitudinal variable in revealing interplay among multiple longitudinal variables and the time-to-event.

Head and Neck Neoplasms↗

Survival curve estimation with partial non-random exposure information.

The objective of this paper is to estimate survival curves for two different exposure groups when the exposure group is not known for all observations, and the data is subject to left truncation and right censoring. The situation we consider is when the probability that the exposure group is missing may depend on whether the observation is censored or uncensored, in which case the exposure is not missing at random. The problem was motivated by a study of Alzheimer's disease to estimate the distribution of ages at diagnosis for individuals with and without an apolipoprotein E4 allele (the exposure group). Genotyping for this risk factor was incomplete and performed more frequently on the cases of Alzheimer's disease (the uncensored observations) than the censored observations. The survival curves are estimated in discrete time using an EM algorithm. A bootstrapping procedure is proposed that guarantees each bootstrap sample has the same proportion of observations with missing exposure. A simulation is performed to evaluate the bias of the estimators and to investigate design and efficiency issues. The methods are applied to the Alzheimer's disease study.

Adult↗

A semi-parametric accelerated failure time cure model.

A cure model is a useful approach for analysing failure time data in which some subjects could eventually experience, and others never experience, the event of interest. A cure model has two components: incidence which indicates whether the event could eventually occur and latency which denotes when the event will occur given the subject is susceptible to the event. In this paper, we propose a semi-parametric cure model in which covariates can affect both the incidence and the latency. A logistic regression model is proposed for the incidence, and the latency is determined by an accelerated failure time regression model with unspecified error distribution. An EM algorithm is developed to fit the model. The procedure is applied to a data set of tonsil cancer patients treated with radiation therapy.

Algorithms↗

Applications of continuous time hidden Markov models to the study of misclassified disease outcomes.

Disease progression in prospective clinical and epidemiological studies is often conceptualized in terms of transitions between disease states. Analysis of data from such studies can be complicated by a number of factors, including the presence of individuals in various prevalent disease states and with unknown prior disease history, interval censored observations of state transitions and misclassified measurements of disease states. We present an approach where the disease states are modelled as the hidden states of a continuous time hidden Markov model using the imperfect measurements of the disease state as observations. Covariate effects on transitions between disease states are incorporated using a generalized regression framework. Parameter estimation and inference are based on maximum likelihood methods and rely on an EM algorithm. In addition, techniques for model assessment are proposed. Applications to two binary disease outcomes are presented: the oral lesion hairy leukoplakia in a cohort of HIV infected men and cervical human papillomavirus (HPV) infection in a cohort of young women. Estimated transition rates and misclassification probabilities for the hairy leukoplakia data agree well with clinical observations on the persistence and diagnosis of this lesion, lending credibility to the interpretation of hidden states as representing the actual disease states. By contrast, interpretation of the results for the HPV data are more problematic, illustrating that successful application of the hidden Markov model may be highly dependent on the degree to which the assumptions of the model are satisfied.

Adolescent↗

A hierarchical Poisson mixture regression model to analyse maternity length of hospital stay.

Inpatient length of stay (LOS) is often considered as a proxy of hospital resource consumption. Using statewide obstetrical delivery data, a two-component Poisson mixture model provides a reasonable fit to the heterogeneous LOS distribution. Adopting the generalized linear mixed model (GLMM) approach, random effects are introduced to the two-component Poisson mixture regression model to account for the inherent correlation of patients clustered within hospitals. An EM algorithm is developed for the joint estimation of regression coefficients and variance component parameters. Related diagnostic measures for assessing model adequacy are derived. When applying the method to analyse maternity LOS, appropriate risk factors for the short-stay and long-stay subgroups can be identified from the respective Poisson components. In addition, predicted random hospital effects enable the comparison of relative efficiencies among hospitals after adjustment for patient case-mix and health provision characteristics.

Adult↗

Joint models for efficient estimation in proportional hazards regression models.

In survival studies, information lost through censoring can be partially recaptured through repeated measures data which are predictive of survival. In addition, such data may be useful in removing bias in survival estimates, due to censoring which depends upon the repeated measures. Here we investigate joint models for survival T and repeated measurements Y, given a vector of covariates Z. Mixture models indexed as f (T/Z) f (Y/T,Z) are well suited for assessing covariate effects on survival time. Our objective is efficiency gains, using non-parametric models for Y in order to avoid introducing bias by misspecification of the distribution for Y. We model (T/Z) as a piecewise exponential distribution with proportional hazards covariate effect. The component (Y/T,Z) has a multinomial model. The joint likelihood for survival and longitudinal data is maximized, using the EM algorithm. The estimate of covariate effect is compared to the estimate based on the standard proportional hazards model and an alternative joint model based estimate. We demonstrate modest gains in efficiency when using the joint piecewise exponential joint model. In a simulation, the estimated efficiency gain over the standard proportional hazards model is 6.4 per cent. In clinical trial data, the estimated efficiency gain over the standard proportional hazards model is 10.2 per cent.

Algorithms↗

Robustness of sample size re-estimation procedure in clinical trials (arbitrary populations).

In clinical trials, one of the main questions that is being asked is how many additional observations, if any, are needed beyond those originally planned. In a two-treatment double-blind clinical experiment, one is interested in testing the null hypothesis of equality of the means against one-sided alternative when the common variance sigma2 is unknown. We wish to determine the required total sample size when the error probabilities alpha and beta are specified at a predetermined alternative. Shih provided a two-stage procedure which is an extension of Stein's one-sample procedure, assuming normal response. He estimates sigma2 by the method of maximum likelihood via the EM algorithm and carries out a simulation study in order to evaluate the effective level of significance and the power. The author proposed a closed-form estimator for sigma2 and showed analytically that the difference between the effective and nominal levels of significance is negligible and that the power exceeds 1-beta when the initial sample size is large. Here we consider responses from arbitrary distributions in which the mean and the variance are not functionally related and show that when the initial sample size is large, the conclusions drawn previously by the author still hold. The effective coverage probability of a fixed-width interval is also evaluated. Proofs of certain assertions are deferred to the Appendix.

Clinical Trials as Topic↗

A classification of Scottish infants using latent class analysis.

This paper illustrates the use of latent class analysis to classify 50,000 infants into a small number of classes or case types, as a preliminary to a study of the allocation of neonatal hospital resources throughout Scotland. Information, extracted from a detailed neonatal discharge record, was summarized by 11 clinical and diagnostic catagorical variables. Statistical models incorporating 1 to 6 latent classes were then estimated using the EM algorithm. The 4 class model was chosen because it provided a good description of the data and the resulting classes had a medical interpretation. The factors influencing the choice of model are discussed and goodness of fit tests are presented. The stability of the classes was also investigated using random halves of the data and an earlier comparable data set.

Classification↗

The use of an extended baseline period in the evaluation of treatment in a longitudinal Duchenne muscular dystrophy trial.

A trial of Duchenne muscular dystrophy involved tracking boys of all ages through a one-year baseline period, followed by a one-year trial of leucine versus placebo treatment. In this paper we develop a model for a total-muscle-strength score that uses the data of the extended baseline period in the evaluation of the leucine treatment. The model is based on a polynomial growth curve in age whose coefficients can vary according to treatment or phase. Maximum likelihood estimates of the parameters of the model are obtained from use of the EM algorithm. We propose tests for the adequacy of the model as well as for treatment effects. A quadratic model appears the most parsimonious fit to the data and there is no evidence of any leucine effect on scores. We examine the asymptotic power of the test for treatment effect and compare it with that of a simpler analysis.

Adolescent↗

A comparative study of three methods for analysing longitudinal pulmonary function data.

We compare three methods of longitudinal analysis of pulmonary function data. Our data set is taken from a study of exposure to toluene diisocyanate (TDI) vapours in a new manufacturing plant. The first two methods are a two-stage weighted regression method and maximum likelihood estimation via the EM algorithm, and these give very similar results. The third method, regression with an autoregressive error structure, was not successfully implemented, and in our view needs better documentation.

Adult↗

The analysis of titration studies in phase III clinical trials.

Clinical trials commonly employ the titration design for certain drugs such as antihypertensives. In a Phase III trial the design has purposes distinct from those of a Phase I or II trial, as well as from those of a trial with a parallel design. In this paper we compare the titration design with the usual parallel design in their respective purposes for Phase III trials, explore the relevant questions addressed, and examine typical data from such trials. We also discuss work which focuses primarily on the Phase I or II titration trials. We formulate the problem in the framework of one-way contingency table augmented with incomplete data and obtain the maximum likelihood estimates of the parameters and their estimated variances/covariances via the EM algorithm. An example of a Phase III study of an antihypertensive agent illustrates the proposed procedure.

Analysis of Variance↗

Adjusting for age-related competing mortality in long-term cancer clinical trials.

Mortality related to causes other than the treated disease may have a significant impact on overall survival in long-term clinical trials. We present a model that adjusts for age-related competing mortality when cause of death is missing or only partially available. Through use of a piecewise exponential survival model, we extend relative survival methods to continuous follow-up data, allowing the competing mortality to differ from that of the general population by a scale parameter. An EM algorithm provides a simple way to compute the maximum likelihood estimators (MLEs) and to test hypotheses using widely available software. We compare the bias and relative efficiency of this model to a piecewise exponential Cox model for overall survival. Theoretical results are confirmed by simulations and illustrated with data from a clinical trial in colorectal cancer. This example also shows how age-related and disease-related mortality can be confounded in an analysis of overall survival. We conclude with a discussion of the advantages and disadvantages of the model.

Adolescent↗

A method of non-parametric back-projection and its application to AIDS data.

The method of back-projection has been used to estimate the unobserved past incidence of infection with the human immunodeficiency virus (HIV) and to obtain projections of future AIDS incidence. Here a new approach to back-projection, which avoids parametric assumptions about the form of the HIV infection intensity, is described. This approach gives the data greater opportunity to determine the shape of the estimated intensity function. The method is based on a modification of an EM algorithm for maximum likelihood estimation that incorporates smoothing of the estimated parameters. It is easy to implement on a computer because the computations are based on explicit formulae. The method is illustrated with applications to AIDS data from Australia, U.S.A. and Japanese haemophiliacs.

Acquired Immunodeficiency Syndrome↗

Computational aspects of analysing random effects/longitudinal models.

Random effects and longitudinal models are becoming increasingly popular in the analysis of many types of data, including medical and biopharmaceutical, because of their richness and flexibility. They can be, however, difficult to fit using traditional statistical tools. Fortunately, there now exists a burgeoning collection of newer computational methods that can be applied to draw inferences with such models. This review attempts to provide an introduction to some of these techniques by describing them as extensions of the EM algorithm, currently a standard tool for the analysis of longitudinal and random effects models. For clarity of exposition, the extensions are classified into three types: large-sample iterative; large-sample simulation, and small-sample simulation.

Longitudinal Studies↗

Methods for the analysis of informatively censored longitudinal data.

This paper describes the problem of informative censoring in longitudinal studies where the primary outcome is rate of change in a continuous variable. Standard approaches based on the linear random effects model are valid only when the data are missing in a non-ignorable fashion. Informative censoring, which is a special type of non-ignorably missing data, occurs when the probability of early termination is related to an individual subject's true rate of change. When present, informative censoring causes bias in standard likelihood-based analyses, as well as in weighted averages of individual least-squares slopes. This paper reviews several methods proposed by others for analysis of informatively censored longitudinal data, and outlines a new approach based on a log-normal survival model. Maximum likelihood estimates may be obtained via the EM algorithm. Advantages of this approach are that it allows general unbalanced data caused by staggered entry and unequally-timed visits, it utilizes all available data, including data from patients with only a single measurement, and it provides a unified method for estimating all model parameters. Issues related to study design when informative censoring may occur are also discussed.

Linear Models↗