PubMed HealthSearch

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

An approximate likelihood procedure for censored data.

An approximate likelihood procedure is suggested for the estimation of the parameters of the density of a single homogeneous sample subject to right censoring. Two examples are given involving the gamma distribution. The method is typically consistent and although it is always inefficient, the efficiency loss is second-order in the degree of censoring when this is small. The relation of the techniques to the results of Reid (1981, Annals of Statistics 9, 78-92) on influence functions for censored data, to the EM algorithm, and to the nonparametric regression techniques of Miller (1976, Biometrika 63, 449-464) and Buckley and James (1979, Biometrika 66, 429-436) are indicated. Simple estimates of standard error are obtained.

Animals

A Bayesian approach to nonlinear random effects models.

Nonlinear random effects models are considered from the Bayesian point of view. The method of analysis follows closely that of Lindley and Smith (1972, Journal of the Royal Statistical Society, Series B 34, 1-42). The numerical method is related to the EM algorithm.

Analysis of Variance

Mixture distributions in psychiatric research.

This paper describes the application of Gaussian mixture distributions to biological marker research in psychiatry. Mixtures of univariate and multivariate normal distributions can be used to determine if diagnostically similar psychiatric patients belong to biologically distinct subpopulations. The resulting biological subtypes may be important in understanding the etiology of psychiatric disorders. The general model and estimation procedure are described (EM algorithm; Dempster, Laird and Rubin 1977). The method is illustrated using two examples of biological data: (1) red cell membranes and monoamine oxidase activity data in normal individuals having no family history of psychiatric illness, the first-degree relatives of bipolar depressed patients and a heterogeneous patient population; and (2) smooth pursuit eye movements that classify relatives of schizophrenics, nonschizophrenics and normal controls into biologically distinct populations.

Bipolar Disorder

Random-effects models for serial observations with binary response.

This paper presents a general mixed model for the analysis of serial dichotomous responses provided by a panel of study participants. Each subject's serial responses are assumed to arise from a logistic model, but with regression coefficients that vary between subjects. The logistic regression parameters are assumed to be normally distributed in the population. Inference is based upon maximum likelihood estimation of fixed effects and variance components, and empirical Bayes estimation of random effects. Exact solutions are analytically and computationally infeasible, but an approximation based on the mode of the posterior distribution of the random parameters is proposed, and is implemented by means of the EM algorithm. This approximate method is compared with a simpler two-step method proposed by Korn and Whittemore (1979, Biometrics 35, 795-804), using data from a panel study of asthmatics originally described in that paper. One advantage of the estimation strategy described here is the ability to use all of the data, including that from subjects with insufficient data to permit fitting of a separate logistic regression model, as required by the Korn and Whittemore method. However, the new method is computationally intensive.

Air Pollution

Nonparametric estimation of the distribution of time to onset for specific diseases in survival/sacrifice experiments.

This paper concerns the analysis of an animal survival/sacrifice experiment designed to investigate the incidence of a particular disease of interest. The disease is assumed to be irreversible, and detectable only at death, for example by a necropsy. Each observation can be of one of three types: (i) death caused by the disease, (ii) death from a competing cause such as sacrifice, with the disease present, or (iii) death with the disease absent. A two-dimensional EM algorithm is proposed for the nonparametric maximum likelihood estimation of the distributions of the time to onset and of the time to death from the disease. These can be compared with nonparametric estimators recently proposed by Kodell , Shaw and Johnson (1982, Biometrics 38, 43-58) and by Dinse and Lagakos (1982, Biometrics 38, 921-932). A slight modification of the algorithm permits the construction of likelihood-based interval estimates of quantiles of the distributions. Some extensions and generalizations are indicated.

Age Factors

Nonparametric estimation of a multivariate distribution in the presence of censoring.

This paper presents examples of situations in which one wishes to estimate a multivariate distribution from data that may be right-censored. A distinction is made between what we term 'homogeneous' and 'heterogeneous' censoring. It is shown how a multivariate empirical survivor function must be constructed in order to be considered a (nonparametric) maximum likelihood estimate of the underlying survivor function. A closed-form solution, similar to the product-limit estimate of Kaplan and Meier, is possible with homogeneous censoring, but an iterative method, such as the EM algorithm, is required with heterogeneous censoring. An example is given in which an anomaly is produced if censored multivariate data are analyzed as a series of univariate variables; this anomaly is shown to disappear if the methods of this paper are used.

Analysis of Variance

Random-effects models for longitudinal data.

Models for the analysis of longitudinal data must recognize the relationship between serial observations on the same unit. Multivariate models with general covariance structure are often difficult to apply to highly unbalanced data, whereas two-stage random-effects models can be used easily. In two-stage models, the probability distributions for the response vectors of different individuals belong to a single family, but some random-effects parameters vary across individuals, with a distribution specified at the second stage. A general family of models is discussed, which includes both growth models and repeated-measures models as special cases. A unified approach to fitting these models, based on a combination of empirical Bayes and maximum likelihood estimation of model parameters and using the EM algorithm, is discussed. Two examples are taken from a current epidemiological study of the health effects of air pollution.

Air Pollution

A multiplicative random effects model for meta-analysis with application to estimation of admixture component.

A multiplicative random effects model is constructed for combining results from different studies. An EM algorithm is developed for the maximum likelihood estimator under the Empirical Bayes framework. The method is applied to estimation of the admixture component in admixed populations. The admixture estimates can be useful in genetic epidemiology and DNA fingerprinting.

Black or African American

Semi-parametric estimation in failure time mixture models.

A mixture model is an attractive approach for analyzing failure time data in which there are thought to be two groups of subjects, those who could eventually develop the endpoint and those who could not develop the endpoint. The proposed model is a semi-parametric generalization of the mixture model of Farewell (1982). A logistic regression model is proposed for the incidence part of the model, and a Kaplan-Meier type approach is used to estimate the latency part of the model. The estimator arises naturally out of the EM algorithm approach for fitting failure time mixture models as described by Larson and Dinse (1985). The procedure is applied to some experimental data from radiation biology and is evaluated in a Monte Carlo simulation study. The simulation study suggests the semi-parametric procedure is almost as efficient as the correct fully parametric procedure for estimating the regression coefficient in the incidence, but less efficient for estimating the latency distribution.

Animals

EM mixed model analysis of data from informatively censored normal distributions.

Maximum likelihood techniques using the EM algorithm are applied to correlated normally distributed survival data from a placebo-controlled, double-blind, dose-ranging crossover study to assess the short-term efficacy of an antianginal drug in patients with chronic stable angina. Censoring was informative and nonterminal and was not due to death or withdrawal from the study. Unlike previous approaches these techniques are mathematically and computationally tractable, do not require computations of high-dimensional integrals, and do not require the inversion of large matrices.

Algorithms

Change points in the series of T4 counts prior to AIDS.

The absolute number of T4 cells has been established as an important clinical marker of disease progression to acquired immunodeficiency syndrome (AIDS) in persons infected with human immunodeficiency virus (HIV). Series of T4 counts are analyzed from the 131 homosexual men who entered the New York Blood Center Study in 1984, mostly seropositive for HIV, and who developed AIDS as participants by 1990. These series exhibit a gradual decline of the log(T4) count followed by a more rapid decline close to the time of the development of AIDS. Empirical Bayes and hierarchical Bayes change point models are proposed to estimate the distribution of the time before AIDS when this rapid decline begins. Results using the EM Algorithm and Markov chain Monte Carlo indicate that the mean change point occurs approximately 1 year before diagnosis with a standard deviation of 9 months. Detection of a change point may indicate that an AIDS diagnosis is increasingly likely for an individual HIV-positive but AIDS-free.

Acquired Immunodeficiency Syndrome

Analysing incomplete longitudinal binary responses: a likelihood-based approach.

In this paper, we describe a likelihood-based method for analysing balanced but incomplete longitudinal binary responses that are assumed to be missing at random. Following the approach outlined in Zhao and Prentice (1990, Biometrika 77, 642-648), we focus on "marginal models" in which the marginal expectation of the response variable is related to a set of covariates. The association between binary responses is modelled in terms of conditional log odds-ratios. We describe a set of scoring equations for jointly estimating both the marginal parameters and the conditional association parameters. An outline of the EM algorithm used to obtain the maximum likelihood estimates is presented. This approach yields valid and efficient estimates when the responses are missing at random, but not necessarily missing completely at random. An example, using data from the Muscatine Coronary Risk Factor Study, is presented to illustrate this methodology.

Age Factors

Models relating the timing of intercourse to the probability of conception and the sex of the baby.

The probability that conception will result from intercourse on particular days of the menstrual cycle is of biological, demographic, and personal interest to many. We describe an existing model for conception as related to the timing of intercourse, and develop a method for fitting it by means of the expectation-maximization (EM) algorithm, using widely available software (GLIM). We then generalize the model to allow for effects on its parameters due to reproductively toxic exposures. A further extension allows one to address the question of whether the pattern of intercourse in a conception cycle has an influence on the sex of the resulting baby. We illustrate the methods by application to data from a prospective study of couples trying to begin a pregnancy.

Abortion, Spontaneous

Monte Carlo estimation of mixed models for large complex pedigrees.

In human quantitative genetics, computational complexity restricts the current methods for estimation of mixed models that include major gene effects to data on small pedigrees. However, large complex pedigrees are not uncommon in practice. Also, large pedigrees tend to provide more information on genetic transmission and are more genetically homogeneous than a pooled sample of many nuclear families. We present a Monte Carlo method, using jointly the EM algorithm and the Gibbs sampler, for estimation of mixed models. The approach also provides a Monte Carlo estimate of the asymptotic variance-covariance matrix of the parameters. The methods are conceptually simple, easy to implement, and can handle multiple heritable/nonheritable random components. A numerical example is given to illustrate the methods.

Algorithms

Estimating incidence and diagnostic error rates for bivariate progressive processes.

Estimating the times until incidences of bivariate progressive processes that are categorical is a common problem in ophthalmology, audiology, pulmonary medicine, and other fields of medical research. We consider study designs in which diagnoses of subject's bivariate status are performed repeatedly across time and when diagnosis is subject to error. In such situations, error confounds the interpretation of the time until an event. A composite model is proposed for parameterizing both the incidence and error distributions, which allows for correlation between sites with respect to both incidence and diagnostic error. An EM algorithm is described for this model, which allows categorical covariates for both incidence and error. The methodology is applied to two examples. The first represents a situation in which bivariate incidence and error can reasonably be assumed symmetric: prospective data concerning the development of ocular lens opacities in a large pharmaceutical clinical trial. The second example represents a situation in which bivariate incidence and error may not be symmetric: clinical evaluations of sexual maturation status with respect to two different anatomical indices in the Cooperative Study of Sickle Cell Disease. The methodology described in this paper is used, in each case, to estimate incidence, characterize error rates, and assess bivariate correlations.

Adolescent

Analysis of infectious disease data from partner studies with unknown source of infection.

Partner studies are useful for estimating the transmission probabilities of infectious diseases. However, it is often not known which partner was the source of the infection (the index case). The objective of this paper is to develop statistical methods for analyzing partner studies when it is uncertain which partners acquired the infection from sources outside the partnership. The approach involves simultaneously modelling the probability of acquisition of infection from outside the partnership, and the probability of transmission within the partnership as a function of covariates. An EM algorithm is presented. Efficiency and simulation results are given in some special situations involving heterosexual transmission studies. In heterosexual partner studies, the methods depend crucially on the availability of a covariate that provides information about which partner was the likely source of infection.

Algorithms

Quantitative SPECT reconstruction of iodine-123 data.

Many clinical and research studies in nuclear medicine require quantitation of iodine-123 (123I) distribution for the determination of kinetics or localization. The objective of this study was to implement several reconstruction methods designed for single-photon emission computed tomography (SPECT) using 123I and to evaluate their performance in terms of quantitative accuracy, image artifacts, and noise. The methods consisted of four attenuation and scatter compensation schemes incorporated into both the filtered backprojection/Chang (FBP) and maximum likelihood-expectation maximization (ML-EM) reconstruction algorithms. The methods were evaluated on data acquired of a phantom containing a hot sphere of 123I activity in a lower level background 123I distribution and nonuniform density media. For both reconstruction algorithms, nonuniform attenuation compensation combined with either scatter subtraction or Metz filtering produced images that were quantitatively accurate to within 15% of the true value. The ML-EM algorithm demonstrated quantitative accuracy comparable to FBP and smaller relative noise magnitude for all compensation schemes.

Humans

Estimation of parameters and missing values under a regression model with non-normally distributed and non-randomly incomplete data.

We carried out a simulation study to compare the performance of three algorithms (complete cases, ALLVALUE, and expectation maximization, EM) in estimating regression parameters and missing values for situations that have varying amounts of missing data, distributions (normal, mixture of normals and lognormal), patterns of incomplete data (random, related and censored), and degrees of correlational structure among the dependent and independent variables. We found that the EM and complete cases algorithms performed equally well regardless of the correlational structure, when the percentage of incomplete data was only 5 per cent. When this percentage increased to 25 per cent, the EM algorithm was generally best for estimation, but the complete cases algorithm was safe and conservative. This finding may be attributed to the study design, which required that the slopes be the same in the population of all cases, and in the population of complete cases. In addition, the one-step imputing method (ALLVALUE) was competitive only for situations with weak correlational structure and/or little missing data. In that situation the bias caused with use of all available information was less than that caused with use of only complete cases. On the other hand, for imputation, the EM algorithm performed optimally, even in situations of censored or log-normally distributed data.

Algorithms