PubMed HealthSearch

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Monte Carlo estimation of variance component models for large complex pedigrees.

Variance component models are widely used in animal and plant breeding. In human genetics, they can be used to identify, among other traits associated with the definition of disease, those that have a significant genetic component in their aetiology. In addition, they can be used in genetic counselling. Most of the methods currently proposed for estimating variance component models often involve repeated inversion of large matrices, resulting in intensive computations, large storage requirements, and numerical instability. Consequently, these methods are restricted to data on nuclear families, to small pedigrees, or to designed pedigrees of simple form. In this paper, the authors propose a method for estimating variance component models for large complex pedigrees using jointly the EM algorithm and the Gibbs sampler. The method can handle variance component models with multiple variance components, without the need for repeated inversion of large matrices even on large complex pedigrees. The method is conceptually simple, numerically stable, and easy to implement.

Algorithms

LeRec: a NN/HMM hybrid for on-line handwriting recognition.

We introduce a new approach for on-line recognition of handwritten words written in unconstrained mixed style. The preprocessor performs a word-level normalization by fitting a model of the word structure using the EM algorithm. Words are then coded into low resolution "annotated images" where each pixel contains information about trajectory direction and curvature. The recognizer is a convolution network that can be spatially replicated. From the network output, a hidden Markov model produces word scores. The entire system is globally trained to minimize word-level errors.

Algorithms

Statistical models for PET and SPECT data.

This article outlines the statistical developments that have taken place in emission tomography during the past decade or so. We discuss the statistical aspects of the modelling of the projection data and define the additive Poisson regression model. This leads to the use of the method of maximum likelihood as a means of estimating the underlying isotope concentration within a given region of a patient's body, and to the use of the EM algorithm to compute the reconstruction. The need for the regulation of the maximum likelihood solution is tackled using Bayesian techniques. A number of algorithms for the computation of regularized solutions are outlined. The issue of parameter estimation is discussed and some open issues are mentioned.

Algorithms

A comparison of continuous- and discrete- time three-state models for rodent tumorigenicity experiments.

The three-state illness-death model provides a useful way to characterize data from a rodent tumorigenicity experiment. Most parametrizations proposed recently in the literature assume discrete time for the death process and either discrete or continuous time for the tumor onset process. We compare these approaches with a third alternative that uses a piecewise continuous model on the hazards for tumor onset and death. All three models assume proportional hazards to characterize tumor lethality and the effect of dose on tumor onset and death rate. All of the models can easily be fitted using an Expectation Maximization (EM) algorithm. The piecewise continuous model is particularly appealing in this context because the complete data likelihood corresponds to a standard piecewise exponential model with tumor presence as a time-varying covariate. It can be shown analytically that differences between the parameter estimates given by each model are explained by varying assumptions about when tumor onsets, deaths, and sacrifices occur within intervals. The mixed-time model is seen to be an extension of the grouped data proportional hazards model [Mutat. Res. 24:267-278 (1981)]. We argue that the continuous-time model is preferable to the discrete- and mixed-time models because it gives reasonable estimates with relatively few intervals while still making full use of the available information. Data from the ED01 experiment illustrate the results.

Animals

Restricted maximum likelihood procedures for the estimation of additive and nonadditive genetic variances and covariances in multibreed populations.

Restricted maximum-likelihood procedures were developed to estimate additive and nonadditive genetic and environmental covariances for multiple traits in multibreed populations. The computational procedure follows the expectation-maximization (EM) algorithm, where the set of equations in the maximization step is solved by successive approximations. This computational procedure does not guarantee convergence to a symmetric positive-definite covariance matrix. Thus, computer programs will need to incorporate restrictions in the maximization step to ensure positive definiteness of each covariance matrix. Additive genetic and environmental covariances were modeled in subclass form (zeros and ones in the design matrices). Nonadditive genetic covariances were modeled in regression form (any value between and including zero and one in the design matrices). Computational requirements will be larger than for intrabreed analyses. Appropriate simplifying assumptions and numerical techniques (e.g., sparse and iterative numerical techniques) will be required for the implementation of these multibreed covariance estimation procedures. Number of iterations (5 to 12) and computing times (57 to 113 min) to achieve convergence when estimating 21 genetic and environmental covariances in five small simulated multibreed data sets (two breeds, 25,200 to 50,400 calves, 120 to 135 unrelated bulls) suggest that these procedures are computationally feasible.

Algorithms

Estimation of variance and covariance components to determine heritabilities and repeatability of weaning weight in American Simmental cattle.

Components of (co)variance for weaning weight were estimated from field data provided by the American Simmental Association. These components were obtained for the observational components of variance corresponding to a sire, maternal grandsire, and dam within maternal grandsire model. From these estimates, direct additive genetic variance (Sigma2A), maternal additive genetic variance (Sigma2M), covariance between direct and maternal additive genetic effects (SigmaAM), variance of permanent environment(Sigma2pe) and temporary environment variance(Sigma2te) were determined. A procedure to approximate restricted maximum likelihood (REML) estimates of the observational components of variance based on the expectation-maximization (EM) algorithm is described. From these results, phenotypic variance ( ) of weaning weight was 667.88 kg2. Values forSigma2A, Sigma2M, Sigma2pe and Sigma2te were 79,30,58,38,49.45, and 469.97 kg2, respectively. Genetic correlation between direct and maternal additive genetic effects was .16.

Animals

Exact likelihood evaluation in a Markov mixture model for time series of seizure counts.

This paper provides an alternative to Albert's (1991), Biometrics 47, 1371-1381) approximation to the E-step when using the EM algorithm for parameter estimation in Markov mixture models. Use of a recursive algorithm of Baum et al. (1970, Annals of Mathematical Statistics 41, 164-171) results in exact evaluation of the likelihood, optimal parameter estimates, and very efficient computation. Applications to time series of seizure counts and fetal movements clearly show the advantages of this exact approach.

Algorithms

The effect of diagnostic misclassification on non-cancer and cancer mortality dose response in A-bomb survivors.

We used the EM algorithm in the context of a joint Poisson regression analysis of cancer and non-cancer mortality in the Radiation Effects Research Foundation (RERF) Life Span Study (LSS) to assess whether the observed increased risk of non-cancer death due to radiation exposure (Shimizu et al., RERF Technical Report 02-91, 1991) can be attributed solely to misclassification of cancer as non-cancer on death certificates. We show that greater levels of dose-independent misclassification than are indicated by a series of autopsies conducted on a subset of LSS members would be required to explain the non-cancer dose response, but that a relatively small amount of dose-dependence in the misclassification of cancer would explain the result. The adjustment for misclassification also results in higher risk estimates for cancer mortality. We review applications of similar statistical methods in other contexts and discuss extensions of the methods to more than two causes of death.

Age Factors

Fitting mixture models to birth weight data: a case study.

Birth weights by gestational age are compared in two birth cohorts from Northern Finland, the first from 1966 and the second from 1985-1986. A curious fact in the data is that mean birth weight before the 39th week was lower in the latter series although the mean birth weight for the total series was higher. Similar findings have been reported in other series. A mixture model with the nonparametric regression function is proposed for studying the hypothesis that the difference was caused by more frequent gross errors in gestational assessment in the earlier cohort. The probability of an error in gestational assessment then greatly depends on the observed gestational age, which makes the mixture model nonstandard. Maximum likelihood solutions to the parameters in the proposed model were computed employing the general expectation-maximization (EM) algorithm. A technique for studying the effect of errors on the intrauterine weight gain curve is proposed and applied to our two birth cohorts. The risk of underestimation of gestational age seems to be larger in the previous series and the differences between the growth curves almost totally vanish when "corrected" by means of the mixture model.

Algorithms

Nonparametric estimation of the size-metastasis relationship in solid cancers.

This paper is concerned with the relationship between the occurrence of metastases and the size of primary cancers. We consider two probabilistic characterizations of this relationship. First is the distribution function of tumor sizes at the point of metastatic transition; second is the probability that detectable metastases are present when the cancer comes to medical attention. The equation relating these two functions is developed and conditions for their being identical are explored. Since the tumor size at the point of metastasis is not usually observable, estimation of the first distribution requires the use of the EM algorithm. Nonparametric methods of estimating both functions are explored, with attention to the fact that tumors often fail to be measured, particularly those that are known to be metastatic. The methods are applied to the estimation of primary tumor size at the point of distant metastasis in lung cancer (epidermoid and adenocarcinoma) and colorectal cancer and at the point of nodal metastasis in breast cancer. Monte Carlo experiments confirm that the bias inherent in the methodology is acceptably small.

Adenocarcinoma

A two-state Markov mixture model for a time series of epileptic seizure counts.

This paper discusses a model for a time series of epileptic seizure counts in which the mean of a Poisson distribution changes according to an underlying two-state Markov chain. The EM algorithm (Dempster, Laird, and Rubin, 1977, Journal of the Royal Statistical Society, Series B 39, 1-38) is used to compute maximum likelihood estimators for the parameters of this two-state mixture model and extensions are made allowing for nonstationarity. The model is illustrated using daily seizure counts for patients with intractable epilepsy and results are compared with a simple Poisson distribution and Poisson regressions. Some simulation results are also presented to demonstrate the feasibility of this model.

Biometry

Mixture models for continuous data in dose-response studies when some animals are unaffected by treatment.

A mixture model is described for dose-response studies where measurements on a continuous variable suggest that some animals are not affected by treatment. The model combines a logistic regression on dose for the probability an animal will "respond" to treatment with a linear regression on dose for the mean of the responders. Maximum likelihood estimation via the EM algorithm is described and likelihood ratio tests are used to distinguish between the full model and meaningful reduced-parameter versions. Use of the model is illustrated with three real-data examples.

Algorithms

Fitting mixture distributions to phenylthiocarbamide (PTC) sensitivity.

A technique for fitting mixture distributions to phenylthiocarbamide (PTC) sensitivity is described. Under the assumptions of Hardy-Weinberg equilibrium, a mixture of three normal components is postulated for the observed distribution, with the mixing parameters corresponding to the proportions of the three genotypes associated with two alleles A and a acting at a single locus. The corresponding genotypes AA, Aa, and aa are then considered to have separate means and variances. This paper is concerned with estimating the parameters of the model, and their standard errors, by using an application of the EM algorithm. This technique also caters for the fact that the sensitivity measurements are only known to lie between the endpoints of certain intervals and that the exact measurement of the attribute is not possible.

Algorithms

Log-linear models in the analysis of disease prevalence data from survival/sacrifice experiments.

This paper considers the problem of analyzing disease prevalence data from survival experiments in which there may also be some serial sacrifice. The assumptions needed for "standard" analyses are reviewed in the context of a general model recently proposed by the authors. This model is then reparametrized in log-linear form, and a generalized EM algorithm is utilized to obtain maximum likelihood estimates of the parameters for a broad class of unsaturated models. Tests based on the relative likelihood are proposed to investigate the effects of treatment, time, and the presence of other diseases on the prevalences and lethalities of specific diseases of interest. An example is given, using data from a large experiment to investigate the effects of low-level radiation on laboratory mice. Finally, some possible directions for future research are indicated.

Animals

Mixed-model analysis of a censored normal distribution with reference to animal breeding.

A mixed-model procedure for analysis of censored data assuming a multivariate normal distribution is described. A Bayesian framework is adopted which allows for estimation of fixed effects and variance components and prediction of random effects when records are left-censored. The procedure can be extended to right- and two-tailed censoring. The model employed is a generalized linear model, and the estimation equations resemble those arising in analysis of multivariate normal or categorical data with threshold models. Estimates of variance components are obtained using expressions similar to those employed in the EM algorithm for restricted maximum likelihood (REML) estimation under normality.

Analysis of Variance

The beta-geometric distribution applied to comparative fecundability studies.

A convenient measure of fecundability is time (number of menstrual cycles) required to achieve pregnancy. Couples attempting pregnancy are heterogeneous in their per-cycle probability of success. If success probabilities vary among couples according to a beta distribution, then cycles to pregnancy will have a beta-geometric distribution. Under this model, the inverse of the cycle-specific conception rate is a linear function of time. Data on cycles to pregnancy can be used to estimate the beta parameters by maximum likelihood in a straightforward manner with a package such as GLIM. The likelihood ratio test can thus be employed in studies of exposures that may impair fecundability. Covariates are incorporated in a natural way. The model is illustrated by applying it to data on cycles to pregnancy in smokers and nonsmokers, with adjustment for covariates. For a cross-sectional study, when length-biased sampling is taken into account, the pre-interview attempt time is shown to follow a beta-geometric distribution, so that the same methods of analysis can be applied even though all of the available data are right-censored. For a cohort followed prospectively, there will be some couples enrolled whose fecundability is effectively 0, and for such applications, the beta could be considered to be contaminated by a distribution degenerate at 0. The mixing parameter (proportion sterile) can be estimated by application of the expectation-maximization (EM) algorithm. This, too, can be carried out using GLIM.

Biometry

An approximate likelihood procedure for censored data.

An approximate likelihood procedure is suggested for the estimation of the parameters of the density of a single homogeneous sample subject to right censoring. Two examples are given involving the gamma distribution. The method is typically consistent and although it is always inefficient, the efficiency loss is second-order in the degree of censoring when this is small. The relation of the techniques to the results of Reid (1981, Annals of Statistics 9, 78-92) on influence functions for censored data, to the EM algorithm, and to the nonparametric regression techniques of Miller (1976, Biometrika 63, 449-464) and Buckley and James (1979, Biometrika 66, 429-436) are indicated. Simple estimates of standard error are obtained.

Animals

A Bayesian approach to nonlinear random effects models.

Nonlinear random effects models are considered from the Bayesian point of view. The method of analysis follows closely that of Lindley and Smith (1972, Journal of the Royal Statistical Society, Series B 34, 1-42). The numerical method is related to the EM algorithm.

Analysis of Variance