PubMed HealthSearch

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

A maximum likelihood method for region-of-interest evaluation in emission tomography.

A maximum likelihood (ML) estimation method, called the ML-ROI algorithm, is presented for the calculation of region-of-interest (ROI) values from emission tomography scans. The EM algorithm is used to directly estimate ROI values from tomographic projection data given the location, size, and shape of all ROIs. The algorithm requires for the specification of a detailed model of the physical factors contributing to projection measurements including resolution and attenuation. The ML-ROI algorithm also provides an estimate of the variability of the ROI estimator (covariance matrix). The algorithm was tested with simulation and phantom data and compared with ROI estimation strategies using filtered backprojection (FBP) images. The ML-ROI estimates were unbiased, i.e., the partial volume effect was eliminated. Except for regions smaller than the detector resolution, the variability of the ML estimates was comparable to or less than the biased FBP estimators. Computation time for the ML-ROI algorithm was between 5 and 10 s/iteration. An evaluation of the sensitivity of the algorithm to misdefinition of the location and size of the ROIs was also performed.

Humans

High diversity of alpha-globin haplotypes in a Senegalese population, including many previously unreported variants.

RFLP haplotypes at the alpha-globin gene complex have been examined in 190 individuals from the Niokolo Mandenka population of Senegal: haplotypes were assigned unambiguously for 210 chromosomes. The Mandenka share with other African populations a sample size-independent haplotype diversity that is much greater than that in any non-African population: the number of haplotypes observed in the Mandenka is typically twice that seen in the non-African populations sampled to date. Of these haplotypes, 17.3% had not been observed in any previous surveys, and a further 19.1% have previously been reported only in African populations. The haplotype distribution shows clear differences between African and non-African peoples, but this is on the basis of population-specific haplotypes combined with haplotypes common to all. The relationship of the newly reported haplotypes to those previously recorded suggests that several mutation processes, particularly recombination as homologous exchange or gene conversion, have been involved in their production. A computer program based on the expectation-maximization (EM) algorithm was used to obtain maximum-likelihood estimates of haplotype frequencies for the entire data set: good concordance between the unambiguous and EM-derived sets was seen for the overall haplotype frequencies. Some of the low-frequency haplotypes reported by the estimation algorithm differ greatly, in structure, from those haplotypes known to be present in human populations, and they may not represent haplotypes actually present in the sample.

Genetic Variation

Empirical estimation of a distribution function with truncated and doubly interval-censored data and its application to AIDS studies.

In this paper we discuss the non-parametric estimation of a distribution function based on incomplete data for which the measurement origin of a survival time or the date of enrollment in a study is known only to belong to an interval. Also the survival time of interest itself is observed from a truncated distribution and is known only to lie in an interval. To estimate the distribution function, a simple self-consistency algorithm, a generalization of Turnbull's (1976, Journal of the Royal Statistical Association, Series B 38, 290-295) self-consistency algorithm, is proposed. This method is then used to analyze two AIDS cohort studies, for which direct use of the EM algorithm (Dempster, Laird and Rubin, 1976, Journal of the Royal Statistical Association, Series B 39, 1-38), which is computationally complicated, has previously been the usual method of the analysis.

Acquired Immunodeficiency Syndrome

Estimating polygenic models for multivariate data on large pedigrees.

We have developed algorithms for the likelihood estimation of additive genetic models for quantitative traits on large pedigrees. The approach uses the expectation L-maximization (EM) algorithm, but avoids intensive computation. In this paper, we focus on extensions of previous work to the case of multivariate data. We exemplify the approach by analyses of bivariate data on a four-generation, 949-member pedigree of the snail Lymnaea elodes, and on a three-generation pedigree of the guppy Poecilia reticulata containing about 400 individuals.

Algorithms

Stochastic models for heterogeneous DNA sequences.

The composition of naturally occurring DNA sequences is often strikingly heterogeneous. In this paper, the DNA sequence is viewed as a stochastic process with local compositional properties determined by the states of a hidden Markov chain. The model used is a discrete-state, discrete-outcome version of a general model for non-stationary time series proposed by Kitagawa (1987). A smoothing algorithm is described which can be used to reconstruct the hidden process and produce graphic displays of the compositional structure of a sequence. The problem of parameter estimation is approached using likelihood methods and an EM algorithm for approximating the maximum likelihood estimate is derived. The methods are applied to sequences from yeast mitochondrial DNA, human and mouse mitochondrial DNAs, a human X chromosomal fragment and the complete genome of bacteriophage lambda.

Base Sequence

Probabilistic linkage of large public health data files.

Probabilistic linkage technology makes it feasible and efficient to link large public health databases in a statistically justifiable manner. The problem addressed by the methodology is that of matching two files of individual data under conditions of uncertainty. Each field is subject to error which is measured by the probability that the field agrees given a record pair matches (called the m probability) and probabilities of chance agreement of its value states (called the u probability). Fellegi and Sunter pioneered record linkage theory. Advances in methodology include use of an EM algorithm for parameter estimation, optimization of matches by means of a linear sum assignment program, and more recently, a probability model that addresses both m and u probabilities for all value states of a field. This provides a means for obtaining greater precision from non-uniformly distributed fields, without the theoretical complications arising from frequency-based matching alone. The model includes an iterative parameter estimation procedure that is more robust than pre-match estimation techniques. The methodology was originally developed and tested by the author at the U.S. Census Bureau for census undercount estimation. The more recent advances and a new generalized software system were tested and validated by linking highway crashes to Emergency Medical Service (EMS) reports and to hospital admission records for the National Highway Traffic Safety Administration (NHTSA).

Adult

Simultaneous emission and transmission measurements for attenuation correction in whole-body PET.

UNLABELLED: We describe a methodology for measuring and correcting for attenuation in whole-body PET using simultaneous emission and transmission (SET) measurements. METHODS: The main components of the methodology are: (a) sinogram windowing of low activity (< or = 50 MBq) rotating 68Ge/Ga rod sources, (b) segmented attenuation correction (SAC) and (c) maximum likelihood reconstruction using the ordered subsets EM (OS-EM) algorithm. The methods were implemented on a whole-body positron emission tomograph. Quantitative accuracy and the signal-to-noise ratio (SNR) were measured for a thorax-tumor phantom as functions of acquisition time (range: 2-20 min per position). RESULTS: When a typical rod source activity (200 MBq 68Ge/Ga) was used, emission SNR was 60% lower in simultaneous than in separate measurements. The difference was only 14% when the rods contained 45 MBq 68Ge/Ga. The SNR was further improved by SAC in conjunction with OS-EM reconstruction and the relative gain increased with increasing acquisition time. Quantitative estimates of tumor, liver and lung radioactivity agreed with values obtained from a separate high count measurement to within 8%, independent of acquisition time. CONCLUSION: Attenuation correction of whole-body PET images is feasible using SET measurements. There is good quantitative agreement with conventional methods and increased noise is offset by the use of SAC and OS-EM reconstruction.

Adult

Three-marker phenotypic analysis of lymphocytes based on two-color immunofluorescence using a multinomial model for flow cytometric counts and maximum likelihood estimation.

The simultaneous flow cytometric study of multiple (> or = 3) markers on individual cells is restricted by technical reasons, e.g., the type of flow cytometer, and the availability of (monoclonal) antibodies (mAb) conjugated with the appropriate fluorochromes. However, a n-way classification (n > or = 3) may be derived from multiple 2-way classifications. The 2-way classifications are obtained by the use of only two fluorochromes, but each fluorochrome may be used for the simultaneous labelling of 2 or more mAb. We present a formal statistical model, based on an underlying multinomial distribution for the observed quadrant counts, by which data from such multiple 2-way classifications can be analyzed. The model is restricted to 3-marker phenotypic analyses, but can be extended to n-way classifications (n > 3). Maximum likelihood estimates are obtained by the application of the EM algorithm. The model was tested on 20 samples of peripheral blood mononuclear cells to study the coexpression of CD45RA, CD45RO, and Leu 8 by lymphocyte subsets defined by the CD56 (MHC-unrestricted cytotoxic cells) and CD3 (T cells) markers. Application of the model gave an excellent fit in all but one cases. In addition, the model reduced the effect of inter-assay variation on the estimates and it provided an analysis of consistency over the data, which allows the detection of outliers due to staining errors.

Cell Adhesion Molecules

Hierarchical Multi-Label Classification With Gene-Environment Interactions in Disease Modeling.

In biomedical studies, gene-environment (G-E) interactions have been demonstrated to have important implications for analyzing disease outcomes beyond the main G and main E effects. Many approaches have been developed for G-E interaction analysis, yielding important findings. However, hierarchical multi-label classification, which provides insightful information on disease outcomes, remains unexplored in G-E analysis literature. Moreover, unlabeled data are commonly observed in practical settings but omitted by many existing methods of hierarchical multi-label classification. In this study, we consider a semi-supervised scenario and develop a novel approach for the two-layer hierarchical response with G-E interactions. A two-step penalized estimation is then proposed using an efficient expectation-maximization (EM) algorithm. Simulation shows that it has superior performance in classification and feature selection. The analysis of The Cancer Genome Atlas (TCGA) data on lung cancer demonstrates the practical utility of the proposed method. Overall, this study can fill the important knowledge gap in G-E interaction analysis by providing a widely applicable framework for hierarchical multi-label classification of complex disease outcomes.

Humans

A classification of Scottish infants using latent class analysis.

This paper illustrates the use of latent class analysis to classify 50,000 infants into a small number of classes or case types, as a preliminary to a study of the allocation of neonatal hospital resources throughout Scotland. Information, extracted from a detailed neonatal discharge record, was summarized by 11 clinical and diagnostic catagorical variables. Statistical models incorporating 1 to 6 latent classes were then estimated using the EM algorithm. The 4 class model was chosen because it provided a good description of the data and the resulting classes had a medical interpretation. The factors influencing the choice of model are discussed and goodness of fit tests are presented. The stability of the classes was also investigated using random halves of the data and an earlier comparable data set.

Classification

The use of an extended baseline period in the evaluation of treatment in a longitudinal Duchenne muscular dystrophy trial.

A trial of Duchenne muscular dystrophy involved tracking boys of all ages through a one-year baseline period, followed by a one-year trial of leucine versus placebo treatment. In this paper we develop a model for a total-muscle-strength score that uses the data of the extended baseline period in the evaluation of the leucine treatment. The model is based on a polynomial growth curve in age whose coefficients can vary according to treatment or phase. Maximum likelihood estimates of the parameters of the model are obtained from use of the EM algorithm. We propose tests for the adequacy of the model as well as for treatment effects. A quadratic model appears the most parsimonious fit to the data and there is no evidence of any leucine effect on scores. We examine the asymptotic power of the test for treatment effect and compare it with that of a simpler analysis.

Adolescent

A comparative study of three methods for analysing longitudinal pulmonary function data.

We compare three methods of longitudinal analysis of pulmonary function data. Our data set is taken from a study of exposure to toluene diisocyanate (TDI) vapours in a new manufacturing plant. The first two methods are a two-stage weighted regression method and maximum likelihood estimation via the EM algorithm, and these give very similar results. The third method, regression with an autoregressive error structure, was not successfully implemented, and in our view needs better documentation.

Adult

The analysis of titration studies in phase III clinical trials.

Clinical trials commonly employ the titration design for certain drugs such as antihypertensives. In a Phase III trial the design has purposes distinct from those of a Phase I or II trial, as well as from those of a trial with a parallel design. In this paper we compare the titration design with the usual parallel design in their respective purposes for Phase III trials, explore the relevant questions addressed, and examine typical data from such trials. We also discuss work which focuses primarily on the Phase I or II titration trials. We formulate the problem in the framework of one-way contingency table augmented with incomplete data and obtain the maximum likelihood estimates of the parameters and their estimated variances/covariances via the EM algorithm. An example of a Phase III study of an antihypertensive agent illustrates the proposed procedure.

Analysis of Variance

Adjusting for age-related competing mortality in long-term cancer clinical trials.

Mortality related to causes other than the treated disease may have a significant impact on overall survival in long-term clinical trials. We present a model that adjusts for age-related competing mortality when cause of death is missing or only partially available. Through use of a piecewise exponential survival model, we extend relative survival methods to continuous follow-up data, allowing the competing mortality to differ from that of the general population by a scale parameter. An EM algorithm provides a simple way to compute the maximum likelihood estimators (MLEs) and to test hypotheses using widely available software. We compare the bias and relative efficiency of this model to a piecewise exponential Cox model for overall survival. Theoretical results are confirmed by simulations and illustrated with data from a clinical trial in colorectal cancer. This example also shows how age-related and disease-related mortality can be confounded in an analysis of overall survival. We conclude with a discussion of the advantages and disadvantages of the model.

Adolescent

A method of non-parametric back-projection and its application to AIDS data.

The method of back-projection has been used to estimate the unobserved past incidence of infection with the human immunodeficiency virus (HIV) and to obtain projections of future AIDS incidence. Here a new approach to back-projection, which avoids parametric assumptions about the form of the HIV infection intensity, is described. This approach gives the data greater opportunity to determine the shape of the estimated intensity function. The method is based on a modification of an EM algorithm for maximum likelihood estimation that incorporates smoothing of the estimated parameters. It is easy to implement on a computer because the computations are based on explicit formulae. The method is illustrated with applications to AIDS data from Australia, U.S.A. and Japanese haemophiliacs.

Acquired Immunodeficiency Syndrome

Computational aspects of analysing random effects/longitudinal models.

Random effects and longitudinal models are becoming increasingly popular in the analysis of many types of data, including medical and biopharmaceutical, because of their richness and flexibility. They can be, however, difficult to fit using traditional statistical tools. Fortunately, there now exists a burgeoning collection of newer computational methods that can be applied to draw inferences with such models. This review attempts to provide an introduction to some of these techniques by describing them as extensions of the EM algorithm, currently a standard tool for the analysis of longitudinal and random effects models. For clarity of exposition, the extensions are classified into three types: large-sample iterative; large-sample simulation, and small-sample simulation.

Longitudinal Studies

Methods for the analysis of informatively censored longitudinal data.

This paper describes the problem of informative censoring in longitudinal studies where the primary outcome is rate of change in a continuous variable. Standard approaches based on the linear random effects model are valid only when the data are missing in a non-ignorable fashion. Informative censoring, which is a special type of non-ignorably missing data, occurs when the probability of early termination is related to an individual subject's true rate of change. When present, informative censoring causes bias in standard likelihood-based analyses, as well as in weighted averages of individual least-squares slopes. This paper reviews several methods proposed by others for analysis of informatively censored longitudinal data, and outlines a new approach based on a log-normal survival model. Maximum likelihood estimates may be obtained via the EM algorithm. Advantages of this approach are that it allows general unbalanced data caused by staggered entry and unequally-timed visits, it utilizes all available data, including data from patients with only a single measurement, and it provides a unified method for estimating all model parameters. Issues related to study design when informative censoring may occur are also discussed.

Linear Models