PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31Linked to original sources

Regression analysis of doubly censored failure time data with frailty.

In doubly censored failure time data, the survival time of interest is defined as the elapsed time between an initial event and a subsequent event, and the occurrences of both events cannot be observed exactly. Instead, only right- or interval-censored observations on the occurrence times are available. For the analysis of such data, a number of methods have been proposed under the assumption that the survival time of interest is independent of the occurrence time of the initial event. This article investigates a different situation where the independence may not be true with the focus on regression analysis of doubly censored data. Cox frailty models are applied to describe the effects of covariates and an EM algorithm is developed for estimation. Simulation studies are performed to investigate finite sample properties of the proposed method and an illustrative example from an acquired immune deficiency syndrome (AIDS) cohort study is provided.

Acquired Immunodeficiency Syndrome↗

Functional hierarchical models for identifying genes with different time-course expression profiles.

Time-course studies of gene expression are essential in biomedical research to understand biological phenomena that evolve in a temporal fashion. We introduce a functional hierarchical model for detecting temporally differentially expressed (TDE) genes between two experimental conditions for cross-sectional designs, where the gene expression profiles are treated as functional data and modeled by basis function expansions. A Monte Carlo EM algorithm was developed for estimating both the gene-specific parameters and the hyperparameters in the second level of modeling. We use a direct posterior probability approach to bound the rate of false discovery at a pre-specified level and evaluate the methods by simulations and application to microarray time-course gene expression data on Caenorhabditis elegans developmental processes. Simulation results suggested that the procedure performs better than the two-way ANOVA in identifying TDE genes, resulting in both higher sensitivity and specificity. Genes identified from the C. elegans developmental data set show clear patterns of changes between the two experimental conditions.

Algorithms↗

Likelihood methods for treatment noncompliance and subsequent nonresponse in randomized trials.

While several new methods that account for noncompliance or missing data in randomized trials have been proposed, the dual effects of noncompliance and nonresponse are rarely dealt with simultaneously. We construct a maximum likelihood estimator (MLE) of the causal effect of treatment assignment for a two-armed randomized trial assuming all-or-none treatment noncompliance and allowing for subsequent nonresponse. The EM algorithm is used for parameter estimation. Our likelihood procedure relies on a latent compliance state covariate that describes the behavior of a subject under all possible treatment assignments and characterizes the missing data mechanism as in Frangakis and Rubin (1999, Biometrika 86, 365-379). Using simulated data, we show that the MLE for normal outcomes compares favorably to the method-of-moments (MOM) and the standard intention-to-treat (ITT) estimators under (1) both normal and non-normal data, and (2) departures from the latent ignorability and compound exclusion restriction assumptions. We illustrate methods using data from a trial to compare the efficacy of two antipsychotics for adults with refractory schizophrenia.

Algorithms↗

Analysis of a partially observed binary covariate process and a censored failure time in the presence of truncation and competing risks.

We develop methods for assessing the association between a binary time-dependent covariate process and a failure time endpoint when the former is observed only at a single time point and the latter is right censored, and when the observations are subject to truncation and competing causes of failure. Using a proportional hazards model for the effect of the covariate process on the failure time of interest, we develop an approach utilizing EM algorithm and profile likelihood for estimating the relative risk parameter and cause-specific hazards for failure. The methods are extended to account for other covariates that can influence the time-dependent covariate process and cause-specific risks of failure. We illustrate the methods with data from a recent study on the association between loss of hepatitis B e antigen and the development of hepatocellular carcinoma in a population of chronic carriers of hepatitis B.

Algorithms↗

Joint modeling of survival and longitudinal data: likelihood approach revisited.

The maximum likelihood approach to jointly model the survival time and its longitudinal covariates has been successful to model both processes in longitudinal studies. Random effects in the longitudinal process are often used to model the survival times through a proportional hazards model, and this invokes an EM algorithm to search for the maximum likelihood estimates (MLEs). Several intriguing issues are examined here, including the robustness of the MLEs against departure from the normal random effects assumption, and difficulties with the profile likelihood approach to provide reliable estimates for the standard error of the MLEs. We provide insights into the robustness property and suggest to overcome the difficulty of reliable estimates for the standard errors by using bootstrap procedures. Numerical studies and data analysis illustrate our points.

Algorithms↗

Statistical cerebrovascular segmentation in three-dimensional rotational angiography based on maximum intensity projections.

Segmentation of three-dimensional rotational angiography (3D-RA) can provide quantitative 3D morphological information of vasculature. The expectation maximization-(EM-) based segmentation techniques have been widely used in the medical image processing community, because of the implementation simplicity, and computational efficiency of the approach. In a brain 3D-RA, vascular regions usually occupy a very small proportion (around 1%) inside an entire image volume. This severe imbalance between the intensity distributions of vessels and background can lead to inaccurate statistical modeling in the EM-based segmentation methods, and thus adversely affect the segmentation quality for 3D-RA. In this paper we present a new method for the extraction of vasculature in 3D-RA images. The new method is fully automatic and computationally efficient. As compared with the original 3D-RA volume, there is a larger proportion (around 20%) of vessels in its corresponding maximum intensity projection (MIP) image. The proposed method exploits this property to increase the accuracy of statistical modeling with the EM algorithm. The algorithm takes an iterative approach to compiling the 3D vascular segmentation progressively with the segmentation of MIP images along the three principal axes, and use a winner-takes-all strategy to combine the results obtained along individual axes. Experimental results on 12 3D-RA clinical datasets indicate that the segmentations obtained by the new method exhibit a high degree of agreement to the ground truth segmentations and are comparable to those produced by the manual optimal global thresholding method.

Algorithms↗

Approximate 3D iterative reconstruction for SPECT.

Compared with slice-by-slice approaches for SPECT reconstruction, three-dimensional iterative methods provide a more accurate physical model and an improved SPECT image. Clinical application of these methods, however, is limited primarily to their computational demands. This paper investigates the methods for approximate 3D iterative reconstruction that greatly reduce this demand by excluding from the reconstruction the smaller magnitude elements of the system matrix. A new method is described which is designed to control the resulting bias in the SPECT image for a given reduction in computation. The approximate methods were compared to fully 3D iterative reconstruction in terms of SPECT image bias and visual quality. All methods were incorporated into the ML-EM algorithm and applied to data from 3D mathematical and experimental brain phantoms. The SPECT images reconstructed by the approximate methods exhibited a positive bias throughout the image that was in general smaller with the new method (in the rage of 2%-6%). The bias was smallest in locally hot regions and largest in locally cold regions. The high quality brain phantom images demonstrated the capability of the new method in realistic imaging contexts. The time per iteration for an entire 3D brain phantom on a modern workstation using the approximate 3D method was 7.0 s.

Algorithms↗

Spontaneous speech recognition using a statistical coarticulatory model for the vocal-tract-resonance dynamics.

A statistical coarticulatory model is presented for spontaneous speech recognition, where knowledge of the dynamic, target-directed behavior in the vocal tract resonance is incorporated into the model design, training, and in likelihood computation. The principal advantage of the new model over the conventional HMM is the use of a compact, internal structure that parsimoniously represents long-span context dependence in the observable domain of speech acoustics without using additional, context-dependent model parameters. The new model is formulated mathematically as a constrained, nonstationary, and nonlinear dynamic system, for which a version of the generalized EM algorithm is developed and implemented for automatically learning the compact set of model parameters. A series of experiments for speech recognition and model synthesis using spontaneous speech data from the Switchboard corpus are reported. The promise of the new model is demonstrated by showing its consistently superior performance over a state-of-the-art benchmark HMM system under controlled experimental conditions. Experiments on model synthesis and analysis shed insight into the mechanism underlying such superiority in terms of the target-directed behavior and of the long-span context-dependence property, both inherent in the designed structure of the new dynamic model of speech.

Algorithms↗

A statistics-based pitch contour model for Mandarin speech.

A statistics-based syllable pitch contour model for Mandarin speech is proposed. This approach takes the mean and the shape of a syllable log-pitch contour as two basic modeling units and considers several affecting factors that contribute to their variations. The affecting factors include the speaker, prosodic state (which essentially represents the high-level linguistic components of F0 and will be explained more clearly in Sec. I), tone, and initial and final syllable classes. The parameters of the two modeling units were automatically estimated using the expectation-maximization (EM) algorithm. Experimental results showed that the root mean squared errors (RMSEs) obtained in the closed and open tests in the reconstructed pitch period were 0.362 and 0.373 ms, respectively. This model provides a way to separate the effects of several major factors. All of the inferred values of the affecting factors were in close agreement with our prior linguistic knowledge. It also gives a quantitative and more complete description of the coarticulation effect of neighboring tones rather than conventional qualitative descriptions of the tone sandhi rules. In addition, the model can provide useful cues to determine the prosodic phrase boundaries, including those occurring at intersyllable locations, with or without punctuation marks.

Adult↗

Speaker identification using time-delay HMEs.

In this paper, we extend the Hierarchical Mixture of Experts (HME) to temporal processing and explore it for a substantial problem, that of text-dependent speaker identification. For a specific multiway classification, we propose a generalized Bernoulli density instead of the multinomial logit density to avoid the instability during training. Time-delay technique is applied for spatio-temporal processing in the HME and a combining scheme is presented for combining multiple time-delay HMEs in order to complete a multi-scale analysis for the temporal data. Using the time-delay HME along with the EM algorithm as well as the combination of multiple time-delay HMEs, the speaker identification system has a good performance and yields significantly fast training. We have also addressed some issues about the time-delay techniques in the HME.

Algorithms↗

A modular neural network architecture for pattern classification based on different feature sets.

We propose a novel connectionist method for the use of different feature sets in pattern classification. Unlike traditional methods, e.g., combination of multiple classifiers and use of a composite feature set, our method copes with the problem based on an idea of soft competition on different feature sets developed in our earlier work. An alternative modular neural network architecture is proposed to provide a more effective implementation of soft competition on different feature sets. The proposed architecture is interpreted as a generalized finite mixture model and, therefore, parameter estimation is treated as a maximum likelihood problem. An EM algorithm is derived for parameter estimation and, moreover, a model selection method is proposed to fit the proposed architecture to a specific problem. Comparative results are presented for the real world problem of speaker identification.

Humans↗

Logos: a modular bayesian model for de novo motif detection.

The complexity of the global organization and internal structure of motifs in higher eukaryotic organisms raises significant challenges for motif detection techniques. To achieve successful de novo motif detection, it is necessary to model the complex dependencies within and among motifs and to incorporate biological prior knowledge. In this paper, we present LOGOS, an integrated LOcal and GlObal motif Sequence model for biopolymer sequences, which provides a principled framework for developing, modularizing, extending and computing expressive motif models for complex biopolymer sequence analysis. LOGOS consists of two interacting submodels: HMDM, a local alignment model capturing biological prior knowledge and positional dependency within the motif local structure; and HMM, a global motif distribution model modeling frequencies and dependencies of motif occurrences. Model parameters can be fit using training motifs within an empirical Bayesian framework. A variational EM algorithm is developed for de novo motif detection. LOGOS improves over existing models that ignore biological priors and dependencies in motif structures and motif occurrences, and demonstrates superior performance on both semi-realistic test data and cis-regulatory sequences from yeast and Drosophila genomes with regard to sensitivity, specificity, flexibility and extensibility.

Algorithms↗

Functional mapping for quantitative trait loci governing growth rates: a parametric model.

Are there-specific quantitative trait loci (QTL) governing growth rates in biology? This is emerging as an exciting but challenging question for contemporary developmental biology, evolutionary biology, and plant and animal breeding. In this article, we present a new statistical model for mapping QTL underlying age-specific growth rates. This model is based on the mechanistic relationship between growth rates and ages established by a variety of mathematical functions. A maximum likelihood approach, implemented with the EM algorithm, is developed to provide the estimates of QTL position, growth parameters characterized by QTL effects, and residual variances and covariances. Based on our model, a number of biologically important hypotheses can be formulated concerning the genetic basis of growth. We use forest trees as an example to demonstrate the power of our model, in which a QTL for stem growth diameter growth rates is successfully mapped to a linkage group constructed from polymorphic markers. The implications of the new model are discussed.

Age Factors↗

A model for estimating joint maternal-offspring effects on seed development in autogamous plants.

We present a statistical model for testing and estimating the effects of maternal-offspring genome interaction on the embryo and endosperm traits during seed development in autogamous plants. Our model is constructed within the context of maximum likelihood implemented with the EM algorithm. Extensive simulations were performed to investigate the statistical properties of our approach. We have successfully identified a quantitative trait locus that exerts a significant maternal-offspring interaction effect on amino acid contents of the endosperm in maize, demonstrating the power of our approach. This approach will be broadly useful in mapping endosperm traits for many agriculturally important crop plants and also make it possible to study the genetic significance of double fertilization in the evolution of higher plants.

Algorithms↗

A note on inference of trait associations with SNP haplotypes and other attributes in generalized linear models.

Recently, Lake et al. [Human Heredity 2003;55:56-65] have proposed an approach based on the EM algorithm for maximum-likelihood inference of trait associations with haplotypes and environmental cofactors in generalized linear models. In this short report, we describe an extension to accommodate missing SNP genotype information. We also discuss differences in the calculation of standard errors between their implementation and our own. Finally, we present results indicating that inference is robust to low levels of dependence between haplotypes and nongenetic factors, but that biased inference can result when there is moderate to strong dependence. Overall, the method is found to perform well in the models we considered.

Algorithms↗

Estimating haplotype effects on dichotomous outcome for unphased genotype data using a weighted penalized log-likelihood approach.

OBJECTIVE: To develop a method to estimate haplotype effects on dichotomous outcomes when phase is unknown, that can also estimate reliable effects of rare haplotypes. METHODS: In short, the method uses a logistic regression approach, with weights attached to all possible haplotype combinations of an individual. An EM-algorithm was used: in the E-step the weights are estimated, and the M-step consists of maximizing the joint log-likelihood. When rare haplotypes were present, a penalty function was introduced. We compared four different penalties. To investigate statistical properties of our method, we performed a simulation study for different scenarios. The evaluation criteria are the mean bias of the parameter estimates, the root of the mean squared error, the coverage probability, power, Type I error rate and the false discovery rate. RESULTS: For the unpenalized approach, mean bias was small, coverage probabilities were approximately 95%, power ranged from 15.2 to 44.7% depending on haplotype frequency, and Type I error rate was around 5%. All penalty functions reduced the standard errors of the rare haplotypes, but introduced bias. This trade-off decreased power. CONCLUSION: The unpenalized weighted log-likelihood approach performs well. A penalty function can help to estimate an effect for rare haplotypes.

Algorithms↗

The APL test: extension to general nuclear families and haplotypes and examination of its robustness.

OBJECTIVE: The Association in the Presence of Linkage test (APL) is a powerful statistical method that allows for missing parental genotypes in nuclear families. However, in its original form, the statistic does not easily extend to mixed nuclear family structures nor to multiple-marker haplotypes. Furthermore, the robustness of APL in practice has not been examined. Here we present a generalization of the APL model and examination of its robustness under a variety of non-standard scenarios. METHODS: The generalization is made possible by incorporating a bootstrap variance estimator instead of the original robust variance estimator. This allows for use of more than two affected siblings. Haplotype analysis was accomplished by combining estimation of haplotype phase into the EM algorithm. Computer simulation was used to examine robustness of the APL to departures from test assumptions. RESULTS: The extended APL tests both single-marker and multiple-marker haplotypes and shows more power than other association methods. Simulation results showed that the single-marker APL test is robust to the departure from HWE. For the haplotype test, violation of the HWE assumption can inflate type I error. We also evaluated general guidelines for the validity of APL with rare alleles and rare haplotypes. Software for the APL test is available from http://www.chg.duke.edu/research/apl.html.

Computer Simulation↗

Candidate-gene association study of mothers with pre-eclampsia, and their infants, analyzing 775 SNPs in 190 genes.

Pre-eclampsia (PE) affects 5-7% of pregnancies in the US, and is a leading cause of maternal death and perinatal morbidity and mortality worldwide. To identify genes with a role in PE, we conducted a large-scale association study evaluating 775 SNPs in 190 candidate genes selected for a potential role in obstetrical complications. SNP discovery was performed by DNA sequencing, and genotyping was carried out in a high-throughput facility using the MassARRAY(TM) System. Women with PE (n = 394) and their offspring (n = 324) were compared with control women (n = 602) and their offspring (n = 631) from the same hospital-based population. Haplotypes were estimated for each gene using the EM algorithm, and empirical p values were obtained for a logistic regression-based score test, adjusted for significant covariates. An interaction model between maternal and offspring genotypes was also evaluated. The most significant findings for association with PE were COL1A1 (p = 0.0011) and IL1A (p = 0.0014) for the maternal genotype, and PLAUR (p = 0.0008) for the offspring genotype. Common candidate genes for PE, including MTHFR and NOS3, were not significantly associated with PE. For the interaction model, SNPs within IGF1 (p = 0.0035) and IL4R (p = 0.0036) gave the most significant results. This study is one of the most comprehensive genetic association studies of PE to date, including an evaluation of offspring genotypes that have rarely been considered in previous studies. Although we did not identify statistically significant evidence of association for any of the candidate loci evaluated here after adjusting for multiple testing using the false discovery rate, additional compelling evidence exists, including multiple SNPs with nominally significant p values in COL1A1 and the IL1A region, and previous reports of association for IL1A, to support continued interest in these genes as candidates for PE. Identification of the genetic regulators of PE may have broader implications, since women with PE are at increased risk of death from cardiovascular diseases later in life.

Adult↗