PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32Linked to original sources

Variational learning for switching state-space models.

We introduce a new statistical model for time series that iteratively segments data into regimes with approximately linear dynamics and learnsthe parameters of each of these linear regimes. This model combines and generalizes two of the most widely used stochastic time-series models -- hidden Markov models and linear dynamical systems -- and is closely related to models that are widely used in the control and econometrics literatures. It can also be derived by extending the mixture of experts neural network (Jacobs, Jordan, Nowlan, & Hinton, 1991) to its fully dynamical version, in which both expert and gating networks are recurrent. Inferring the posterior probabilities of the hidden states of this model is computationally intractable, and therefore the exact expectation maximization (EM) algorithm cannot be applied. However, we present a variational approximation that maximizes a lower bound on the log-likelihood and makes use of both the forward and backward recursions for hidden Markov models and the Kalman filter recursions for linear dynamical systems. We tested the algorithm on artificial data sets and a natural data set of respiration force from a patient with sleep apnea. The results suggest that variational approximations are a viable method for inference and learning in switching state-space models.

Algorithms↗

LeRec: a NN/HMM hybrid for on-line handwriting recognition.

We introduce a new approach for on-line recognition of handwritten words written in unconstrained mixed style. The preprocessor performs a word-level normalization by fitting a model of the word structure using the EM algorithm. Words are then coded into low resolution "annotated images" where each pixel contains information about trajectory direction and curvature. The recognizer is a convolution network that can be spatially replicated. From the network output, a hidden Markov model produces word scores. The entire system is globally trained to minimize word-level errors.

Algorithms↗

Dynamic model of visual recognition predicts neural response properties in the visual cortex.

The responses of visual cortical neurons during fixation tasks can be significantly modulated by stimuli from beyond the classical receptive field. Modulatory effects in neural responses have also been recently reported in a task where a monkey freely views a natural scene. In this article, we describe a hierarchical network model of visual recognition that explains these experimental observations by using a form of the extended Kalman filter as given by the minimum description length (MDL) principle. The model dynamically combines input-driven bottom-up signals with expectation-driven top-down signals to predict current recognition state. Synaptic weights in the model are adapted in a Hebbian manner according to a learning rule also derived from the MDL principle. The resulting prediction-learning scheme can be viewed as implementing a form of expectation-maximization (EM) algorithm. The architecture of the model posits an active computational role of the reciprocal connections between adjoining visual cortical areas in determining neural response properties. In particular, the model demonstrates the possible role of feedback from higher cortical areas in mediating neurophysiological effects due to stimuli from beyond the classical receptive field. Simulations of the model are provided that help explain the experimental observations regarding neural responses in both free viewing and fixation conditions.

Animals↗

Statistical models for PET and SPECT data.

This article outlines the statistical developments that have taken place in emission tomography during the past decade or so. We discuss the statistical aspects of the modelling of the projection data and define the additive Poisson regression model. This leads to the use of the method of maximum likelihood as a means of estimating the underlying isotope concentration within a given region of a patient's body, and to the use of the EM algorithm to compute the reconstruction. The need for the regulation of the maximum likelihood solution is tackled using Bayesian techniques. A number of algorithms for the computation of regularized solutions are outlined. The issue of parameter estimation is discussed and some open issues are mentioned.

Algorithms↗

Maximum likelihood analysis of generalized linear models with missing covariates.

Missing data is a common occurrence in most medical research data collection enterprises. There is an extensive literature concerning missing data, much of which has focused on missing outcomes. Covariates in regression models are often missing, particularly if information is being collected from multiple sources. The method of weights is an implementation of the EM algorithm for general maximum-likelihood analysis of regression models, including generalized linear models (GLMs) with incomplete covariates. In this paper, we will describe the method of weights in detail, illustrate its application with several examples, discuss its advantages and limitations, and review extensions and applications of the method.

Algorithms↗

Imputation strategies for blood pressure data nonignorably missing due to medication use.

BACKGROUND: Underlying or untreated blood pressure (BP) is often an outcome of interest, but is unobservable when study participants are on anti-hypertensive medications. Untreated levels are not missing at random but would be higher among those on such medication. In such cases, standard methods of analysis may lead to bias. PURPOSE: BPs obtained at the private physician's office (out-of-study BPs) at the time of prescription of anti-hypertensive medications were available from Phase II of the Trials of Hypertension Prevention (TOHP) and were used to adjust for the potential bias. METHODS: Observed out-of-study BPs were used to estimate the conditional expectation and variance of the unobserved unmedicated study BPs. For those with no physician data, imputation from bootstrap samples of out-of-study BPs was used. An iterative method based on the EM algorithm was used to estimate the unknown study parameters in a random-effects model using multiple imputations. This was compared to an alternative model for the out-of-study BPs based on a theoretical truncated normal distribution, and to standard analyses, including both multivariate repeated measures and last-observation-carried-forward (LOCF) analyses, using data from Phase II of TOHP. RESULTS: Differences between methods were seen in the decline in BP over time in the reference group, where the changes from baseline to 36 months were 3.0 in univariate analyses, 2.4 using LOCF, and 2.6 in the multivariate analysis, compared to 2.0 or 1.7 in the imputation analyses, depending on the number of physician visits. Estimated intervention effects tended to be slightly larger using the imputation methods. LIMITATIONS: out-of-study measures may not be available for other studies. CONCLUSIONS: Because the proposed strategy was based on an empirically observed distribution for out-of-study BP, fewer assumptions about the missing data were made. These data may be useful in suggesting imputation strategies for other studies.

Blood Pressure↗

Modelling the growth curve of Maine-Anjou beef cattle using heteroskedastic random coefficients models.

A heteroskedastic random coefficients model was described for analyzing weight performances between the 100th and the 650th days of age of Maine-Anjou beef cattle. This model contained both fixed effects, random linear regression and heterogeneous variance components. The objective of this study was to analyze the difference of growth curves between animals born as twin and single bull calves. The method was based on log-linear models for residual and individual variances expressed as functions of explanatory variables. An expectation-maximization (EM) algorithm was proposed for calculating restricted maximum likelihood (REML) estimates of the residual and individual components of variances and covariances. Likelihood ratio tests were used to assess hypotheses about parameters of this model. Growth of Maine-Anjou cattle was described by a third order regression on age for a mean growth curve, two correlated random effects for the individual variability and independent errors. Three sources of heterogeneity of residual variances were detected. The difference of weight performance between bulls born as single and twin bull calves was estimated to be equal to about 15 kg for the growth period considered.

Algorithms↗

XRate: a fast prototyping, training and annotation tool for phylo-grammars.

BACKGROUND: Recent years have seen the emergence of genome annotation methods based on the phylo-grammar, a probabilistic model combining continuous-time Markov chains and stochastic grammars. Previously, phylo-grammars have required considerable effort to implement, limiting their adoption by computational biologists. RESULTS: We have developed an open source software tool, xrate, for working with reversible, irreversible or parametric substitution models combined with stochastic context-free grammars. xrate efficiently estimates maximum-likelihood parameters and phylogenetic trees using a novel "phylo-EM" algorithm that we describe. The grammar is specified in an external configuration file, allowing users to design new grammars, estimate rate parameters from training data and annotate multiple sequence alignments without the need to recompile code from source. We have used xrate to measure codon substitution rates and predict protein and RNA secondary structures. CONCLUSION: Our results demonstrate that xrate estimates biologically meaningful rates and makes predictions whose accuracy is comparable to that of more specialized tools.

Algorithms↗

A hierarchical statistical model for estimating population properties of quantitative genes.

BACKGROUND: Earlier methods for detecting major genes responsible for a quantitative trait rely critically upon a well-structured pedigree in which the segregation pattern of genes exactly follow Mendelian inheritance laws. However, for many outcrossing species, such pedigrees are not available and genes also display population properties. RESULTS: In this paper, a hierarchical statistical model is proposed to monitor the existence of a major gene based on its segregation and transmission across two successive generations. The model is implemented with an EM algorithm to provide maximum likelihood estimates for genetic parameters of the major locus. This new method is successfully applied to identify an additive gene having a large effect on stem height growth of aspen trees. The estimates of population genetic parameters for this major gene can be generalized to the original breeding population from which the parents were sampled. A simulation study is presented to evaluate finite sample properties of the model. CONCLUSIONS: A hierarchical model was derived for detecting major genes affecting a quantitative trait based on progeny tests of outcrossing species. The new model takes into account the population genetic properties of genes and is expected to enhance the accuracy, precision and power of gene detection.

Algorithms↗

A multilocus likelihood approach to joint modeling of linkage, parental diplotype and gene order in a full-sib family.

BACKGROUND: Unlike a pedigree initiated with two inbred lines, a full-sib family derived from two outbred parents frequently has many different segregation types of markers whose linkage phases are not known prior to linkage analysis. RESULTS: We formulate a general model of simultaneously estimating linkage, parental diplotype and gene order through multi-point analysis in a full-sib family. Our model is based on a multinomial mixture model taking into account different diplotypes and gene orders, weighted by their corresponding occurring probabilities. The EM algorithm is implemented to provide the maximum likelihood estimates of the linkage, parental diplotype and gene order over any type of markers. CONCLUSIONS: Through simulation studies, this model is found to be more computationally efficient compared with existing models for linkage mapping. We discuss the extension of the model and its implications for genome mapping in outcrossing species.

Alleles↗

Linkage disequilibrium mapping via cladistic analysis of phase-unknown genotypes and inferred haplotypes in the Genetic Analysis Workshop 14 simulated data.

We recently described a method for linkage disequilibrium (LD) mapping, using cladistic analysis of phased single-nucleotide polymorphism (SNP) haplotypes in a logistic regression framework. However, haplotypes are often not available and cannot be deduced with certainty from the unphased genotypes. One possible two-stage approach is to infer the phase of multilocus genotype data and analyze the resulting haplotypes as if known. Here, haplotypes are inferred using the expectation-maximization (EM) algorithm and the best-guess phase assignment for each individual analyzed. However, inferring haplotypes from phase-unknown data is prone to error and this should be taken into account in the subsequent analysis. An alternative approach is to analyze the phase-unknown multilocus genotypes themselves. Here we present a generalization of the method for phase-known haplotype data to the case of unphased SNP genotypes. Our approach is designed for high-density SNP data, so we opted to analyze the simulated dataset. The marker spacing in the initial screen was too large for our method to be effective, so we used the answers provided to request further data in regions around the disease loci and in null regions. Power to detect the disease loci, accuracy in localizing the true site of the locus, and false-positive error rates are reported for the inferred-haplotype and unphased genotype methods. For this data, analyzing inferred haplotypes outperforms analysis of genotypes. As expected, our results suggest that when there is little or no LD between a disease locus and the flanking region, there will be no chance of detecting it unless the disease variant itself is genotyped.

Chromosome Mapping↗

Multinomial logistic regression approach to haplotype association analysis in population-based case-control studies.

BACKGROUND: The genetic association analysis using haplotypes as basic genetic units is anticipated to be a powerful strategy towards the discovery of genes predisposing human complex diseases. In particular, the increasing availability of high-resolution genetic markers such as the single-nucleotide polymorphisms (SNPs) has made haplotype-based association analysis an attractive alternative to single marker analysis. RESULTS: We consider haplotype association analysis under the population-based case-control study design. A multinomial logistic model is proposed for haplotype analysis with unphased genotype data, which can be decomposed into a prospective logistic model for disease risk as well as a model for the haplotype-pair distribution in the control population. Environmental factors can be readily incorporated and hence the haplotype-environment interaction can be assessed in the proposed model. The maximum likelihood estimation with unphased genotype data can be conveniently implemented in the proposed model by applying the EM algorithm to a prospective multinomial logistic regression model and ignoring the case-control design. We apply the proposed method to the hypertriglyceridemia study and identifies 3 haplotypes in the apolipoprotein A5 gene that are associated with increased risk for hypertriglyceridemia. A haplotype-age interaction effect is also identified. Simulation studies show that the proposed estimator has satisfactory finite-sample performances. CONCLUSION: Our results suggest that the proposed method can serve as a useful alternative to existing methods and a reliable tool for the case-control haplotype-based association analysis.

Algorithms↗

Identification of novel functional sequence variants in the gene for peptidase inhibitor 3.

BACKGROUND: Peptidase inhibitor 3 (PI3) inhibits neutrophil elastase and proteinase-3, and has a potential role in skin and lung diseases as well as in cancer. Genome-wide expression profiling of chorioamniotic membranes revealed decreased expression of PI3 in women with preterm premature rupture of membranes. To elucidate the molecular mechanisms contributing to the decreased expression in amniotic membranes, the PI3 gene was searched for sequence variations and the functional significance of the identified promoter variants was studied. METHODS: Single nucleotide polymorphisms (SNPs) were identified by direct sequencing of PCR products spanning a region from 1,173 bp upstream to 1,266 bp downstream of the translation start site. Fourteen SNPs were genotyped from 112 and nine SNPs from 24 unrelated individuals. Putative transcription factor binding sites as detected by in silico search were verified by electrophoretic mobility shift assay (EMSA) using nuclear extract from Hela and amnion cell nuclear extract. Deviation from Hardy-Weinberg equilibrium (HWE) was tested by chi2 goodness-of-fit test. Haplotypes were estimated using expectation maximization (EM) algorithm. RESULTS: Twenty-three sequence variations were identified by direct sequencing of polymerase chain reaction (PCR) products covering 2,439 nt of the PI3 gene (-1,173 nt of promoter sequences and all three exons). Analysis of 112 unrelated individuals showed that 20 variants had minor allele frequencies (MAF) ranging from 0.02 to 0.46 representing "true polymorphisms", while three had MAF < or = 0.01. Eleven variants were in the promoter region; several putative transcription factor binding sites were found at these sites by database searches. Differential binding of transcription factors was demonstrated at two polymorphic sites by electrophoretic mobility shift assays, both in amniotic and HeLa cell nuclear extracts. Differential binding of the transcription factor GATA1 at -689C>G site was confirmed by a supershift. CONCLUSION: The promoter sequences of PI3 have a high degree of variability. Functional promoter variants provide a possible mechanism for explaining the differences in PI3 mRNA expression levels in the chorioamniotic membranes, and are also likely to be useful in elucidating the role of PI3 in other diseases.

Binding Sites↗

Evidence for the association of the SLC22A4 and SLC22A5 genes with type 1 diabetes: a case control study.

BACKGROUND: Type 1 diabetes (T1D) is a chronic, autoimmune and multifactorial disease characterized by abnormal metabolism of carbohydrate and fat. Diminished carnitine plasma levels have been previously reported in T1D patients and carnitine increases the sensitivity of the cells to insulin. Polymorphisms in the carnitine transporters, encoded by the SLC22A4 and SLC22A5 genes, have been involved in susceptibility to two other autoimmune diseases, rheumatoid arthritis and Crohn's disease. For these reasons, we investigated for the first time the association with T1D of six single nucleotide polymorphisms (SNPs) mapping to these candidate genes: slc2F2, slc2F11, T306I, L503F, OCTN2-promoter and OCTN2-intron. METHODS: A case-control study was performed in the Spanish population with 295 T1D patients and 508 healthy control subjects. Maximum-likelihood haplotype frequencies were estimated by applying the Expectation-Maximization (EM) algorithm implemented by the Arlequin software. RESULTS: When independently analyzed, one of the tested polymorphisms in the SLC22A4 gene at 1672 showed significant association with T1D in our Spanish cohort. The overall comparison of the inferred haplotypes was significantly different between patients and controls (chi2 = 10.43; p = 0.034) with one of the haplotypes showing a protective effect for T1D (rs3792876/rs1050152/rs2631367/rs274559, CCGA: OR = 0.62 (0.41-0.93); p = 0.02). CONCLUSION: The haplotype distribution in the carnitine transporter locus seems to be significantly different between T1D patients and controls; however, additional studies in independent populations would allow to confirm the role of these genes in T1D risk.

Adolescent↗

Techniques for incorporating longitudinal measurements into analyses of survival data from clinical trials.

This article reviews existing approaches for joint analysis of longitudinal measurements, possibly measured with error or incompletely observed, and event-time data, possibly censored. The models take the form of selection or pattern-mixture models; estimation proceeds via the EM algorithm or Bayesian sampling techniques. The models are compared, their estimation and inferential procedures described, and advantages and disadvantages noted. Examples are discussed from several disease areas, including cancer and AIDS.

Algorithms↗

Interpreting the 13C-urea breath test among a large population of young children from a developing country.

The 13C-urea breath test is a noninvasive tool for the diagnosis of gastric Helicobacter pylori infection. However, it has not been validated in young children from the developing world, where infection is very common. 13C urea breath tests were performed on 1532 occasions on 247 Gambian infants and children aged from 3 to 48 mo. The means and variances of the separate sub-populations of 13C enrichment results contained within the overall dataset were estimated by a Genstat procedure using the EM algorithm, thereby identifying a cut-off value to discriminate positive from negative results. To illustrate the appropriateness of this calculated cut-off value, 13C urea breath tests were performed upon a small group of 14 patients aged 6 to 28 mo undergoing diagnostic upper endoscopy. Fixed gastric antral biopsies were examined to identify H. pylori. Two subpopulations were identified within the large dataset. A cut-off value of 5.47 delta per thousand relative to Pee Dee Belemnite limestone above baseline at 30 min identified 95% of the normally distributed negative sub-population and 99.4% of the log normal distributed positive sub-population. Comparison with endoscopic data confirmed that this cut-off value was appropriate for this population, as 7/7 children without H. pylori on their gastric biopsies had negative urea breath tests, and 6/7 children with gastric H. pylori colonization had positive urea breath tests. These findings confirm the value of the urea breath test as a diagnostic tool in young children from developing countries. They also offer a way to calculate the most appropriate cut-off value for use in different populations and the likelihood that it will correctly assign any value into the appropriate sub-population, without the need for endoscopy.

Breath Tests↗

A pharmacokinetic/pharmacodynamic model for a platelet activating factor antagonist based on data arising from Phase I studies.

A nonlinear mixed-effects modelling approach was used to analyse pharmacokinetic and pharmacodynamic data from two Phase I studies of a platelet activating factor (PAF) antagonist under development for the treatment of seasonal allergic rhinitis. Data for single-dose (8 subjects) and multiple-dose (9 subjects) administration were available for analysis with a program based on an EM algorithm. Pharmacokinetic analyses of plasma drug concentrations were performed using a biexponential model with first-order absorption. PAF response data were modelled with a hyperbolic Emax model. The drug showed nonlinear pharmacokinetics, with the clearance decreasing from 46.0 to 27.1 L h(-1) over a dose range of 160-480 mg. There was an apparent dose dependency within the C50 (concentration producing 50% of the maximum effect) but at higher doses most of the data was above the estimated C50 and when the data was analysed simultaneously a value of 17.57 ng mL(-1) was obtained for C50, with considerable intersubject variability (103%). Consistent results were obtained from the two studies and the population and individual pharmacodynamic parameter estimates from the analyses provided predicted responses that were in good agreement with the observed data. The results were used to simulate a 320-mg twice-daily dosing regimen.

Administration, Oral↗

Imputation of missing data when measuring physical activity by accelerometry.

PURPOSE: We consider the issue of summarizing accelerometer activity count data accumulated over multiple days when the time interval in which the monitor is worn is not uniform for every subject on every day. The fact that counts are not being recorded during periods in which the monitor is not worn means that many common estimators of daily physical activity are biased downward. METHODS: Data from the Trial for Activity in Adolescent Girls (TAAG), a multicenter group-randomized trial to reduce the decline in physical activity among middle-school girls, were used to illustrate the problem of bias in estimation of physical activity due to missing accelerometer data. The effectiveness of two imputation procedures to reduce bias was investigated in a simulation experiment. Count data for an entire day, or a segment of the day were deleted at random or in an informative way with higher probability of missingness at upper levels of body mass index (BMI) and lower levels of physical activity. RESULTS: When data were deleted at random, estimates of activity computed from the observed data and those based on a data set in which the missing data have been imputed were equally unbiased; however, imputation estimates were more precise. When the data were deleted in a systematic fashion, the bias in estimated activity was lower using imputation procedures. Both imputation techniques, single imputation using the EM algorithm and multiple imputation (MI), performed similarly, with no significant differences in bias or precision. CONCLUSIONS: Researchers are encouraged to take advantage of software to implement missing value imputation, as estimates of activity are more precise and less biased in the presence of intermittent missing accelerometer data than those derived from an observed data analysis approach.

Acceleration↗