PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 415 records · Page 23Linked to original sources

Are fMRI event-related response constant in time? A model selection answer.

An accurate estimation of the hemodynamic response function (HRF) in functional magnetic resonance imaging (fMRI) is crucial for a precise spatial and temporal estimate of the underlying neuronal processes. Recent works have proposed non-parametric estimation of the HRF under the hypotheses of linearity and stationarity in time. Biological literature suggests, however, that response magnitude may vary with attention or ongoing activity. We therefore test a more flexible model that allows for the variation of the magnitude of the HRF with time in a maximum likelihood framework. Under this model, the magnitude of the HRF evoked by a single event may vary across occurrences of the same type of event. This model is tested against a simpler model with a fixed magnitude using information theory. We develop a standard EM algorithm to identify the event magnitudes and the HRF. We test this hypothesis on a series of 32 regions (4 ROIS on eight subjects) of interest and find that the more flexible model is better than the usual model in most cases. The important implications for the analysis of fMRI time series for event-related neuroimaging experiments are discussed.

Algorithms↗

A graphical model for estimating stimulus-evoked brain responses from magnetoencephalography data with large background brain activity.

This paper formulates a novel probabilistic graphical model for noisy stimulus-evoked MEG and EEG sensor data obtained in the presence of large background brain activity. The model describes the observed data in terms of unobserved evoked and background factors with additive sensor noise. We present an expectation maximization (EM) algorithm that estimates the model parameters from data. Using the model, the algorithm cleans the stimulus-evoked data by removing interference from background factors and noise artifacts and separates those data into contributions from independent factors. We demonstrate on real and simulated data that the algorithm outperforms benchmark methods for denoising and separation. We also show that the algorithm improves the performance of localization with beamforming algorithms.

Algorithms↗

M385T polymorphism in the factor V gene, but not Leiden mutation, is associated with placental abruption in Finnish women.

This study determines whether genetic variability in the gene encoding factor V contributes to differences in susceptibility to placental abruption. Allele and genotype frequencies of three single nucleotide polymorphisms (SNPs) in the factor V gene leading to nonsynonymous changes (M385T in exon 8, and R485K and R506Q [Leiden mutation] in exon 10) were studied in 116 Caucasian women with placental abruption and 112 healthy controls. Single-point analysis was expanded to haplotype analysis and haplotype frequencies were estimated using an expectation-maximisation (EM) algorithm. Comparison of single-point allele and genotype distributions of SNPs in exon 8 and exon 10 of the factor V gene revealed statistically significant differences in M385T allele (P = 0.021) and genotype ( P = 0.013) frequencies between the patients and the control subjects. The C allele of SNP M385T was significantly less frequent among the patients (7%) vs. the control subjects (13%), at an odds ratio of 0.48 (95% CI 0.25-0.91). Allele and genotype differences between the patients and control subjects as regards R485K and Leiden mutation were not significant. In haplotype estimation analysis, there was a significantly lower frequency of haplotype T-R-R encoding the T385-R485-R506 variant in the group with placental abruption vs. the control group (P = 0.038) at an odds ratio of 0.519 (95% CI 0.272-0.987). We conclude that T385 is less frequent among the patient group than in the control group. The M385T variant in the factor V gene other than the Leiden mutation may play a role in disease susceptibility.

Abruptio Placentae↗

Mixture models of serum iron measures in population screening for hemochromatosis and iron overload.

Homozygosity for the C282Y mutation of the hemochromatosis gene on chromosome 6p (HFE) is a common genetic trait that increases susceptibility to iron overload. The authors describe and apply methodology developed for the analysis of phenotypic and genotypic data from 46,136 non-Hispanic Caucasians, a subset of the multi-ethnic cohort enrolled in the Hemochromatosis and Iron Overload Screening (HEIRS) Study. For analysis of the distribution of transferrin saturation (TS), mixtures of normal distributions were considered and the expectation-maximization (EM) algorithm was applied for parameter estimation. Maximized log-likelihoods were compared, and significance was assessed by resampling. Sensitivity, specificity, and predictive values from the modeled subpopulations were compared with the actual observed genotypes for C282Y and H63D mutations in the HFE gene. A strong association between HFE genotype and TS subpopulations was found in these data collected from different geographic regions, confirming the external validity of the statistical approach when applied to population-based data. It was concluded that mixture modeling of phenotypic data may provide a clinical guide for screening with gender-specific thresholds to identify potential samples for genetic testing.

Adult↗

Direct determination of MUC5B promoter haplotypes based on the method of single-strand conformation polymorphism and their statistical estimation.

Haplotype-based human genome research is important in identifying disease susceptibility genes efficiently. Although haplotype reconstruction by statistical methods is widely used, direct haplotype determination by molecular techniques has also been developed as a complementary method for statistical estimation. In this study, we demonstrate a molecular haplotyping method making use of single-strand conformation polymorphism (SSCP) gels. We identified 10 common SNPs and a dinucleotide insertion/deletion polymorphism within 2-kb region upstream of the transcription initiation site of MUC5B and determined haplotype structure, dividing the region into two DNA fragments. Real haplotypes were determined unambiguously by our SSCP-based analysis with fragments longer than 1 kb. Haplotypes reconstructed from diploid genotypes in the same region by the statistical methods including EM algorithm were also evaluated. Direct comparison between statistical estimation and direct determination of haplotypes revealed that major haplotypes containing multiple marker sites showing strong LD are estimated in great accuracy but that a variety of haplotypes reflecting weak LD are not reconstructed precisely enough. Our data can be helpful in implementing molecular haplotyping or statistical estimation, since usage of these methods may be determined depending on the haplotype structures.

Base Sequence↗

Model for mapping imprinted quantitative trait loci in an inbred F2 design.

The role of imprinting in shaping development has been ubiquitously observed in plants, animals, and humans. However, a statistical method that can detect and estimate the effects of imprinted quantitative trait loci (iQTL) over the genome has not been extensively developed. In this article, we propose a maximum likelihood approach for testing and estimating the imprinted effects of iQTL that contribute to variation in a quantitative trait. This approach, implemented with the EM algorithm, allows for a genome-wide scan for the existence of iQTL. This approach was used to reanalyze published data in an F(2) family derived from the LG/S and SM/S mouse strains. Several iQTL that regulate the growth of body weight by expressing paternally inherited alleles were identified. Our approach provides a standard procedure for testing the statistical significance of iQTL involved in the genetic control of complex traits.

Algorithms↗

Sharpening spots: correcting for bleedover in cDNA array images.

For cDNA array methods that depend on imaging of a radiolabel, we show that bleedover of one spot onto another, due to the gap between the array and the imaging media, can be a major problem. The images can be sharpened, however, using a blind convolution method based on the EM algorithm. The sharpened images look like a set of donuts, which concurs with our knowledge of the spotting process. Oversharpened images are actually useful as well, in locating the centers of each spot.

Algorithms↗

Maximum likelihood and Bayesian methods for estimating the distribution of selective effects among classes of mutations using DNA polymorphism data.

Maximum likelihood and Bayesian approaches are presented for analyzing hierarchical statistical models of natural selection operating on DNA polymorphism within a panmictic population. For analyzing Bayesian models, we present Markov chain Monte-Carlo (MCMC) methods for sampling from the joint posterior distribution of parameters. For frequentist analysis, an Expectation-Maximization (EM) algorithm is presented for finding the maximum likelihood estimate of the genome wide mean and variance in selection intensity among classes of mutations. The framework presented here provides an ideal setting for modeling mutations dispersed through the genome and, in particular, for the analysis of how natural selection operates on different classes of single nucleotide polymorphisms (SNPs).

Bayes Theorem↗

The loss of statistical power to distinguish populations when certain samples are ambiguous.

Case-control studies are used to map loci associated with a genetic disease. The usual case-control study tests for significant differences in frequencies of alleles at marker loci. In this paper, we consider the problem of comparing two or more marker loci simultaneously and testing for significant differences in haplotype rather than allele frequencies. We consider two situations. In the first, genotypes at marker loci are resolved into haplotypes by making use of biochemical methods or by genotyping family members. In the second, genotypes at marker loci are not resolved into haplotypes, but, by assuming random mating, haplotypes can be inferred using a likelihood method such as the expectation-maximization (EM) algorithm. We assume that a causative locus has two alleles with a multiplicative effect on the penetrance of a disease, with one allele increasing the penetrance by a factor pi. We find, for small values of pi-1 and large sample sizes, asymptotic results that predict the statistical power of a test for significant differences in haplotype frequencies between cases and a random sample of the population, both when haplotypes can be resolved and when haplotypes have to be inferred. The increase in power when haplotypes can be resolved can be expressed as a ratio R, which is the increase in sample size needed to achieve the same power when haplotypes are resolved over when they are not resolved. In general, R depends on the pattern of linkage disequilibrium between the causative allele and the marker haplotypes but is independent of the frequency of the causative allele and, to a first approximation, is independent of pi. For the special situation of two di-allelic marker loci, we obtain a simple expression for R and its upper bound.

Alleles↗

A SAS macro for sample size re-estimation.

The assessment of sample size in clinical trials comparing means requires a variance estimate of the main efficacy variable. If no reliable information about the variance of the key response is available at the beginning of a clinical trial, the use of data from the first 'few' patients entered in the trial ('internal pilot') may be appropriate to estimate the variance and thus to recalculate the required sample size. A SAS macro that implements the EM algorithm for carrying out and simulating such interim power evaluations without unblinding the treatment status is presented.

Algorithms↗

A zero-inflated Poisson mixed model to analyze diagnosis related groups with majority of same-day hospital stays.

With increasing trend of same-day procedures and operations performed for hospital admissions, it is important to analyze those Diagnosis Related Groups (DRGs) consisting of mainly same-day separations. A zero-inflated Poisson (ZIP) mixed model is presented to identify health- and patient-related characteristics associated with length of stay (LOS) and to model variations in LOS within such DRGs. Random effects are introduced to account for inter-hospital variations and the dependence of clustered LOS observations via the generalized linear mixed models (GLMM) approach. Parameter estimation is achieved by maximizing an appropriate log-likelihood function using the EM algorithm to obtain approximate residual maximum likelihood (REML) estimates. An S-Plus macro is developed to provide a unified ZIP modeling approach. The determination of pertinent factors would benefit hospital administrators and clinicians to manage LOS and expenditures efficiently.

Algorithms↗

[Imputation of the date of HIV seroconversion in cohorts of haemophiliacs].

OBJECTIVES: To describe the methods used to impute HIV seroconversion date in the haemophiliac cohorts from GEMES project and to validate its use. METHOD: 632 haemophiliacs coming from three hemophilia units identified as HIV+ and 1.092 individuals coming from 5 project GEMES cohorts with a seroconversion window (time among test HIV and HIV+) less than 3 years where mid point (PM) was assumed as seroconversion date. For both groups, seroconversion date was imputed after estimating the probability distribution of seroconversion by means of the EM algorithm. Two imputation methods are used: one obtained from the expected value and the other from the geometric mean of 5 random samples. from the estimated distribution. Imputations have been validated in the non haemophiliacs cohorts comparing with the PM seroconversion date. Also AIDS free time and survival from the different seroconversion imputed dates were compared. RESULTS: Median seroconversion date is located in May of 1993 for the non haemophiliacs and in 1982 for the haemophiliacs. Not big differences are observed among the imputed seroconversion dates and the mid-point seroconversion date in the non-haemophiliac cohorts. Similar results are found for the haemophiliac cohorts. Also no differences are observed in the estimated AIDS-free time for both groups of cohorts. CONCLUSIONS: Geometric mean imputation from several random samples provides a good estimate of the HIV seroconversion date that can be used to estimate AIDS-free time and survival in haemophiliac cohorts where seroconversion date is ignored.

Cohort Studies↗

The accuracy of DNA sequences: estimating sequence quality.

In this paper we describe a method for the statistical reconstruction of a large DNA sequence from a set of sequenced fragments. We assume that the fragments have been assembled and address the problem of determining the degree to which the reconstructed sequence is free from errors, i.e., its accuracy. A consensus distribution is derived from the assembled fragment configuration based upon the rates of sequencing errors in the individual fragments. The consensus distribution can be used to find a minimally redundant consensus sequence that meets a prespecified confidence level, either base by base or across any region of the sequence. A likelihood-based procedure for the estimation of the sequencing error rates, which utilizes an iterative EM algorithm, is described. Prior knowledge of the error rates is easily incorporated into the estimation procedure. The methods are applied to a set of assembled sequence fragments from the human G6PD locus. We close the paper with a brief discussion of the relevance and practical implications of this work.

Algorithms↗

Data smoothing regularization, multi-sets-learning, and problem solving strategies.

First, we briefly introduce the basic idea of data smoothing regularization, which was firstly proposed by Xu [Brain-like computing and intelligent information systems (1997) 241] for parameter learning in a way similar to Tikhonov regularization but with an easy solution to the difficulty of determining an appropriate hyper-parameter. Also, the roles of this regularization are demonstrated on Gaussian-mixture via smoothed versions of the EM algorithm, the BYY model selection criterion, adaptive harmony algorithm as well as its related Rival penalized competitive learning. Second, these studies are extended to a mixture of reconstruction errors of Gaussian types, which provides a new probabilistic formulation for the multi-sets learning approach [Proc. IEEE ICNN94 I (1994) 315] that learns multiple objects in typical geometrical structures such as points, lines, hyperplanes, circles, ellipses, and templates of given shapes. Finally, insights are provided on three problem solving strategies, namely the competition-penalty adaptation based learning, the global evidence accumulation based selection, and the guess-test based decision, with a general problem solving paradigm suggested.

Learning↗

Neural Networks for Predicting Conditional Probability Densities: Improved Training Scheme Combining EM and RVFL.

Predicting conditional probability densities with neural networks requires complex (at least two-hidden-layer) architectures, which normally leads to rather long training times. By adopting the RVFL concept and constraining a subset of the parameters to randomly chosen initial values (such that the EM-algorithm can be applied), the training process can be accelerated by about two orders of magnitude. This allows training of a whole ensemble of networks at the same computational costs as would be required otherwise for training a single model. The simulations performed suggest that in this way a significant improvement of the generalization performance can be achieved. Copyright 1997 Elsevier Science Ltd.

Journal Article↗

Plasma and tonsillar tissue pharmacokinetics of teicoplanin following intramuscular administration to children.

The population pharmacokinetics of teicoplanin in plasma and tonsillar tissue in children was determined following intramuscular administration. Thirty seven patients in all received either a single 5 mg/kg dose; 2 doses of 5 mg/kg, 12 h apart; 3 doses of 5 mg/kg, 12 h apart; or, a single 10 mg/kg dose. Limited data, comprising a maximum of 2 blood samples and 1 tonsillar sample were taken from each patient, with the maximum time being 48 h after the first dose of teicoplanin (in the 3 x 5 mg/kg dosing schedule). All plasma data were analyzed simultaneously by a maximum likelihood method employing a modified EM algorithm. A first-order absorption, one-compartment disposition model was fitted to the data. Mean parameter values (with lower and upper 95% confidence intervals) were: clearance/bioavailability, 0.024 L h(-1) kg(-1) (0.020-0.027); volume of distribution/bioavailability, 0.61 L kg(-1) (0.54-0.70); absorption rate constant, 0.43 h(-1) (0.31-0.61). A first-order transfer model for distribution of teicoplanin between plasma and tonsillar tissue was fitted to the tonsil data. The mean parameter values (95% confidence intervals) were: transfer rate constant between plasma and tonsils 0.49 h(-1) (0.35-0.67); transfer rate constant between tonsils and plasma 0.73 h(-1) (0.52-1.03). These rate constants correspond to a distribution half-life of 0.95 h and an equilibrium distribution concentration ratio between tonsillar tissue and plasma of 0.67. After normalising clearance and volume of distribution for body weight, there was no further influence of body weight on the pharmacokinetic parameters. Also, there was no effect of dose, and as two formulations were used, one for the 5 mg/kg dose and the other for the 10 mg/kg dose, no effect of formulation on the pharmacokinetics of teicoplanin after im (intramuscular) administration was found.

Algorithms↗

Fusing speed and phase information for vascular segmentation of phase contrast MR angiograms.

This paper presents a statistical approach to aggregating speed and phase (directional) information for vascular segmentation of phase contrast magnetic resonance angiograms (PC-MRA). Rather than relying on speed information alone, as done by others and in our own work, we demonstrate that including phase information as a priori knowledge in a Markov random field (MRF) model can improve the quality of segmentation. This is particularly true in the region within an aneurysm where there is a heterogeneous intensity pattern and significant vascular signal loss. We propose to use a Maxwell-Gaussian mixture density to model the background signal distribution and combine this with a uniform distribution for modelling vascular signal to give a Maxwell-Gaussian-uniform (MGU) mixture model of image intensity. The MGU model parameters are estimated by the modified expectation-maximisation (EM) algorithm. In addition, it is shown that the Maxwell-Gaussian mixture distribution (a) models the background signal more accurately than a Maxwell distribution, (b) exhibits a better fit to clinical data and (c) gives fewer false positive voxels (misclassified vessel voxels) in segmentation. The new segmentation algorithm is tested on an aneurysm phantom data set and two clinical data sets. The experimental results show that the proposed method can provide a better quality of segmentation when both speed and phase information are utilised.

Algorithms↗

A general mixture model approach for mapping quantitative trait loci from diverse cross designs involving multiple inbred lines.

Most current statistical methods developed for mapping quantitative trait loci (QTL) based on inbred line designs apply to crosses from two inbred lines. Analysis of QTL in these crosses is restricted by the parental genetic differences between lines. Crosses from multiple inbred lines or multiple families are common in plant and animal breeding programmes, and can be used to increase the efficiency of a QTL mapping study. A general statistical method using mixture model procedures and the EM algorithm is developed for mapping QTL from various cross designs of multiple inbred lines. The general procedure features three cross design matrices, W, that define the contribution of parental lines to a particular cross and a genetic design matrix, D, that specifies the genetic model used in multiple line crosses. By appropriately specifying W matrices, the statistical method can be applied to various cross designs, such as diallel, factorial, cyclic, parallel or arbitrary-pattern cross designs with two or multiple parental lines. Also, with appropriate specification for the D matrix, the method can be used to analyse different kinds of cross populations, such as F2 backcross, four-way cross and mixed crosses (e.g. combining backcross and F2). Simulation studies were conducted to explore the properties of the method, and confirmed its applicability to diverse experimental designs.

Animals↗