PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 361 records · Page 20Linked to original sources

A Monte Carlo investigation of dual-energy-window scatter correction for volume-of-interest quantification in 99Tcm SPECT.

Using Monte Carlo simulation of 99Tcm single-photon-emission computed tomography (SPECT), we investigate the effects of tissue-background activity, tumour location, patient size, uncertainty of energy windows, and definition of tumour region on the accuracy of quantification. The dual-energy-window method of correction for Compton scattering is employed and the multiplier which yields correct activity for the VI as a whole calculated. The model is usually a sphere containing radioactive water located within a cylinder filled with a more dilute solution of radioactivity. Two simulation codes are employed. Reconstruction is by ML-EM algorithm with attenuation compensation. The scatter multiplier depends only slightly on the sphere location or the cylinder diameter. It also depends little on whether correction is before or after reconstruction. At low background level, it changes with VOI size, but not at higher background. For a geometrical VOI, it is 1.25 at zero background, decreases sharply to 0.56 for equal concentrations, and is 0.44 when the background concentration is very large. Quantification is accurate (less than 9% error) if the test background is reasonably close to that used in setting the universal scatter-multiplier value, or if the test backgrounds are always large and so is the universal-value background, but not if the test backgrounds cover a large range of values including zero. Results largely agree with those from experiment after the experimental data with background is re-evaluated with prejudice.

Humans↗

Validation of the central-ray approximation for attenuated depth-dependent convolution in quantitative SPECT reconstruction.

In order to model photon attenuation and detector resolution variation as a depth-dependent convolution for efficient reconstruction of quantitative SPECT, a central-ray approximation is necessary. This work investigates the impact of the approximation upon reconstruction accuracy and computational efficiency. A patient chest CT image was acquired and converted into an object-specific attenuation map. From a segmentation of the map, an emission thorax phantom was constructed with a cardiac insert. To generate a system-specific resolution-variant kernal, a point source was measured at several depths from the surface of a low-energy, high-resolution, parallel-hole collimator of a SPECT system. Projections of parallel-beam geometry were simulated from the phantom, the map, and the kernel on an elliptical orbit. Reconstruction was performed by the ML-EM algorithm with and without the central-ray approximation. The approximation cuts down dramatically (more than 100 fold) the computing time with a negligible loss (less than 1%) of reconstruction accuracy.

Algorithms↗

Scatter compensation methods in 3D iterative SPECT reconstruction: a simulation study.

Effects of different scatter compensation methods incorporated in fully 3D iterative reconstruction are investigated. The methods are: (i) the inclusion of an 'ideal scatter estimate' (ISE); (ii) like (i) but with a noiseless scatter estimate (ISE-NF); (iii) incorporation of scatter in the point spread function during iterative reconstruction ('ideal scatter model', ISM); (iv) no scatter compensation (NSC); (v) ideal scatter rejection (ISR), as can be approximated by using a camera with a perfect energy resolution. The iterative method used was an ordered subset expectation maximization (OS-EM) algorithm. A cylinder containing small cold spheres was used to calculate contrast-to-noise curves. For a brain study, global errors between reconstruction and 'true' distributions were calculated. Results show that ISR is superior to all other methods. In all cases considered, ISM is superior to ISE and performs approximately as well as (brain study) or better than (cylinder data) ISE-NF. Both ISM and ISE improve contrast-to-noise curves and reduce global errors, compared with NSC. In the case of ISE, blurring of the scatter estimate with a Gaussian kernel results in slightly reduced errors in brain studies, especially at low count levels. The optimal Gaussian kernel size is strongly dependent on the noise level.

Algorithms↗

Megavoltage CT on a tomotherapy system.

A megavoltage computed tomography (MVCT) system was developed on the University of Wisconsin tomotherapy benchtop. This system can operate either axially or helically, and collect transmission data without any bounds on delivered dose. Scan times as low as 12 s per slice are possible, and scans were run with linac output rates of 100 MU min(-1), although the system can be tuned to deliver arbitrarily low dose rates. Images were reconstructed with clinically reasonable doses ranging from 8 to 12 cGy. These images delineate contrasts below 2% and resolutions of 3.0 mm. Thus, the MVCT image quality of this system should be sufficient for verifying the patient's position and anatomy prior to radiotherapy. Additionally, synthetic data were used to test the potential for improved MVCT contrast using maximum-likelihood (ML) reconstruction. Specifically, the maximum-likelihood expectation-maximization (ML-EM) algorithm and a transmission ML algorithm were compared with filtered backprojection (FBP). It was found that for expected clinical MVCT doses enough imaging photons are used such that little benefit is conferred by the improved noise model of ML algorithms. For significantly lower doses, some quantitative improvement is achieved through ML reconstruction. Nonetheless, the image quality at those lower doses is not satisfactory for radiotherapy verification.

Algorithms↗

Detection of recombination in DNA multiple alignments with hidden Markov models.

Conventional phylogenetic tree estimation methods assume that all sites in a DNA multiple alignment have the same evolutionary history. This assumption is violated in data sets from certain bacteria and viruses due to recombination, a process that leads to the creation of mosaic sequences from different strains and, if undetected, causes systematic errors in phylogenetic tree estimation. In the current work, a hidden Markov model (HMM) is employed to detect recombination events in multiple alignments of DNA sequences. The emission probabilities in a given state are determined by the branching order (topology) and the branch lengths of the respective phylogenetic tree, while the transition probabilities depend on the global recombination probability. The present study improves on an earlier heuristic parameter optimization scheme and shows how the branch lengths and the recombination probability can be optimized in a maximum likelihood sense by applying the expectation maximization (EM) algorithm. The novel algorithm is tested on a synthetic benchmark problem and is found to clearly outperform the earlier heuristic approach. The paper concludes with an application of this scheme to a DNA sequence alignment of the argF gene from four Neisseria strains, where a likely recombination event is clearly detected.

Algorithms↗

RMDP: a dedicated package for 131I SPECT quantification, registration and patient-specific dosimetry.

The limitations of traditional targeted radionuclide therapy (TRT) dosimetry can be overcome by using voxel-based techniques. All dosimetry techniques are reliant on a sequence of quantitative emission and transmission data. The use of (131)I, for example, with NaI or mIBG, presents additional quantification challenges beyond those encountered in low-energy NM diagnostic imaging, including dead-time correction and additional photon scatter and penetration in the camera head. The Royal Marsden Dosimetry Package (RMDP) offers a complete package for the accurate processing and analysis of raw emission and transmission patient data. Quantitative SPECT reconstruction is possible using either FBP or OS-EM algorithms. Manual, marker- or voxel-based registration can be used to register images from different modalities and the sequence of SPECT studies required for 3-D dosimetry calculations. The 3-D patient-specific dosimetry routines, using either a beta-kernel or voxel S-factor, are included. Phase-fitting each voxel's activity series enables more robust maps to be generated in the presence of imaging noise, such as is encountered during late, low-count scans or when there is significant redistribution within the VOI between scans. Error analysis can be applied to each generated dose-map. Patients receiving (131)I-mIBG, (131)I-NaI, and (186)Re-HEDP therapies have been analyzed using RMDP. A Monte-Carlo package, developed specifically to address the problems of (131)I quantification by including full photon interactions in a hexagonal-hole collimator and the gamma camera crystal, has been included in the dosimetry package. It is hoped that the addition of this code will lead to improved (131)I image quantification and will contribute towards more accurate 3-D dosimetry.

Humans↗

Cancer mortality in Ireland, 1926-1995.

BACKGROUND: Investigation of long time series of cancer data can still be very useful in helping to identify Cancer Control priorities and achievements. Since the partition of Ireland into the independent Republic of Ireland and Northern Ireland, which remained part of the United Kingdom, cancer mortality data have been published in an essentially similar format in both countries. The information presented here will contribute to providing a basis for the collaborative Cancer Research programme initiated recently. PATIENTS AND METHODS: Cancer mortality data have been assembled and analysed separately for the Republic of Ireland and Northern Ireland: the data have then been combined to present mortality rates for the whole of Ireland, covering the period from 1926 to 1995. Several rubrics had to be aggregated to provide data continuously over the time span (e.g. colon and rectum and cervix and body of the uterus). When data were only available in 10-year classes of age, the EM algorithm was employed to obtain 5-year age-specific rates. All rates presented are age-standardised, employing the World Standard Population. RESULTS: In women, the death rate from all neoplasms combined increased very slightly from 117 per 100 000 in 1946-1950 to 120 per 100 000 in 1991-1995. In men, the death rate increased from 127 per 100 000 to 172 per 100 000 over the same time period. The overall cancer death rate in Ireland is currently similar to the European average in men, although in women it is among the top fifth of national cancer mortality rates in European countries. While cancer is a major cause of death in Ireland, there is no evidence of an evolving epidemic building up: the death rates from most forms of cancer are declining towards the end of the time period considered. CONCLUSIONS: As demonstrated by falling death rates from Hodgkin's disease and testicular cancer, major treatment advances appear to have been incorporated effectively into clinical practice in Ireland. Progress is apparent in tobacco control and further initiatives in this area must be undertaken since tobacco appears to be the only major new carcinogen introduced recently into the Irish environment during the period covered by this study. Effective population-based screening programmes for cervix and breast cancer and, more controversially, consideration of a National Prostate Cancer Screening programme, offer scope for further improvement in mortality. Examination of this long time series of mortality data from Ireland provides information about the evolving cancer pattern and provides the necessary background to evaluate the impact of the cross-border cancer research activities now being launched.

Adult↗

Meta-MEME: motif-based hidden Markov models of protein families.

MOTIVATION: Modeling families of related biological sequences using Hidden Markov models (HMMs), although increasingly widespread, faces at least one major problem: because of the complexity of these mathematical models, they require a relatively large training set in order to accurately recognize a given family. For families in which there are few known sequences, a standard linear HMM contains too many parameters to be trained adequately. RESULTS: This work attempts to solve that problem by generating smaller HMMs which precisely model only the conserved regions of the family. These HMMs are constructed from motif models generated by the EM algorithm using the MEME software. Because motif-based HMMs have relatively few parameters, they can be trained using smaller data sets. Studies of short chain alcohol dehydrogenases and 4Fe-4S ferredoxins support the claim that motif-based HMMs exhibit increased sensitivity and selectivity in database searches, especially when training sets contain few sequences.

Alcohol Dehydrogenase↗

A dissimilarity matrix between protein atom classes based on Gaussian mixtures.

MOTIVATION: Previously, Rantanen et al. (2001; J. Mol. Biol., 313, 197-214) constructed a protein atom-ligand fragment interaction library embodying experimentally solved, high-resolution three-dimensional (3D) structural data from the Protein Data Bank (PDB). The spatial locations of protein atoms that surround ligand fragments were modeled with Gaussian mixture models, the parameters of which were estimated with the expectation-maximization (EM) algorithm. In the validation analysis of this library, there was strong indication that the protein atom classification, 24 classes, was too large and that a reduction in the classes would lead to improved predictions. RESULTS: Here, a dissimilarity (distance) matrix that is suitable for comparison and fusion of 24 pre-defined protein atom classes has been derived. Jeffreys' distances between Gaussian mixture models are used as a basis to estimate dissimilarities between protein atom classes. The dissimilarity data are analyzed both with a hierarchical clustering method and independently by using multidimensional scaling analysis. The results provide additional insight into the relationships between different protein atom classes, giving us guidance on, for example, how to readjust protein atom classification and, thus, they will help us to improve protein--ligand interaction predictions. CONTACT: vira@utu.fi

Cluster Analysis↗

Clustering of time-course gene expression data using a mixed-effects model with B-splines.

MOTIVATION: Time-course gene expression data are often measured to study dynamic biological systems and gene regulatory networks. To account for time dependency of the gene expression measurements over time and the noisy nature of the microarray data, the mixed-effects model using B-splines was introduced. This paper further explores such mixed-effects model in analyzing the time-course gene expression data and in performing clustering of genes in a mixture model framework. RESULTS: After fitting the mixture model in the framework of the mixed-effects model using an EM algorithm, we obtained the smooth mean gene expression curve for each cluster. For each gene, we obtained the best linear unbiased smooth estimate of its gene expression trajectory over time, combining data from that gene and other genes in the same cluster. Simulated data indicate that the methods can effectively cluster noisy curves into clusters differing in either the shapes of the curves or the times to the peaks of the curves. We further demonstrate the proposed method by clustering the yeast genes based on their cell cycle gene expression data and the human genes based on the temporal transcriptional response of fibroblasts to serum. Clear periodic patterns and varying times to peaks are observed for different clusters of the cell-cycle regulated genes. Results of the analysis of the human fibroblasts data show seven distinct transcriptional response profiles with biological relevance. AVAILABILITY: Matlab programs are available on request from the authors.

Algorithms↗

Discovering molecular pathways from protein interaction and gene expression data.

In this paper, we describe an approach for identifying 'pathways' from gene expression and protein interaction data. Our approach is based on the assumption that many pathways exhibit two properties: their genes exhibit a similar gene expression profile, and the protein products of the genes often interact. Our approach is based on a unified probabilistic model, which is learned from the data using the EM algorithm. We present results on two Saccharomyces cerevisiae gene expression data sets, combined with a binary protein interaction data set. Our results show that our approach is much more successful than other approaches at discovering both coherent functional groups and entire protein complexes.

Algorithms↗

Genome-wide discovery of transcriptional modules from DNA sequence and gene expression.

In this paper, we describe an approach for understanding transcriptional regulation from both gene expression and promoter sequence data. We aim to identify transcriptional modules--sets of genes that are co-regulated in a set of experiments, through a common motif profile. Using the EM algorithm, our approach refines both the module assignment and the motif profile so as to best explain the expression data as a function of transcriptional motifs. It also dynamically adds and deletes motifs, as required to provide a genome-wide explanation of the expression data. We evaluate the method on two Saccharomyces cerevisiae gene expression data sets, showing that our approach is better than a standard one at recovering known motifs and at generating biologically coherent modules. We also combine our results with binding localization data to obtain regulatory relationships with known transcription factors, and show that many of the inferred relationships have support in the literature.

Algorithms↗

Gene networks inference using dynamic Bayesian networks.

This article deals with the identification of gene regulatory networks from experimental data using a statistical machine learning approach. A stochastic model of gene interactions capable of handling missing variables is proposed. It can be described as a dynamic Bayesian network particularly well suited to tackle the stochastic nature of gene regulation and gene expression measurement. Parameters of the model are learned through a penalized likelihood maximization implemented through an extended version of EM algorithm. Our approach is tested against experimental data relative to the S.O.S. DNA Repair network of the Escherichia coli bacterium. It appears to be able to extract the main regulations between the genes involved in this network. An added missing variable is found to model the main protein of the network. Good prediction abilities on unlearned data are observed. These first results are very promising: they show the power of the learning algorithm and the ability of the model to capture gene interactions.

Algorithms↗

Statistical modeling of sequencing errors in SAGE libraries.

MOTIVATION: Sequencing errors may bias the gene expression measurements made by Serial Analysis of Gene Expression (SAGE). They may introduce non-existent tags at low abundance and decrease the real abundance of other tags. These effects are increased in the longer tags generated in LongSAGE libraries. Current sequencing technology generates quite accurate estimates of sequencing error rates. Here we make use of the sequence neighborhood of SAGE tags and error estimates from the base-calling software to correct for such errors. RESULTS: We introduce a statistical model for the propagation of sequencing errors in SAGE and suggest an Expectation-Maximization (EM) algorithm to correct for them given observed sequences in a library and base-calling error estimates. We tested our method using simulated and experimental SAGE libraries. When comparing SAGE libraries, we found that sequencing errors can introduce considerable bias. High abundance tags may be falsely called as significantly differentially expressed, especially when comparing libraries with different levels of sequencing errors and/or of different size. Truly, differentially expressed tags have decreased significance as 'true'-tag counts are generally underestimated. This may alter if tags near the threshold of differential expression are called significant. Moreover, the number of different transcripts present in a library is overestimated as false tags are introduced at low abundance. Our correction method adjusts the tag counts to be closer to the true counts and is able to partly correct for biases introduced by sequencing errors. AVAILABILITY: An implementation using R is distributed as an R package. An online version is available at http://tagcalling.mbgproject.org

Algorithms↗

A meta-analysis of case-control and cohort studies with interval-censored exposure data: application to chorionic villus sampling.

Chorionic villus sampling (CVS) is a valued method of prenatal diagnosis that is often preferred over amniocentesis because it can be performed earlier, but which has also raised concern over a possible association with increased risk of terminal transverse limb deficiency (TTLD). We present and apply a meta-analytic method for estimating a combined dose-response effect from a series of case-control and cohort studies in which the exposure variable is interval-censored. Assuming coarsening at random for the interval-censoring, and calling upon the familiar result of Cornfield to pool case-control and cohort information on the association between a rare binary outcome and a multilevel exposure variable, we form a likelihood-based model to assess the effect of gestational age at the time of CVS on the presence or absence of a rare birth defect. Effect estimates are computed with a variant of the EM algorithm termed the method of weights, which enables the use of standard weighted regression software. Our findings suggest that CVS exposure at early gestational age leads to an increased risk of TTLD.

Journal Article↗

Maximum likelihood estimation in random effects cure rate models with nonignorable missing covariates.

We introduce a method of parameter estimation for a random effects cure rate model. We also propose a methodology that allows us to account for nonignorable missing covariates in this class of models. The proposed method corrects for possible bias introduced by complete case analysis when missing data are not missing completely at random and is motivated by data from a pair of melanoma studies conducted by the Eastern Cooperative Oncology Group in which clustering by cohort or time of study entry was suspected. In addition, these models allow estimation of cure rates, which is desirable when we do not wish to assume that all subjects remain at risk of death or relapse from disease after sufficient follow-up. We develop an EM algorithm for the model and provide an efficient Gibbs sampling scheme for carrying out the E-step of the algorithm.

Journal Article↗

The joint modeling of a longitudinal disease progression marker and the failure time process in the presence of cure.

In this paper we present an extension of cure models: to incorporate a longitudinal disease progression marker. The model is motivated by studies of patients with prostate cancer undergoing radiation therapy. The patients are followed until recurrence of the prostate cancer or censoring, with the PSA marker measured intermittently. Some patients are cured by the treatment and are immune from recurrence. A joint-cure model is developed for this type of data, in which the longitudinal marker and the failure time process are modeled jointly, with a fraction of patients assumed to be immune from the endpoint. A hierarchical nonlinear mixed-effects model is assumed for the marker and a time-dependent Cox proportional hazards model is used to model the time to endpoint. The probability of cure is modeled by a logistic link. The parameters are estimated using a Monte Carlo EM algorithm. Importance sampling with an adaptively chosen t-distribution and variable Monte Carlo sample size is used. We apply the method to data from prostate cancer and perform a simulation study. We show that by incorporating the longitudinal disease progression marker into the cure model, we obtain parameter estimates with better statistical properties. The classification of the censored patients into the cure group and the susceptible group based on the estimated conditional recurrence probability from the joint-cure model has a higher sensitivity and specificity, and a lower misclassification probability compared with the standard cure model. The addition of the longitudinal data has the effect of reducing the impact of the identifiability problems in a standard cure model and can help overcome biases due to informative censoring.

Journal Article↗

Generalized additive models for cancer mapping with incomplete covariates.

Maps depicting cancer incidence rates have become useful tools in public health research, giving valuable information about the spatial variation in rates of disease. Typically, these maps are generated using count data aggregated over areas such as counties or census blocks. However, with the proliferation of geographic information systems and related databases, it is becoming easier to obtain exact spatial locations for the cancer cases and suitable control subjects. The use of such point data allows us to adjust for individual-level covariates, such as age and smoking status, when estimating the spatial variation in disease risk. Unfortunately, such covariate information is often subject to missingness. We propose a method for mapping cancer risk when covariates are not completely observed. We model these data using a logistic generalized additive model. Estimates of the linear and non-linear effects are obtained using a mixed effects model representation. We develop an EM algorithm to account for missing data and the random effects. Since the expectation step involves an intractable integral, we estimate the E-step with a Laplace approximation. This framework provides a general method for handling missing covariate values when fitting generalized additive models. We illustrate our method through an analysis of cancer incidence data from Cape Cod, Massachusetts. These analyses demonstrate that standard complete-case methods can yield biased estimates of the spatial variation of cancer risk.

Algorithms↗