PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36Linked to original sources

Nonparametric estimation of a multivariate distribution in the presence of censoring.

This paper presents examples of situations in which one wishes to estimate a multivariate distribution from data that may be right-censored. A distinction is made between what we term 'homogeneous' and 'heterogeneous' censoring. It is shown how a multivariate empirical survivor function must be constructed in order to be considered a (nonparametric) maximum likelihood estimate of the underlying survivor function. A closed-form solution, similar to the product-limit estimate of Kaplan and Meier, is possible with homogeneous censoring, but an iterative method, such as the EM algorithm, is required with heterogeneous censoring. An example is given in which an anomaly is produced if censored multivariate data are analyzed as a series of univariate variables; this anomaly is shown to disappear if the methods of this paper are used.

Analysis of Variance↗

Random-effects models for longitudinal data.

Models for the analysis of longitudinal data must recognize the relationship between serial observations on the same unit. Multivariate models with general covariance structure are often difficult to apply to highly unbalanced data, whereas two-stage random-effects models can be used easily. In two-stage models, the probability distributions for the response vectors of different individuals belong to a single family, but some random-effects parameters vary across individuals, with a distribution specified at the second stage. A general family of models is discussed, which includes both growth models and repeated-measures models as special cases. A unified approach to fitting these models, based on a combination of empirical Bayes and maximum likelihood estimation of model parameters and using the EM algorithm, is discussed. Two examples are taken from a current epidemiological study of the health effects of air pollution.

Air Pollution↗

A multiplicative random effects model for meta-analysis with application to estimation of admixture component.

A multiplicative random effects model is constructed for combining results from different studies. An EM algorithm is developed for the maximum likelihood estimator under the Empirical Bayes framework. The method is applied to estimation of the admixture component in admixed populations. The admixture estimates can be useful in genetic epidemiology and DNA fingerprinting.

Black or African American↗

Semi-parametric estimation in failure time mixture models.

A mixture model is an attractive approach for analyzing failure time data in which there are thought to be two groups of subjects, those who could eventually develop the endpoint and those who could not develop the endpoint. The proposed model is a semi-parametric generalization of the mixture model of Farewell (1982). A logistic regression model is proposed for the incidence part of the model, and a Kaplan-Meier type approach is used to estimate the latency part of the model. The estimator arises naturally out of the EM algorithm approach for fitting failure time mixture models as described by Larson and Dinse (1985). The procedure is applied to some experimental data from radiation biology and is evaluated in a Monte Carlo simulation study. The simulation study suggests the semi-parametric procedure is almost as efficient as the correct fully parametric procedure for estimating the regression coefficient in the incidence, but less efficient for estimating the latency distribution.

Animals↗

EM mixed model analysis of data from informatively censored normal distributions.

Maximum likelihood techniques using the EM algorithm are applied to correlated normally distributed survival data from a placebo-controlled, double-blind, dose-ranging crossover study to assess the short-term efficacy of an antianginal drug in patients with chronic stable angina. Censoring was informative and nonterminal and was not due to death or withdrawal from the study. Unlike previous approaches these techniques are mathematically and computationally tractable, do not require computations of high-dimensional integrals, and do not require the inversion of large matrices.

Algorithms↗

Change points in the series of T4 counts prior to AIDS.

The absolute number of T4 cells has been established as an important clinical marker of disease progression to acquired immunodeficiency syndrome (AIDS) in persons infected with human immunodeficiency virus (HIV). Series of T4 counts are analyzed from the 131 homosexual men who entered the New York Blood Center Study in 1984, mostly seropositive for HIV, and who developed AIDS as participants by 1990. These series exhibit a gradual decline of the log(T4) count followed by a more rapid decline close to the time of the development of AIDS. Empirical Bayes and hierarchical Bayes change point models are proposed to estimate the distribution of the time before AIDS when this rapid decline begins. Results using the EM Algorithm and Markov chain Monte Carlo indicate that the mean change point occurs approximately 1 year before diagnosis with a standard deviation of 9 months. Detection of a change point may indicate that an AIDS diagnosis is increasingly likely for an individual HIV-positive but AIDS-free.

Acquired Immunodeficiency Syndrome↗

Analysing incomplete longitudinal binary responses: a likelihood-based approach.

In this paper, we describe a likelihood-based method for analysing balanced but incomplete longitudinal binary responses that are assumed to be missing at random. Following the approach outlined in Zhao and Prentice (1990, Biometrika 77, 642-648), we focus on "marginal models" in which the marginal expectation of the response variable is related to a set of covariates. The association between binary responses is modelled in terms of conditional log odds-ratios. We describe a set of scoring equations for jointly estimating both the marginal parameters and the conditional association parameters. An outline of the EM algorithm used to obtain the maximum likelihood estimates is presented. This approach yields valid and efficient estimates when the responses are missing at random, but not necessarily missing completely at random. An example, using data from the Muscatine Coronary Risk Factor Study, is presented to illustrate this methodology.

Age Factors↗

Models relating the timing of intercourse to the probability of conception and the sex of the baby.

The probability that conception will result from intercourse on particular days of the menstrual cycle is of biological, demographic, and personal interest to many. We describe an existing model for conception as related to the timing of intercourse, and develop a method for fitting it by means of the expectation-maximization (EM) algorithm, using widely available software (GLIM). We then generalize the model to allow for effects on its parameters due to reproductively toxic exposures. A further extension allows one to address the question of whether the pattern of intercourse in a conception cycle has an influence on the sex of the resulting baby. We illustrate the methods by application to data from a prospective study of couples trying to begin a pregnancy.

Abortion, Spontaneous↗

Monte Carlo estimation of mixed models for large complex pedigrees.

In human quantitative genetics, computational complexity restricts the current methods for estimation of mixed models that include major gene effects to data on small pedigrees. However, large complex pedigrees are not uncommon in practice. Also, large pedigrees tend to provide more information on genetic transmission and are more genetically homogeneous than a pooled sample of many nuclear families. We present a Monte Carlo method, using jointly the EM algorithm and the Gibbs sampler, for estimation of mixed models. The approach also provides a Monte Carlo estimate of the asymptotic variance-covariance matrix of the parameters. The methods are conceptually simple, easy to implement, and can handle multiple heritable/nonheritable random components. A numerical example is given to illustrate the methods.

Algorithms↗

Estimating incidence and diagnostic error rates for bivariate progressive processes.

Estimating the times until incidences of bivariate progressive processes that are categorical is a common problem in ophthalmology, audiology, pulmonary medicine, and other fields of medical research. We consider study designs in which diagnoses of subject's bivariate status are performed repeatedly across time and when diagnosis is subject to error. In such situations, error confounds the interpretation of the time until an event. A composite model is proposed for parameterizing both the incidence and error distributions, which allows for correlation between sites with respect to both incidence and diagnostic error. An EM algorithm is described for this model, which allows categorical covariates for both incidence and error. The methodology is applied to two examples. The first represents a situation in which bivariate incidence and error can reasonably be assumed symmetric: prospective data concerning the development of ocular lens opacities in a large pharmaceutical clinical trial. The second example represents a situation in which bivariate incidence and error may not be symmetric: clinical evaluations of sexual maturation status with respect to two different anatomical indices in the Cooperative Study of Sickle Cell Disease. The methodology described in this paper is used, in each case, to estimate incidence, characterize error rates, and assess bivariate correlations.

Adolescent↗

Analysis of infectious disease data from partner studies with unknown source of infection.

Partner studies are useful for estimating the transmission probabilities of infectious diseases. However, it is often not known which partner was the source of the infection (the index case). The objective of this paper is to develop statistical methods for analyzing partner studies when it is uncertain which partners acquired the infection from sources outside the partnership. The approach involves simultaneously modelling the probability of acquisition of infection from outside the partnership, and the probability of transmission within the partnership as a function of covariates. An EM algorithm is presented. Efficiency and simulation results are given in some special situations involving heterosexual transmission studies. In heterosexual partner studies, the methods depend crucially on the availability of a covariate that provides information about which partner was the likely source of infection.

Algorithms↗

Fitting a multiplicative incidence model to age- and time-specific prevalence data.

We discuss the assessment of age- and time-specific disease incidence using prevalence data. A method is described for conveniently fitting a discrete-time multiplicative model, subject to positivity constraints, using the EM-algorithm. Together with smoothing, it allows essentially nonparametric assessment of incidence trends. The method is illustrated using previously analyzed data on toxoplasmosis.

Adolescent↗

[Assessment of pharmacokinetic parameters of amikacin in a group of neutropenic patients in onco-hematology].

The pharmacokinetics of Amikacin were studied in 56 febrile episodes for 45 patients with severe neutropenia while using the USC*Pack PC Clinical Programs for adaptive control of their dosage regimens [223 drug levels]. The purpose of this study are: i] to estimate the pharmacokinetic parameters in this neutropenic population [56 episodes, I], ii] to evaluate the effect of the dosage regimen: once-a-day [22 episodes, II] versus bid or tid [34 episodes, III]. Patients [mean age 53.3 +/- 17.9], 23 men and 22 women, received amikacin [17.7 +/- 3.6 mg/kg/d at day 1] in a 30 minutes infusion. The mean estimated creatinine clearance [CCr] was 76 +/- 22.5 ml/min/1.73 m2 at day 1. The method used for the population modeling was the Non Parametric EM algorithm [NPEM2] which computes the complete probability density function for a 1 or a 2 compartment model. The parametrizations studied are: Clearance/Volume [CL/VOL], Elimination rate constant/Volume [Kel/VOL] and KS/VS with Kel = KS * CCr + 0.00693, VS = VOL/Weight for a 1 compartment pharmacokinetic model. The main results concerned CL and VS with: CL[I] = 4.94 +/- 2.71, CL[II] = 4.74 +/- 2.65, CL[III] = 5.14 +/- 2.75 l/h and VS[I] = 0.31 +/- 0.11, VS[II] = 0.34 +/- 0.10, VS[III] = 0.30 +/- 0.11 l/kg. Volume of distribution VS is not so large as expected and a slight difference appears between II and III. The pharmacokinetic parameters obtained for this population of neutropenic patients will be used thereafter for the daily adaptive control of Amikacine therapy in our haematologic/oncologic patients. The variability observed remains important and requires an individualization of the dosage regimen for each patient.

Adult↗

Parameter estimation from incomplete data in binomial regression when the missing data mechanism is nonignorable.

We propose a method for estimating parameters in binomial regression models when the response variable is missing and the missing data mechanism is nonignorable. We assume throughout that the covariates are fully observed. Using a logit model for the missing data mechanism, we show how parameter estimation can be accomplished using the EM algorithm by the method of weights proposed in Ibrahim (1990, Journal of the American Statistical Association 85, 765-769). An example from the Six Cities Study (Ware et al., 1984, American Review of Respiratory Diseases 129, 366-374) is presented to illustrate the method.

Air Pollution↗

A competing risks analysis of presenting AIDS diagnoses trends.

The proportions of gay men presenting with various AIDS diagnoses display temporal trends. In particular, the proportion of initial diagnoses reported as Kaposi's sarcoma (KS) has declined over time. Epidemiologists have hypothesized that (a) KS may require a cofactor, whose prevalence has declined over time, or (b) KS may have a shorter incubation period than other presenting diagnoses. We examine whether this latter hypothesis, considered in a competing risks framework, could account for the observed decline in KS. We nonparametrically estimate the relevant cause-specific hazard functions from the doubly-censored data of the San Francisco City Clinic Cohort by maximizing a roughness penalized likelihood using an EM algorithm. These estimates suggest that differences in the underlying cause-specific hazard functions account for a substantial portion of the observed diagnoses trends.

Acquired Immunodeficiency Syndrome↗

Semiparametric estimation of major gene and family-specific random effects for age of onset.

Analysis of familial diseases with variable age of onset is a common problem in human genetics. Most existing methods make some parametric distributional assumption on age of onset, and few methods have been designed with the goal of testing the hypothesis of a Mendelian gene against other hypotheses of familial dependence. We introduce the Cox model with major genetic and random familial effects to model age-of-onset dependence patterns among family members and to incorporate family heterogeneity. This model allows testing for and estimating major gene effects in the presence of residual correlations. Generalized maximum likelihood estimation using a Monte Carlo EM algorithm is used for parameter estimation. The methods are illustrated by a simulated data set and a data set from a case-control family study of breast cancer.

Adult↗

Myocardial perfusion imaging with a combined x-ray CT and SPECT system.

UNLABELLED: We evaluated a novel combined x-ray CT and SPECT medical imaging system for quantitative in vivo measurements of 99mTc-sestamibi uptake in an animal model of myocardial perfusion. METHODS: Correlated emission-transmission myocardial images were obtained from 7- to 10-kg pigs. The x-ray CT image was used to generate an object-specific attenuation map that was incorporated into an iterative ML-EM algorithm for reconstruction and attenuation correction of the coregistered SPECT images. The pixel intensities in the SPECT images were calibrated in units of radionuclide concentrations (MBq/g), then compared against in vitro 99mTc activity concentration measured from the excised myocardium. In addition, the coregistered x-ray CT image was used to determine anatomical boundaries for quantitation of myocardial regions with low perfusion. RESULTS: The accuracy of the quantitative measurement of in vivo activity concentration in the porcine myocardium was improved by object-specific attenuation correction. However, an additional correction for partial volume errors was required to retrieve the true activity concentration from the reconstructed SPECT images. CONCLUSION: Accurate absolute SPECT quantitation required object-specific correction for attenuation and partial volume effects. Additional anatomical information from the x-ray CT image was helpful in defining regions of interest for quantitation of the SPECT images.

Algorithms↗

A method for assessing age-time disease incidence using serial prevalence data.

This paper considers nonparametric estimation of age- and time-specific trends in disease incidence using serial prevalence data collected from multiple cross-sectional samples of a population over time. The methodology accounts for differential selection of diseased and undiseased individuals resulting, for example, from differences in mortality. It is shown that when a log-linear incidence odds model is adopted, an EM algorithm provides a convenient method for carrying out maximum likelihood estimation, primarily using existing generalized linear models software. The procedure is quite general, allowing a range of age-time incidence models to be fitted under the same framework. Furthermore, by making use of existing software for fitting generalized additive models, the procedure can be generalized with virtually no extra complexity to allow maximization of a penalized likelihood for smooth nonparametric estimation. Automatic choice of smoothing level for the penalized likelihood estimates is discussed, using generalized cross-validation. The method is applied to a data set on serial toxoplasmosis prevalence, which has previously been analyzed under the assumption of nondifferential selection. A variety of age-time incidence models are fitted, and the sensitivity to plausible differential selection patterns is considered. It is found that nonmultiplicative models are unnecessary and that qualitative incidence trends are fairly robust to differential selection.

Algorithms↗