PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 433 records · Page 24Linked to original sources

Clustering expressed genes on the basis of their association with a quantitative phenotype.

Cluster analyses of gene expression data are usually conducted based on their associations with the phenotype of a particular disease. Many disease traits have a clearly defined binary phenotype (presence or absence), so that genes can be clustered based on the differences of expression levels between the two contrasting phenotypic groups. For example, cluster analysis based on binary phenotype has been successfully used in tumour research. Some complex diseases have phenotypes that vary in a continuous manner and the method developed for a binary trait is not immediately applicable to a continuous trait. However, understanding the role of gene expression in these complex traits is of fundamental importance. Therefore, it is necessary to develop a new statistical method to cluster expressed genes based on their association with a quantitative trait phenotype. We developed a model-based clustering method to classify genes based on their association with a continuous phenotype. We used a linear model to describe the relationship between gene expression and the phenotypic value. The model effects of the linear model (linear regression coefficients) represent the strength of the association. We assumed that the model effects of each gene follow a mixture of several multivariate Gaussian distributions. Parameter estimation and cluster assignment were accomplished via an Expectation-Maximization (EM) algorithm. The method was verified by analysing two simulated datasets, and further demonstrated using real data generated in a microarray experiment for the study of gene expression associated with Alzheimer's disease.

Algorithms↗

Modelling epistatic effects of embryo and endosperm QTL on seed quality traits.

Coordinated expression of embryo and endosperm tissues is required for proper seed development. The coordination among these two tissues is controlled by the interaction between multiple genes expressed in the embryo and endosperm genomes. In this article, we present a statistical model for testing whether quantitative trait loci (QTL) active in different genomes, diploid embryo and triploid endosperm, epistatically affect a trait expressed on the endosperm tissue. The maximum likelihood approach, implemented with the EM algorithm, was derived to provide the maximum likelihood estimates of the locations of embryo- and endosperm-specific QTL and their main effects and epistatic effects. This model was used in a real example for rice in which two QTL, one from the embryo genome and the other from the endosperm genome, exert a significant interaction effect on gel consistency on the endosperm. Our model has successfully detected Waxy, a candidate gene in the embryo genome known to regulate one of the major steps of amylose biosynthesis in the endosperm. This model will have great implications for agricultural and evolutionary genetic research.

Algorithms↗

Statistical approaches to estimating mean water quality concentrations with detection limits.

We review statistical methodology for estimating mean concentrations of potentially toxic pollutants in water, for small samples that are not normally distributed and often contain substantial numbers of nondetects, i.e. samples that are only known to be below some set of fixed thresholds. Maximum likelihood estimation (MLE) and regression on order statistics (ROS) are two main approaches that dominate the literature, with transformation bias under non-normality that increases with the severity of censoring being the main problem. We consider exact maximum likelihood estimators in conjunction with the Box-Cox transformation and propose the Quenouille-Tukey Jackknife as a method for bias reduction and variance estimation. Exact maximum likelihood estimators resulting from the expectation-maximization (EM) algorithm are exhibited in a simple heuristic form that also provides estimated values for the nondetects as subsidiary outputs. We show in simulationsthatthetwo main approaches perform well for the log-normal and gamma distributions as long as the jackknife is employed to reduce bias. Bias corrections to MLE used in the literature are shown to correct in the wrong direction under severe censoring. The jackknife is also used for estimating the variance of the both the MLE and ROS estimators. Robustness is improved by searching a class of power transformations (Box-Cox) for the best approximating normal distribution. We conclude that both the exact MLE and ROS procedures can be useful under varying experimental conditions. Limited simulations indicate that the ROS procedure is unbiased and has a smaller variance than the MLE under the log-normal distribution and is robust. The MLE performed better in simulations involving the gamma as the underlying distribution. We also compare the estimators for the mean and variance that one obtains from typical sets of water quality data, analyzing for copper, alumnium, arsenic, chromium, nickel, and lead.

Forecasting↗

Multivariate life testing in variably scaled environments.

This paper examines modeling and inference questions for experiments in which different subsets of a set of kappa possibly dependent components are tested in r different environments. In each environment, the failure times of the set of components on test is assumed to be governed by a particular type of multivariate exponential (MVE) distribution. For any given component tested in several environments, it is assumed that its marginal failure rate varies from one environment to another via a change of scale between the environments, resulting in a joint MVE model which links in a natural way the applicable MVE distributions describing component behavior in each fixed environment. This study thus extends the work of Proschan and Sullo (1976) to multiple environments and the work of Kvam and Samaniego (1993) to dependent data. The problem of estimating model parameters via the method of maximum likelihood is examined in detail. First, necessary and sufficient conditions for the identifiability of model parameters are established. We then treat the derivation of the MLE via a numerically-augmented application of the EM algorithm. The feasibility of the estimation method is demonstrated in an example in which the likelihood ratio test of the hypothesis of equal component failure rates within any given environment is carried out.

Algorithms↗

Estimation of distribution functions using data from different environments.

Suppose that when a unit operates in a certain environment, its lifetime has distribution G, and when the unit operates in another environment, its lifetime has a different distribution, say F. Moreover, suppose the unit is operated for a certain period of time in the first environment and is then transferred to the second environment. Thus we observe a censored lifetime in the first environment and a failure time of a "used" unit in the second environment. We propose an EM algorithm approach for obtaining a self-consistent estimator of F using observations from both environments. The case where failure times are subject to right censoring is considered as well. We also establish the maximum likelihood estimator of F when the unit is repairable. Application and simulation studies are presented to illustrate the methods derived.

Algorithms↗

A parametric estimation procedure for relapse time distributions.

In a relapse clinical trial patients who have recovered from some recurrent disease (e.g., ulcer or cancer) are examined at a number of predetermined times. A relapse can be detected either at one of these planned inspections or at a spontaneous visit initiated by the patient because of symptoms. In the first case the observations of the time to relapse, X, is interval-censored by two predetermined time-points. In the second case the upper endpoint of the interval is an observation of the time to symptoms, Y. To model the progression of the disease we use a partially observable Markov process. This approach results in a bivariate phase-type distribution for the joint distribution of (X, Y). It is a flexible model which contains several natural distributions for X, and allows the conditional distributions of the marginals to smoothly depend on each other. To estimate the distributions involved we develop an EM-algorithm. The estimation procedure is evaluated and compared with a non-parametric method in a couple of example based on simulated data.

Algorithms↗

Joint modeling of event time and nonignorable missing longitudinal data.

Survival studies usually collect on each participant, both duration until some terminal event and repeated measures of a time-dependent covariate. Such a covariate is referred to as an internal time-dependent covariate. Usually, some subjects drop out of the study before occurrence of the terminal event of interest. One may then wish to evaluate the relationship between time to dropout and the internal covariate. The Cox model is a standard framework for that purpose. Here, we address this problem in situations where the value of the covariate at dropout is unobserved. We suggest a joint model which combines a first-order Markov model for the longitudinally measured covariate with a time-dependent Cox model for the dropout process. We consider maximum likelihood estimation in this model and show how estimation can be carried out via the EM-algorithm. We state that the suggested joint model may have applications in the context of longitudinal data with nonignorable dropout. Indeed, it can be viewed as generalizing Diggle and Kenward's model (1994) to situations where dropout may occur at any point in time and may be censored. Hence we apply both models and compare their results on a data set concerning longitudinal measurements among patients in a cancer clinical trial.

Algorithms↗

A pharmacokinetic model for tenidap in normal volunteers and rheumatoid arthritis patients.

PURPOSE: To develop a pharmacokinetic model for tenidap and to identify important relationships between the pharmacokinetic parameters and available covariates. METHODS: Plasma concentration data from several phase I and phase II studies were used to develop a pharmacokinetic model for tenidap, a novel anti-rheumatic drug. An appropriate pharmacokinetic model was selected on the basis of individual nonlinear regression analyses and an EM algorithm was used to perform a nonlinear mixed-effects analysis. Scatter plots of posterior individual pharmacokinetic parameters were used to identify possible covariate effects. RESULTS: Predicted responses were in good agreement with the observed data. A bi-exponential model with zero order absorption was subsequently used to develop the mixed-effects model. Covariate relationships selected on the basis of differences in the objective function, although statistically significant, were not particularly strong. CONCLUSIONS: The pharmacokinetics of tenidap can be described by a bi-exponential model with zero order absorption. Based on differences in the log-likelihood, significant covariate-parameter relationships were identified between smoking and CL, and between gender and Vss and CLd. Simulated sparse data analyses indicated that the model would be robust for the analysis of sparse data generated in observational studies.

Adult↗

An additive genetic gamma frailty model for linkage analysis of diseases with variable age of onset using nuclear families.

Many late-onset complex diseases exhibit variable age of onset. Efficiently incorporating age of onset information into linkage analysis can potentially increase the power of dissecting complex diseases. In this paper, we treat age of onset as a genetic trait with censored observations. We use multiple markers to infer the inheritance vector at the disease susceptibility (DS) locus in order to extract information about the inheritance pattern of the disease allele in a pedigree. Given the inheritance distribution at the DS locus, we define the genetic frailty for each individual within a nuclear family as the sum of frailties due to a putative major disease gene and a polygenic effect due to any remaining DS loci. Conditioning on these frailties we use the proportional hazards model for the risk of developing disease. We show that a test of linkage can be formulated as a test of zero variance due to a specific locus of the additive gamma frailties. Maximum likelihood estimation, using the EM algorithm, and likelihood ratio tests are employed for parameter estimation and tests of linkage. A simulation study presented indicates that the proposed method is well behaved and can be more powerful than the currently available allele-sharing based linkage methods. A breast cancer data example is used for illustration.

Adult↗

Maximum penalized likelihood estimation in a gamma-frailty model.

The shared frailty models allow for unobserved heterogeneity or for statistical dependence between observed survival data. The most commonly used estimation procedure in frailty models is the EM algorithm, but this approach yields a discrete estimator of the distribution and consequently does not allow direct estimation of the hazard function. We show how maximum penalized likelihood estimation can be applied to nonparametric estimation of a continuous hazard function in a shared gamma-frailty model withright-censored and left-truncated data. We examine the problem of obtaining variance estimators for regression coefficients, the frailty parameter and baseline hazard functions. Some simulations for the proposed estimation procedure are presented. A prospective cohort (Paquid) with grouped survival data serves to illustrate the method which was used to analyze the relationship between environmental factors and the risk of dementia.

Algorithms↗

Regression modeling with recurrent events and time-dependent interval-censored marker data.

In life history studies involving patients with chronic diseases it is often of interest to study the relationship between a marker process and a more clinically relevant response process. This interest may arise from a desire to gain a better understanding of the underlying pathophysiology, a need to evaluate the utility of the marker as a potential surrogate outcome, or a plan to conduct inferences based on joint models. We consider data from a trial of breast cancer patients with bone metastases. In this setting, the marker process is a point process which records the onset times and cumulative number of bone lesions which reflects the extent of metastatic bone involvement The response is also a point process, which records the times patients experience skeletal complications resulting from these bone lesions. Interest lies in assessing how the development of new bone lesions affects the incidence of skeletal complications. By considering the marker as an internal time-dependent covariate in the point process model for skeletal complications we develop and apply methods which allow one to express the association via regression. A complicating feature of this study is that new bone lesions are only detected upon periodic radiographic examination, which makes the marker processes subject to interval-censoring. A modified EM algorithm is used to deal with this incomplete data problem.

Algorithms↗

Correcting the bias in estimation of genetic variances contributed by individual QTL.

In addition to locating chromosomal positions of quantitative trait loci (QTL), estimating the sizes of identified QTL is also an important component in QTL mapping. The size of a QTL is usually measured by the proportion of the phenotypic variance contributed by the QTL. However, the genetic variance may be overestimated in a small line crossing experiment. In this study, we investigate this bias and develop a simple method to correct the bias. The bias correction, however, requires the error of the estimated genetic effect, which is not trivial if the genetic effect is estimated using the Expectation and Maximization (EM) algorithm. Therefore, we also develop a simple method to estimate the standard error of the estimated genetic effect, which is subsequently used to correct the bias in the variance estimate.

Algorithms↗

Two exonic single nucleotide polymorphisms in the microsomal epoxide hydrolase gene are jointly associated with preeclampsia.

This study determined whether genetic variability in exons 3 and 4 of the microsomal epoxide hydrolase gene jointly modifies individual preeclampsia risk. The study also determined whether genetic variability in the gene encoding for microsomal epoxide hydrolase (EPHX) contributes to individual differences in susceptibility to the development of preeclampsia. The study involved 133 preeclamptic and 115 healthy control pregnant women who were genotyped for two single nucleotide polymorphisms (SNPs), T-->C (Tyr113His) in exon 3 and A-->G (His139Arg) in exon 4, in the EPHX gene. Chi-square analysis was used to assess genotype and allele frequency differences between the preeclamptic and control groups. In addition, single-point analysis was expanded to pair of loci haplotype analysis to examine the estimated haplotype frequencies of the two SNPs, of unknown phase, among the preeclamptic and control groups. Estimated haplotype frequencies were assessed using the maximum-likelihood method, employing an expectation-maximization (EM) algorithm. Single-point allele and genotype distributions in exons 3 and 4 of the EPHX gene were not statistically different between the groups. However, according to the haplotype estimation analysis, we observed a significantly elevated frequency of haplotype T-A (Tyr113-His139) among the preeclampsia group vs the control group (P=0.01). The odds ratio for preeclampsia associated with the high-activity haplotype T-A (Tyr113-His139) was 1.61 (95% CI: 1.12-2.32). The use of two intragenic SNPs jointly in haplotype analysis of association demonstrated that the genetically determined high-activity haplotype T-A (Tyr113-His139) was significantly associated with preeclampsia.

Adult↗

Genetic variability of the marine mussel Mytilus galloprovincialis assessed using two-dimensional electrophoresis.

Two-dimensional electrophoresis (2-DE) has been used to measure the degree of genetic variability of the marine mussel Mytilus galloprovincialis. Genetic polymorphisms were detected in 33 of a total of 86 polypeptides scored among the most abundant proteins from foot samples in 38 individuals. Estimates of average heterozygosity were 0.101+/-0.018 and 0.114+/-0.021 in a natural and a cultured population, respectively, from the NW of the Iberian Peninsula. These are the highest estimates of average heterozygosity reported by 2-DE in an animal species to date. We consider that these data throw open the question of the level of genetic variability detectable by two-dimensional electrophoresis. Multilocus genotype data were used to infer haplotypic frequencies by means of the EM algorithm in order to detect linkage disequilibrium between loci coding abundant proteins. Significant associations were found in 22.7% of the 406 two-locus pairs analysed. Also, clusters of loci in which all pairwise combinations exhibit statistically significant associations were detected and physical linkage between some of these loci is postulated from the linkage disequilibrium data.

Animals↗

Sequencing drug response with HapMap.

The information about how DNA sequence varies across the human genome is crucial for unravelling the genetic basis of drug response. A haplotype map, or HapMap, intended to reveal such a variation pattern, has been recently developed by the International HapMap Consortium. Here, we present a conceptual model for directly characterizing specific DNA sequence variants that are responsible for drug response based on the haplotype structure provided by HapMap. Our model is developed in the maximum likelihood context, incorporated by clinically meaningful mathematical functions that model drug response and implemented with the EM algorithm. Our model is employed to a pharmacogenetic study of cardiovascular disease with 107 patients. We found that the haplotype constituted by allele Gly16 (G) at codon 16 and allele Glu27 (G) at codon 27 genotyped within the beta2AR candidate gene exhibits a different effect on heart rate curve from the rest haplotypes. Parents with the diplotype consisting of two copies of haplotype GG are more sensitive in heart rate to increasing dosages of dobutamine than those with other haplotypes. This model provides a powerful tool for elucidating the genetic variants of drug response and ultimately designing personalized medications based on each patient's genetic constitution.

Algorithms↗

The additive genetic gamma frailty model for linkage analysis of age-of-onset variation.

Age of onset is a key factor in the linkage analysis of many complex diseases. Current methods in nonparametric linkage analysis are mainly concentrated on the affected relative pairs or affected family members with age of onset information either ignored or taken into account by specifying age-dependent penetrances for liability classes. On the other hand, gamma frailty models were developed in the biostatistics literature to model familial aggregation of age of onset. However, these frailty models cannot be used directly for linkage analysis. This paper extends the gamma frailty model by incorporating inheritance vector information and provides a semiparametric approach for linkage testing. For a given inheritance vector at the putative disease locus, we construct an additive genetic gamma frailty for each individual within a nuclear family and use the Cox proportional hazard model to model age of onset. We derive the conditional hazard ratio parameter for sib pairs and define a likelihood ratio based LOD score statistic under our model. The EM algorithm is used for estimating the parameters and the maximum likelihood functions. Simulated data sets are used to illustrate these new statistical methods.

Age Factors↗

Accelerated gene counting for haplotype frequency estimation.

Current implementations of the EM algorithm for estimating haplotype frequencies from genotypes on proximal loci require computational resources that grow as nh2k, where n is the number of individuals genotyped and h is the number of haplotypes possible on k loci. For diallelic loci hk=2k. We present an approach whose computational requirement grows as n2t where t is the largest number of loci at which an individual in the sample is heterozygous. The method is illustrated by haplotype frequency estimation from a sample of 45 individuals genotyped at 26 single nucleotide polymorphisms in the PIK3R1 gene.

1-Phosphatidylinositol 4-Kinase↗

Association of single nucleotide polymorphisms of the bile salt export pump gene with intrahepatic cholestasis of pregnancy.

BACKGROUND: We determined whether genetic variability in the gene encoding the bile salt export pump (BSEP) contributes to individual differences in susceptibility to the development of intrahepatic cholestasis of pregnancy (ICP). METHODS: The study involved 57 affected and 115 healthy control pregnant women who were genotyped for two single nucleotide polymorphisms (SNPs) in the BSEP gene. Chi-square analysis was used to assess genotype and allele frequency differences between the cholestatic and control groups. In addition, single locus analysis was expanded to pair of loci haplotype analysis to examine the estimated haplotype frequencies of the two SNPs, of unknown phase, among the cholestatic and control groups. Estimated haplotype frequencies were assessed using the maximum-likelihood method, employing an expectation-maximization (EM) algorithm. RESULTS: The genotype and allele frequency distribution of the two intragenic SNPs in the ICP and control groups revealed significant evidence of association with the exon 28 SNP (P=0.04 and P=0.02, respectively). In addition, a borderline allele association was noted with the intron 19 SNP (P=0.08). Although the overall distribution of estimated haplotypes of intron 19 and exon 28 SNPs did not differ between the ICP and control groups, the most common haplotype, A-G, was significantly overrepresented in the ICP group (P=0.02), at an odds ratio of 1.73 (95% CI: 1.08-2.74). CONCLUSIONS: The use of two intragenic SNPs in both single locus and haplotype analyses of association suggests that the BSEP gene is a susceptibility gene in intrahepatic cholestasis of pregnancy.

ATP Binding Cassette Transporter, Subfamily B, Mem↗