PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 595 records · Page 33Linked to original sources

A statistical framework for genetic association studies of power curves in bird flight.

How the power required for bird flight varies as a function of forward speed can be used to predict the flight style and behavioral strategy of a bird for feeding and migration. A U-shaped curve was observed between the power and flight velocity in many birds, which is consistent to the theoretical prediction by aerodynamic models. In this article, we present a general genetic model for fine mapping of quantitative trait loci (QTL) responsible for power curves in a sample of birds drawn from a natural population. This model is developed within the maximum likelihood context, implemented with the EM algorithm for estimating the population genetic parameters of QTL and the simplex algorithm for estimating the QTL genotype-specific parameters of power curves. Using Monte Carlo simulation derived from empirical observations of power curves in the European starling (Sturnus vulgaris), we demonstrate how the underlying QTL for power curves can be detected from molecular markers and how the QTL detected affect the most appropriate flight speeds used to design an optimal migration strategy. The results from our model can be directly integrated into a conceptual framework for understanding flight origin and evolution.

Journal Article↗

A comparison of continuous- and discrete- time three-state models for rodent tumorigenicity experiments.

The three-state illness-death model provides a useful way to characterize data from a rodent tumorigenicity experiment. Most parametrizations proposed recently in the literature assume discrete time for the death process and either discrete or continuous time for the tumor onset process. We compare these approaches with a third alternative that uses a piecewise continuous model on the hazards for tumor onset and death. All three models assume proportional hazards to characterize tumor lethality and the effect of dose on tumor onset and death rate. All of the models can easily be fitted using an Expectation Maximization (EM) algorithm. The piecewise continuous model is particularly appealing in this context because the complete data likelihood corresponds to a standard piecewise exponential model with tumor presence as a time-varying covariate. It can be shown analytically that differences between the parameter estimates given by each model are explained by varying assumptions about when tumor onsets, deaths, and sacrifices occur within intervals. The mixed-time model is seen to be an extension of the grouped data proportional hazards model [Mutat. Res. 24:267-278 (1981)]. We argue that the continuous-time model is preferable to the discrete- and mixed-time models because it gives reasonable estimates with relatively few intervals while still making full use of the available information. Data from the ED01 experiment illustrate the results.

Animals↗

Polymorphism of the C-reactive protein (CRP) gene is related to serum CRP Level and arterial pulse wave velocity in healthy elderly Japanese.

The objective of this study was to clarify relationships between the C-reactive protein (CRP) gene and both the serum level of CRP and arterial pulse wave velocity (PWV), using haplotype analysis of healthy elderly Japanese. Five single-nucleotide polymorphisms (SNPs) of the human CRP gene (rs1341665, rs3091244, rs1800947, rs1130864 and rs1205) were used to genotype 315 healthy elderly Japanese subjects (mean age, 77.9 +/- 4.1 years; male/female ratio, 0.96). Linkage disequilibrium was analyzed for the five SNPs. The frequency of each haplotype and diplotype was estimated using the expectation/maximization (EM) algorithm. There were statistically significant associations between the CRP level and two CRP genotypes; the p value for the T allele of rs3091244 (CT + AT + TT vs. CC + CA + AA) was 0.002 (95% confidential interval [CI], 2.1-24), and the p value for the T allele of rs1130864 (TT + TC vs. CC) was 0.002 (95% CI, 2.1-24). The only genotype that was significantly associated with arterial PWV was the C allele of rs1800947, with a p value of 0.039. The haplotype was constructed using rs1341665, rs3091244 and rs1800947, in that order. There was a significant association between the CRP level and the T-T-G haplotype, with a p value of 0.002 (95% CI, 2.1-24). There was a significant association between arterial PWV and the C-C-C haplotype, with a p value of 0.039. We concluded that rs3091244, rs1130864 and the T-T-G haplotype are genetic markers for elevated basal CRP levels. rs1800947 and the C-C-C haplotype appear to be susceptibility markers for atherosclerosis, but this requires confirmation.

Age Factors↗

Individual differences in paired comparison data.

Thurstonian models provide a flexible framework for the analysis of multiple paired comparison judgments because they allow a wide range of hypotheses about the judgments' mean and covariance structures to be tested. However, applications have been limited to a large extent by the computational intractability involved in fitting this class of models. This paper demonstrates that the Monte Carlo EM algorithm facilitates maximum likelihood estimation of Thurstonian paired comparison models even when the number of items is large. A paired comparison study is presented in detail to illustrate the estimation approach.

Algorithms↗

Stochastic EM for estimating the parameters of a multilevel IRT model.

An item response theory (IRT) model is used as a measurement error model for the dependent variable of a multilevel model. The dependent variable is latent but can be measured indirectly by using tests or questionnaires. The advantage of using latent scores as dependent variables of a multilevel model is that it offers the possibility of modelling response variation and measurement error and separating the influence of item difficulty and ability level. The two-parameter normal ogive model is used for the IRT model. It is shown that the stochastic EM algorithm can be used to estimate the parameters which are close to the maximum likelihood estimates. This algorithm is easily implemented. The estimation procedure will be compared to an implementation of the Gibbs sampler in a Bayesian framework. Examples using real data are given.

Educational Status↗

Local influence analysis of structural equation models with continuous and ordinal categorical variables.

This paper proposes a method to assess the local influence of minor perturbations for a structural equation model with continuous and ordinal categorical variables. The key idea is to treat the latent variables as hypothetical missing data and then apply Cook's approach to the conditional expectation of the complete-data log-likelihood function in the corresponding EM algorithm for deriving the normal curvature and the conformal normal curvature. Building blocks for achieving the diagnostic measures are computed via observations generated by the Gibbs sampler. It is shown that the proposed methodology is relatively simple to implement, computationally efficient, and feasible for a wide variety of perturbation schemes. Two illustrative real examples are presented.

Acquired Immunodeficiency Syndrome↗

Joint mapping of quantitative trait Loci for multiple binary characters.

Joint mapping for multiple quantitative traits has shed new light on genetic mapping by pinpointing pleiotropic effects and close linkage. Joint mapping also can improve statistical power of QTL detection. However, such a joint mapping procedure has not been available for discrete traits. Most disease resistance traits are measured as one or more discrete characters. These discrete characters are often correlated. Joint mapping for multiple binary disease traits may provide an opportunity to explore pleiotropic effects and increase the statistical power of detecting disease loci. We develop a maximum-likelihood method for mapping multiple binary traits. We postulate a set of multivariate normal disease liabilities, each contributing to the phenotypic variance of one disease trait. The underlying liabilities are linked to the binary phenotypes through some underlying thresholds. The new method actually maps loci for the variation of multivariate normal liabilities. As a result, we are able to take advantage of existing methods of joint mapping for quantitative traits. We treat the multivariate liabilities as missing values so that an expectation-maximization (EM) algorithm can be applied here. We also extend the method to joint mapping for both discrete and continuous traits. Efficiency of the method is demonstrated using simulated data. We also apply the new method to a set of real data and detect several loci responsible for blast resistance in rice.

Algorithms↗

Mapping quantitative trait Loci interactions from the maternal and offspring genomes.

The expression of most developmental or behavioral traits involves complex interactions between quantitative trait loci (QTL) from the maternal and offspring genomes. The maternal-offspring interactions play a pivotal role in shaping the direction and rate of evolution in terms of their substantial contribution to quantitative genetic (co)variation. To study the genetics and evolution of maternal-offspring interactions, a unifying statistical framework that embraces both the direct and indirect genetic effects of maternal and offspring QTL on any complex trait is developed. This model is derived for a simple backcross design within the maximum-likelihood context, implemented with the EM algorithm. Results from extensive simulations suggest that this model can provide reasonable estimation of additive and dominant effects of the QTL at different generations and their interaction effects derived from the maternal and offspring genomes. Although our model is framed to characterize the actions and interactions of maternal and offspring QTL affecting offspring traits, the idea can be readily extended to decipher the genetic machinery of maternal traits, such as maternal care. Our model provides a powerful means for studying the evolutionary significance of indirect genetic effects in any sexually reproductive organisms.

Animals↗

Sequencing complex diseases With HapMap.

Determining the patterns of DNA sequence variation in the human genome is a useful first step toward identifying the genetic basis of a common disease. A haplotype map (HapMap), aimed at describing these variation patterns across the entire genome, has been recently developed by the International HapMap Consortium. In this article, we present a novel statistical model for directly characterizing specific sequence variants that are responsible for disease risk based on the haplotype structure provided by HapMap. Our model is developed in the maximum-likelihood context, implemented with the EM algorithm. We perform simulation studies to investigate the statistical properties of this disease-sequencing model. A worked example from a human obesity study with 155 patients was used to validate this model. In this example, we found that patients carrying a haplotype constituted by allele Gly16 at codon 16 and allele Gln27 at codon 27 genotyped within the beta2AR candidate gene display significantly lower body mass index than patients carrying the other haplotypes. The implications and extensions of our model are discussed.

Body Mass Index↗

Reconstituting the frequency spectrum of ascertained single-nucleotide polymorphism data.

Most of the available SNP data have eluded valid population genetic analysis because most population genetical methods do not correctly accommodate the special discovery process used to identify SNPs. Most of the available SNP data have allele frequency distributions that are biased by the ascertainment protocol. We here show how this problem can be corrected by obtaining maximum-likelihood estimates of the true allele frequency distribution. In simple cases, the ML estimate of the true allele frequency distribution can be obtained analytically, but in other cases computational methods based on numerical optimization or the EM algorithm must be used. We illustrate the new correction method by analyzing some previously published SNP data from the SNP Consortium. Appropriate treatment of SNP ascertainment is vital to our ability to make correct inferences from the data of the International HapMap Project.

Alleles↗

A unified statistical model for functional mapping of environment-dependent genetic expression and genotype x environment interactions for ontogenetic development.

The effects of quantitative trait loci (QTL) on phenotypic development may depend on the environment (QTL x environment interaction), other QTL (genetic epistasis), or both. In this article, we present a new statistical model for characterizing specific QTL that display environment-dependent genetic expressions and genotype x environment interactions for developmental trajectories. Our model was derived within the maximum-likelihood-based mixture model framework, incorporated by biologically meaningful growth equations and environment-dependent genetic effects of QTL, and implemented with the EM algorithm. With this model, we can characterize the dynamic patterns of genetic effects of QTL governing growth curves and estimate the global effect of the underlying QTL during the course of growth and development. In a real example with rice, our model has successfully detected several QTL that produce differences in their genetic expression between two contrasting environments. These detected QTL cause significant genotype x environment interactions for some fundamental aspects of growth trajectories. The model provides the basis for deciphering the genetic architecture of trait expression adjusted to different biotic and abiotic environments and genetic relationships for growth rates and the timing of life-history events for any organism.

Chromosome Mapping↗

A model selection-based interval-mapping method for autopolyploids.

While extensive progress has been made in quantitative trait locus (QTL) mapping for diploid species, similar progress in QTL mapping for polyploids has been limited due to the complex genetic architecture of polyploids. To date, QTL mapping in polyploids has focused mainly on tetraploids with dominant and/or codominant markers. Here, we extend this view to include any even ploidy level under a dominant marker system. Our approach first selects the most likely chromosomal marker configurations using a Bayesian selection criterion and then fits an interval-mapping model to each candidate. Profiles of the likelihood-ratio test statistic and the maximum-likelihood estimates (MLEs) of parameters including QTL effects are obtained via the EM algorithm. Putative QTL are then detected using a resampling-based significance threshold, and the corresponding parental configuration is identified to be the underlying parental configuration from which the data are observed. Although presented via pseudo-doubled backcross experiments, this approach can be readily extended to other breeding systems. Our method is applied to single-dose restriction fragment autotetraploid alfalfa data, and the performance is investigated through simulation studies.

Algorithms↗

A general framework for statistical linkage analysis in multivalent tetraploids.

In multivalent polyploids, simultaneous pairings among homologous chromosomes at meiosis result in a unique cytological phenomenon-double reduction. Double reduction casts an impact on chromosome evolution in higher plants, but because of its confounded effect on the pattern of gene cosegregation, it complicates linkage analysis and map construction with polymorphic molecular markers. In this article, we have proposed a general statistical model for simultaneously estimating the frequencies of double reduction, the recombination fraction, and optimal parental linkage phases between any types of markers, both fully and partially informative, or dominant and codominant, for a tetraploid species that undergoes only multivalent pairing. This model provides an in-depth extension of our earlier linkage model that was built upon Fisher's classifications for different gamete formation modes during the polysomic inheritance of a multivalent polyploid. By implementing a two-stage hierarchical EM algorithm, we derived a closed-form solution for estimating the frequencies of double reduction through the estimation of gamete mode frequencies and the recombination fraction. We performed different settings of simulation studies to demonstrate the statistical properties of our model for estimating and testing double reduction and the linkage in multivalent tetraploids. As shown by a comparative analysis, our model provides a general framework that covers existing statistical approaches for linkage mapping in polyploids that are predominantly multivalent. The model will have great implications for understanding the genome structure and organization of polyploid species.

Algorithms↗

Theoretical basis for the identification of allelic variants that encode drug efficacy and toxicity.

Almost all drugs that produce a favorable response (efficacy) may also produce adverse effects (toxicity). The relative strengths of drug efficacy and toxicity that vary in human populations are controlled by the combined influences of multiple genes and environmental influences. Genetic mapping has proven to be a powerful tool for detecting and identifying specific DNA sequence variants on the basis of the haplotype map (HapMap) constructed from single-nucleotide polymorphisms (SNPs). In this article, we present a novel statistical model for sequence mapping of two different but related drug responses. This model is incorporated by mathematical functions of drug response to varying doses or concentrations and the statistical device used to model the correlated structure of the residual (co)variance matrix. We implement a closed-form solution for the EM algorithm to estimate the population genetic parameters of SNPs and the simplex algorithm to estimate the curve parameters describing the pharmacodynamic changes of different genetic variants and matrix-structuring parameters. Extensive simulations are performed to investigate the statistical properties of our model. The implications of our model in pharmacogenetic and pharmacogenomic research are discussed.

Algorithms↗

Functional mapping of quantitative trait loci that interact with the hg mutation to regulate growth trajectories in mice.

The high growth (hg) mutation increases body size in mice by 30-50%. Given the complexity of the genetic regulation of animal growth, it is likely that the effect of this major locus is mediated by other quantitative trait loci (QTL) with smaller effects within a web of gene interactions. In this article, we extend our functional mapping model to characterize modifier QTL that interact with the hg locus during ontogenetic growth. Our model is derived within the maximum-likelihood context, incorporated by mathematical aspects of growth laws and implemented with the EM algorithm. In an F2 population founded by a congenic high growth (HG) line and non-HG line, a highly additive effect due to the hg gene was detected on growth trajectories. Three QTL located on chromosomes 2 and X were identified to trigger significant additive and/or dominant effects on the process of growth. The most significant finding made from our model is that these QTL interact with the hg locus to affect the shapes of the growth process. Our model provides a powerful means for understanding the genetic architecture and regulation of growth rate and body size in mammals.

Age Factors↗

Improvement of mapping accuracy by unifying linkage and association analysis.

It is well known that pedigree/family data record information on the coexistence in founder haplotypes of alleles at nearby loci and the cotransmission from parent to offspring that reveal different, but complementary, profiles of the genetic architecture. Either conventional linkage analysis that assumes linkage equilibrium or family-based association tests (FBATs) capture only partial information, leading to inefficiency. For example, FBATs will fail to detect even very tight linkage in the case where no allelic association exists, while a violation of the assumption of linkage equilibrium will result in biased estimation and reduced efficiency in linkage mapping. In this article, by using a data augmentation technique and the EM algorithm, we propose a likelihood-based approach that embeds both linkage and association analyses into a unified framework for general pedigree data. Relative to either linkage or association analysis, the proposed approach is expected to have greater estimation accuracy and power. Monte Carlo simulations support our theoretical expectations and demonstrate that our new methodology: (1) is more powerful than either FBATs or classic linkage analysis; (2) can unbiasedly estimate genetic parameters regardless of whether association exists, thus remedying the bias and less precision of traditional linkage analysis in the presence of association; and (3) is capable of identifying tight linkage alone. The new approach also holds the theoretical advantage that it can extract statistical information to the maximum extent and thereby improve mapping accuracy and power because it integrates multilocus population-based association study and pedigree-based linkage analysis into a coherent framework. Furthermore, our method is numerically stable and computationally efficient, as compared to existing parametric methods that use the simplex algorithm or Newton-type methods to maximize high-order multidimensional likelihood functions, and also offers the computation of Fisher's information matrix. Finally, we apply our methodology to a genetic study on bone mineral density (BMD) for the vitamin D receptor (VDR) gene and find that VDR is significantly linked to BMD at the one-third region of the wrist.

Algorithms↗

Mapping quantitative trait loci for longitudinal traits in line crosses.

Quantitative traits whose phenotypic values change over time are called longitudinal traits. Genetic analyses of longitudinal traits can be conducted using any of the following approaches: (1) treating the phenotypic values at different time points as repeated measurements of the same trait and analyzing the trait under the repeated measurements framework, (2) treating the phenotypes measured from different time points as different traits and analyzing the traits jointly on the basis of the theory of multivariate analysis, and (3) fitting a growth curve to the phenotypic values across time points and analyzing the fitted parameters of the growth trajectory under the theory of multivariate analysis. The third approach has been used in QTL mapping for longitudinal traits by fitting the data to a logistic growth trajectory. This approach applies only to the particular S-shaped growth process. In practice, a longitudinal trait may show a trajectory of any shape. We demonstrate that one can describe a longitudinal trait with orthogonal polynomials, which are sufficiently general for fitting any shaped curve. We develop a mixed-model methodology for QTL mapping of longitudinal traits and a maximum-likelihood method for parameter estimation and statistical tests. The expectation-maximization (EM) algorithm is applied to search for the maximum-likelihood estimates of parameters. The method is verified with simulated data and demonstrated with experimental data from a pseudobackcross family of Populus (poplar) trees.

Chromosome Mapping↗

On the generalized poisson regression mixture model for mapping quantitative trait loci with count data.

Statistical methods for mapping quantitative trait loci (QTL) have been extensively studied. While most existing methods assume normal distribution of the phenotype, the normality assumption could be easily violated when phenotypes are measured in counts. One natural choice to deal with count traits is to apply the classical Poisson regression model. However, conditional on covariates, the Poisson assumption of mean-variance equality may not be valid when data are potentially under- or overdispersed. In this article, we propose an interval-mapping approach for phenotypes measured in counts. We model the effects of QTL through a generalized Poisson regression model and develop efficient likelihood-based inference procedures. This approach, implemented with the EM algorithm, allows for a genomewide scan for the existence of QTL throughout the entire genome. The performance of the proposed method is evaluated through extensive simulation studies along with comparisons with existing approaches such as the Poisson regression and the generalized estimating equation approach. An application to a rice tiller number data set is given. Our approach provides a standard procedure for mapping QTL involved in the genetic control of complex traits measured in counts.

Algorithms↗