PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 613 records · Page 34Linked to original sources

A bivalent polyploid model for mapping quantitative trait loci in outcrossing tetraploids.

Two major aspects have made the genetic and genomic study of polyploids extremely difficult. First, increased allelic or nonallelic combinations due to multiple alleles result in complex gene actions and interactions for quantitative trait loci (QTL) in polyploids. Second, meiotic configurations in polyploids undergo a complex biological process including either bivalent or multivalent formation, or both. For bivalent polyploids, different degrees of preferential chromosome pairings may occur during meiosis. In this article, we develop a maximum-likelihood-based model for mapping QTL in tetraploids by considering the quantitative inheritance and meiotic mechanism of bivalent polyploids. This bivalent polyploid model is implemented with the EM algorithm to simultaneously estimate QTL position, QTL effects, and QTL-marker linkage phases by incorporating the impact of a cytological parameter determining bivalent chromosome pairings (the preferential pairing factor). Simulation studies are performed to investigate the performance and robustness of our statistical method for parameter estimation. The implication and extension of the bivalent polyploid model are discussed.

Algorithms↗

A general framework for analyzing the genetic architecture of developmental characteristics.

The genetic architecture of growth traits plays a central role in shaping the growth, development, and evolution of organisms. While a limited number of models have been devised to estimate genetic effects on complex phenotypes, no model has been available to examine how gene actions and interactions alter the ontogenetic development of an organism and transform the altered ontogeny into descendants. In this article, we present a novel statistical model for mapping quantitative trait loci (QTL) determining the developmental process of complex traits. Our model is constructed within the traditional maximum-likelihood framework implemented with the EM algorithm. We employ biologically meaningful growth curve equations to model time-specific expected genetic values and the AR(1) model to structure the residual variance-covariance matrix among different time points. Because of a reduced number of parameters being estimated and the incorporation of biological principles, the new model displays increased statistical power to detect QTL exerting an effect on the shape of ontogenetic growth and development. The model allows for the tests of a number of biological hypotheses regarding the role of epistasis in determining biological growth, form, and shape and for the resolution of developmental problems at the interface with evolution. Using our newly developed model, we have successfully detected significant additive x additive epistatic effects on stem height growth trajectories in a forest tree.

Algorithms↗

Mapping quantitative trait loci in F2 incorporating phenotypes of F3 progeny.

In plants and laboratory animals, QTL mapping is commonly performed using F(2) or BC individuals derived from the cross of two inbred lines. Typical QTL mapping statistics assume that each F(2) individual is genotyped for the markers and phenotyped for the trait. For plant traits with low heritability, it has been suggested to use the average phenotypic values of F(3) progeny derived from selfing F(2) plants in place of the F(2) phenotype itself. All F(3) progeny derived from the same F(2) plant belong to the same F(2:3) family, denoted by F(2:3). If the size of each F(2:3) family (the number of F(3) progeny) is sufficiently large, the average value of the family will represent the genotypic value of the F(2) plant, and thus the power of QTL mapping may be significantly increased. The strategy of using F(2) marker genotypes and F(3) average phenotypes for QTL mapping in plants is quite similar to the daughter design of QTL mapping in dairy cattle. We study the fundamental principle of the plant version of the daughter design and develop a new statistical method to map QTL under this F(2:3) strategy. We also propose to combine both the F(2) phenotypes and the F(2:3) average phenotypes to further increase the power of QTL mapping. The statistical method developed in this study differs from published ones in that the new method fully takes advantage of the mixture distribution for F(2:3) families of heterozygous F(2) plants. Incorporation of this new information has significantly increased the statistical power of QTL detection relative to the classical F(2) design, even if only a single F(3) progeny is collected from each F(2:3) family. The mixture model is developed on the basis of a single-QTL model and implemented via the EM algorithm. Substantial computer simulation was conducted to demonstrate the improved efficiency of the mixture model. Extension of the mixture model to multiple QTL analysis is developed using a Bayesian approach. The computer program performing the Bayesian analysis of the simulated data is available to users for real data analysis.

Algorithms↗

[Impact of antiretroviral therapy on the magnitude of the HIV/AIDS epidemic in Brazil: various scenarios].

We applied the back-calculation method to estimate the magnitude of the HIV epidemic in Brazil, using the EM and EMS algorithms. Under certain assumptions regarding the behavior of infected patients towards combined antiretroviral therapy, we discuss five different scenarios applied to the Brazilian epidemic. Our objective was to illustrate the impact of combined antiretroviral treatment on the incubation period and thus on estimates of the size of the HIV-infected population, based on reported AIDS cases.

Acquired Immunodeficiency Syndrome↗

Lack of an association between a newly identified promoter polymorphism (-1702G > A) of the leukotriene C4 synthase gene and aspirin-intolerant asthma in a Korean population.

Aspirin-intolerant asthma (AIA) is a distinct clinical syndrome that refers to the development of bronchoconstriction in asthmatic individuals following the ingestion of aspirin and other nonsteroidal anti-inflammatory drugs (NSAIDs). It is widely recognized that increased cysteinyl leukotriene (cysLT) biosynthesis is associated with the development and progression of AIA. Leukotriene C4 synthase (LTC4S) is the terminal enzyme in cysLT production and is a strong candidate gene in the pathogenesis of aspirin-intolerant asthma (AIA). In this paper, we report a new single nucleotide polymorphism (SNP) of the LTC4S promoter, -1702G>A, in AIA patients and evaluate its genetic role in the association with the LTC4S-444 A>C polymorphism. We enrolled 110 AIA patients, 125 aspirin-tolerant asthma (ATA) patients, and 125 normal controls. SNP genotyping of the LTC4S-1702G>A and -444A>C polymorphisms was performed using SNP-IT assays. Haplotype analyses were performed using Haploview version 2.05, which is based on an estimation-maximization (EM) algorithm. There were no significant differences in the allele or genotype frequencies of the LTC4S-1702G>A and -444A>C polymorphisms among the three groups (p > 0.05), with no significant differences in the observed haplotype frequencies (p > 0.05). Moreover, no significant associations were found between the genotype of each SNP in AIA patients with the clinical characteristics, including a forced expiratory volume in one second (FEV1) %, a provocation concentration of methacholine to induce more than 20% decrease of FEV1 (PC20) to methacholine, and serum total IgE levels (p > 0.05). These results indicate that there is no association between these two promoter polymorphisms of LTC4S and the phenotype of AIA in a Korean population.

Adult↗

On-line estimation of concentration parameters in fermentation processes.

It has long been thought that bioprocess, with their inherent measurement difficulties and complex dynamics, posed almost insurmountable problems to engineers. A novel software sensor is proposed to make more effective use of those measurements that are already available, which enable improvement in fermentation process control. The proposed method is based on mixtures of Gaussian processes (GP) with expectation maximization (EM) algorithm employed for parameter estimation of mixture of models. The mixture model can alleviate computational complexity of GP and also accord with changes of operating condition in fermentation processes, i.e., it would certainly be able to examine what types of process-knowledge would be most relevant for local models' specific operating points of the process and then combine them into a global one. Demonstrated by on-line estimate of yeast concentration in fermentation industry as an example, it is shown that soft sensor based state estimation is a powerful technique for both enhancing automatic control performance of biological systems and implementing on-line monitoring and optimization.

Algorithms↗

Renal elimination of amikacin and the aging process.

OBJECTIVE: Although amikacin is primarily eliminated via glomerular filtration, drug concentrations are not consistently predicted in all patients. To better describe the relationship between amikacin clearance and both age and renal function, we used a new heuristic approach involving statistical analysis of dependence. DESIGN AND SETTING: Retrospective pharmacokinetic study using data from seven centres in France. PARTICIPANTS: 634 patients with sepsis aged between 18 and 98 years of age who received intravenous amikacin. METHODS: Clearance of amikacin was modelled using the NonParametric EM algorithm for a two-compartment model (NPEM2) with intravenous infusion. RESULTS: A total of 2499 serum amikacin determinations was available for analysis. The relationship between the clearance of amikacin and age was weak. Interestingly, the Z method, which filters data based on dependence criteria, selected data that were best fitted by a polynomial function (r = 0.90; p < 0.001). This representation of the polynomial function was similar to a previously proposed theoretical model describing covariations between the clearance of amikacin and age. However, the polynomial function applied to only 33% of the patients that were selected by the Z method. The correlation between the clearance of amikacin and renal function was also relatively low (r = 0.39). The Z method exhibited a continuous and strong dependence pattern between the clearance of amikacin and age for 49% of the patients. CONCLUSIONS: The Z methodology, which filters data using dependence criteria, confirms that age, renal function and amikacin clearance are strongly related, but only in less than half of a large sample of patients with sepsis without renal pathology. These results suggest that other variables should be taken into account in order to improve the description of the behaviour of amikacin. The Z methodology improved the classical description of relationships between variables, and should be applied to better select pertinent variables in pharmacokinetic studies.

Adult↗

A method for evaluating the impact of individual haplotypes on disease incidence in molecular epidemiology studies.

Estimation of the association between haplotypes and disease from a case-control study is considered. Assuming a single "disease haplotype'' leads to the increased risk, attention focusses on the relative risks associated with a single copy, or two copies of the disease haplotype, relative to individuals with no copies. In this setting, case frequencies of the haplotype pairs are in Hardy-Weinberg Equilibrium (HWE) only if the combined influence of the two copies of the disease haplotype on risk is multiplicative. Thus, imputation cannot rely on the assumption of HWE for cases. A method is presented for obtaining estimates of the relative risks, making use of the EM algorithm and the assumption of HWE only for controls. The method accounts for the additional variation in the estimates due to the imputation of expected frequencies of haplotype pairs from ambiguous genotypes. A simulation study shows that the resulting confidence intervals have nominal coverage, and that the methods based on the assumption of HWE for both cases and controls can lead to bias.

Journal Article↗

Estimating motifs under order restrictions.

Transcription factors and many other DNA-binding proteins recognize more than one specific sequence. Among sequences recognized by a given DNA-binding protein, different positions exhibit varying degrees of conservation. The reason is that base pairs that are more extensively contacted by the protein tend to be more conserved. This observation can be used in the discovery of transcription factor binding sites. Here we present a rigorous means to accomplish this. In particular, we constrain the order of the information (entropy) in the columns of the position specific weight matrix (PWM) which characterizes the motif being sought. We then show how to compute the maximum likelihood estimate of a PWM under such order restrictions. This computation is easily integrated with the EM algorithm or the Gibbs sampler to enhance performance in the search for motifs in unaligned sequences. We demonstrate our method on a well-known data set of binding sites of the transcription factor Crp in E. coli.

Journal Article↗

Pseudo-likelihood for non-reversible nucleotide substitution models with neighbour dependent rates.

In the field of molecular evolution genome substitution models with neighbour dependent substitution rates have recently received much attention. It is well-known that substitution of nucleotides does not occur independently of neighbouring nucleotides, but there has been less focus on the phenomenon that this substitution process is also not time-reversible. In this paper I construct a pseudo-likelihood type method for inference in non-reversible substitution models with neighbour dependent substitution rates. I also construct an EM-algorithm for maximising the pseudo-likelihood. For human-mouse aligned sequence data a number of different models are investigated, where I show that strand-symmetric models are appropriate, and that overlapping di-nucleotide models do not fit the data well.

Algorithms↗

Restricted maximum likelihood procedures for the estimation of additive and nonadditive genetic variances and covariances in multibreed populations.

Restricted maximum-likelihood procedures were developed to estimate additive and nonadditive genetic and environmental covariances for multiple traits in multibreed populations. The computational procedure follows the expectation-maximization (EM) algorithm, where the set of equations in the maximization step is solved by successive approximations. This computational procedure does not guarantee convergence to a symmetric positive-definite covariance matrix. Thus, computer programs will need to incorporate restrictions in the maximization step to ensure positive definiteness of each covariance matrix. Additive genetic and environmental covariances were modeled in subclass form (zeros and ones in the design matrices). Nonadditive genetic covariances were modeled in regression form (any value between and including zero and one in the design matrices). Computational requirements will be larger than for intrabreed analyses. Appropriate simplifying assumptions and numerical techniques (e.g., sparse and iterative numerical techniques) will be required for the implementation of these multibreed covariance estimation procedures. Number of iterations (5 to 12) and computing times (57 to 113 min) to achieve convergence when estimating 21 genetic and environmental covariances in five small simulated multibreed data sets (two breeds, 25,200 to 50,400 calves, 120 to 135 unrelated bulls) suggest that these procedures are computationally feasible.

Algorithms↗

Estimation of variance and covariance components to determine heritabilities and repeatability of weaning weight in American Simmental cattle.

Components of (co)variance for weaning weight were estimated from field data provided by the American Simmental Association. These components were obtained for the observational components of variance corresponding to a sire, maternal grandsire, and dam within maternal grandsire model. From these estimates, direct additive genetic variance (Sigma2A), maternal additive genetic variance (Sigma2M), covariance between direct and maternal additive genetic effects (SigmaAM), variance of permanent environment(Sigma2pe) and temporary environment variance(Sigma2te) were determined. A procedure to approximate restricted maximum likelihood (REML) estimates of the observational components of variance based on the expectation-maximization (EM) algorithm is described. From these results, phenotypic variance ( ) of weaning weight was 667.88 kg2. Values forSigma2A, Sigma2M, Sigma2pe and Sigma2te were 79,30,58,38,49.45, and 469.97 kg2, respectively. Genetic correlation between direct and maternal additive genetic effects was .16.

Animals↗

Comparison of IgM capture ELISA with a commercial rapid immunochromatographic card test & IgM microwell ELISA for the detection of antibodies to dengue viruses.

BACKGROUND & OBJECTIVES: There is a need for a reliable test for the early diagnosis of dengue fever (DF), which is now active in many parts of India especially in the post monsoon months. This study evaluated two commercial tests with an assay available from a national laboratory in India to obtain information and to make a comparison among the three tests as to which will be the most suited for the detection of IgM antibodies to dengue virus. METHODS: An IgM capture ELISA (National Institute of Virology, Pune, India) was compared with two commercial tests, the PanBio Rapid Immunochromatographic Card Test (Brisbane, Australia) and the PanBio Microwell IgM ELISA for the detection of IgM antibodies to dengue virus. We tested 154 samples from individuals with febrile illnesses having DF--like symptoms. RESULTS: The NIV IgM capture ELISA (MAC-ELISA) showed a high positivity rate (38.9%) as compared to the PanBio Rapid (22.7%) and the PanBio IgM ELISA (20.7%). The true prevalence of disease, sensitivity and specificity of the three tests were estimated using 2LC latent class models using expectation-maximization (EM) algorithm. The NIV MAC-ELISA showed a high sensitivity (96%) as compared to PanBio Rapid (73%) and PanBio IgM ELISA (72%). A subset of 68 samples (of the 154 tested) were analyzed by the NIV MAC-ELISA for IgM antibodies additionally to Japanese encephalitis (JE) and West Nile (WN) of which 31 samples showed positivity to either one, two or all three flaviviruses. Out of the 8 samples which were positive for dengue IgM alone by the NIV MAC-ELISA, only 2 (25%) each were picked up by the other 2 tests. While out of 7 samples positive for IgM to all three flaviviruses IgM by the NIV MAC-ELISA, 5 (71%) were picked up by the other 2 tests. Of the 5 that were picked up by the PanBio tests, 3 had the highest absorbance values to WN by the NIV MAC-ELISA, indicating cross reactivity by PanBio tests. INTERPRETATION & CONCLUSION: The MAC-ELISA though a 3 day procedure, would be a valuable screening test for the detection of IgM to dengue in routine diagnostic laboratories because of its high sensitivity and specificity rates. The test uses specific viral antigens to detect IgM antibodies not only to dengue but also to JE and West Nile as a result of which IgM antibodies to all the 3 commonly encountered flaviviruses can be detected in a single run. It also has the advantage in that depending on the strength of the antibody units obtained to a specific flaviviral antigen, presumptive diagnosis as to which particular arboviral infection has occurred can be made in conjunction with clinical presentation.

Antibodies, Viral↗

Adverse effects of sulfasalazine in patients with rheumatoid arthritis are associated with diplotype configuration at the N-acetyltransferase 2 gene.

OBJECTIVE: N-acetyltransferase 2 (NAT2) is a key enzyme for the acetylation of sulfasalazine (SSZ). We examine whether there was a correlation between diplotype configurations (combinations of 2 haplotypes for a subject) at the NAT2 gene and the adverse effects of SSZ used for the treatment of rheumatoid arthritis (RA). METHODS: The findings from 144 patients with RA who had been treated with SSZ were collected from our outpatient department and used for a retrospective study. Haplotype analysis was performed by the maximum-likelihood estimation based on the EM algorithm using the obtained polymorphism data. RESULTS: Sixteen patients (11.1%) had experienced adverse effects from SSZ, the most common being allergic reactions including rash and fever. The slow acetylators who had no NAT2*4 haplotype had experienced adverse effects more frequently (62.5%) than the fast acetylators who had at least one NAT2*4 haplotype (8.1%) (p < 0.001, OR 7.73, 95% CI 3.54-16.86). In 25% of the slow acetylators, the adverse effects were so severe that they were hospitalized. CONCLUSION: Genotyping the NAT2 gene followed by estimation of diplotype configuration before administration of SSZ is likely to reduce the frequency of adverse effects in Japanese patients with RA.

Acetylation↗

[The relationship between haplotypes of angiotensinogen gene and essential hypertension].

OBJECTIVE: To investigate the relationship between the polymorphism of angiotensinogen gene (AGT) and the risk for hypertension in a Chinese population. METHODS: Three polymorphisms of AGT gene were analyzed in 335 patients with documented essential hypertension and 196 control subjects by using PCR-restriction fragment length polymorphism. Expectation maximization(EM) algorithm was then used for pairwise linkage disequilibrium test and haplotype analysis of AGT polymorphisms. RESULTS: Linkage disequilibrium between M235T and A-20C, between M235T and A-6G, between A-20C and A-6G was observed (P<10(-4)). The case-control analysis revealed that the frequency of T235 is significantly higher in essential hypertension patients than in control subjects. But all haplotype frequencies showed no significant difference between the patient and control groups. CONCLUSION: No association was noted between the haplotypes of AGT gene and hypertension in tested people, but T235 allele might play an important role in increased risk for essential hypertension.

Alleles↗

[Maximum likelihood analysis for mapping dynamic trait QTL in outbred population. I . Methodology].

The quantitative traits whose phenotypic values change with time in life or other quantitative factors, were defined as dynamic traits. Based on the idea about random regression test-day model for estimating breeding values in animal evaluation, a mathematic model was constructed for mapping dynamic trait loci by using Legendre polynomials to model dynamic changes of each genetic effect. The Maximum likelihood analysis implemented via EM algorithm was used to estimate the parameters, including QTL position, fixed genetic regression effects for dynamic trait QTL mapping in outbred population. Compared with the existing method for dynamic trait mapping QTL, the new method presented here not only allowed to sample dynamic trait in disequillibrium way, but also can achieve to map dynamic traits QTL in any resource population by just one step. The further study on dynamic trait mapping QTL was theoretically discussed by incorporating genetic analysis of dynamic trait into general mapping QTL.

Likelihood Functions↗

[The haplotypes of three single nucleotide polymorphisms in caspase-3 gene in Han nationality of Zhejiang province in China].

OBJECTIVE: To investigate the single nucleotide polymorphisms (SNPs) and the distribution of their haplotypes in caspase-3 gene in Zhejiang Han nationality in China. METHODS: Denaturing high-performance liquid chromatography (DHPLC) and DNA sequencing were used to detect the SNPs in the regulatory region and the exons 2-7 and their flanking sequences in caspase-3 gene. Expectation maximization (EM) algorithm was used for haplotype frequencies analysis and pairwise linkage disequilibrium test. RESULTS: (1) Three SNPs were identified in caspase-3 gene; the three sites C829A, A17532C and C20541T were located in 5' regulatory region, intron 4 and 3' regulatory region, respectively. (2) Strong linkage disequilibrium was found among these SNPs; site A17532C and C20541T were in complete linkage disequilibrium. (3) C-829/A-17532/C-20541 (54.3%) was the main haplotype of Zhejiang Han nationality. CONCLUSION: The above findings indicated there is strong linkage disequilibrium among the three SNPs in caspase-3 gene in Han nationality of Zhejiang province and the main haplotype of Han nationality is obviously different from that of North American.

Asian People↗

Exact likelihood evaluation in a Markov mixture model for time series of seizure counts.

This paper provides an alternative to Albert's (1991), Biometrics 47, 1371-1381) approximation to the E-step when using the EM algorithm for parameter estimation in Markov mixture models. Use of a recursive algorithm of Baum et al. (1970, Annals of Mathematical Statistics 41, 164-171) results in exact evaluation of the likelihood, optimal parameter estimates, and very efficient computation. Applications to time series of seizure counts and fetal movements clearly show the advantages of this exact approach.

Algorithms↗