PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “EM algorithm”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 487 records · Page 27Linked to original sources

Construction of a genetic linkage map in tetraploid species using molecular markers.

This article presents methodology for the construction of a linkage map in an autotetraploid species, using either codominant or dominant molecular markers scored on two parents and their full-sib progeny. The steps of the analysis are as follows: identification of parental genotypes from the parental and offspring phenotypes; testing for independent segregation of markers; partition of markers into linkage groups using cluster analysis; maximum-likelihood estimation of the phase, recombination frequency, and LOD score for all pairs of markers in the same linkage group using the EM algorithm; ordering the markers and estimating distances between them; and reconstructing their linkage phases. The information from different marker configurations about the recombination frequency is examined and found to vary considerably, depending on the number of different alleles, the number of alleles shared by the parents, and the phase of the markers. The methods are applied to a simulated data set and to a small set of SSR and AFLP markers scored in a full-sib population of tetraploid potato.

Algorithms↗

A multivalent pairing model of linkage analysis in autotetraploids.

Polyploidy has been recognized as an important step in the evolutionary diversification of flowering plants and may have a significant impact on plant breeding. Statistical analyses for linkage mapping in polyploid species can be difficult due to considerable complexities in polysomic inheritance. In this article, we develop a novel statistical method for linkage analysis of polymorphic markers in a full-sib family of autotetraploids. This method is established on multivalent pairings of homologous chromosomes at meiosis and can provide a simultaneous maximum-likelihood estimation of the double reduction frequencies of and recombination fraction between two markers. The EM algorithm is implemented to provide a tractable way for estimating relative proportions of different modes of gamete formation that generate identical gamete genotypes due to multivalent pairings. Extensive simulation studies were performed to demonstrate the statistical properties of this method. The implications of the new method for understanding the genome structure and organization of polyploid species are discussed.

Algorithms↗

Functional mapping of quantitative trait loci underlying the character process: a theoretical framework.

Unlike a character measured at a finite set of landmark points, function-valued traits are those that change as a function of some independent and continuous variable. These traits, also called infinite-dimensional characters, can be described as the character process and include a number of biologically, economically, or biomedically important features, such as growth trajectories, allometric scalings, and norms of reaction. Here we present a new statistical infrastructure for mapping quantitative trait loci (QTL) underlying the character process. This strategy, termed functional mapping, integrates mathematical relationships of different traits or variables within the genetic mapping framework. Logistic mapping proposed in this article can be viewed as an example of functional mapping. Logistic mapping is based on a universal biological law that for each and every living organism growth over time follows an exponential growth curve (e.g., logistic or S-shaped). A maximum-likelihood approach based on a logistic-mixture model, implemented with the EM algorithm, is developed to provide the estimates of QTL positions, QTL effects, and other model parameters responsible for growth trajectories. Logistic mapping displays a tremendous potential to increase the power of QTL detection, the precision of parameter estimation, and the resolution of QTL localization due to the small number of parameters to be estimated, the pleiotropic effect of a QTL on growth, and/or residual correlations of growth at different ages. More importantly, logistic mapping allows for testing numerous biologically important hypotheses concerning the genetic basis of quantitative variation, thus gaining an insight into the critical role of development in shaping plant and animal evolution and domestication. The power of logistic mapping is demonstrated by an example of a forest tree, in which one QTL affecting stem growth processes is detected on a linkage group using our method, whereas it cannot be detected using current methods. The advantages of functional mapping are also discussed.

Chromosome Mapping↗

A general statistical framework for mapping quantitative trait loci in nonmodel systems: issue for characterizing linkage phases.

Because of uncertainty about linkage phases of founders, linkage mapping in nonmodel, outcrossing systems using molecular markers presents one of the major statistical challenges in genetic research. In this article, we devise a statistical method for mapping QTL affecting a complex trait by incorporating all possible QTL-marker linkage phases within a mapping framework. The advantage of this model is the simultaneous estimation of linkage phases and QTL location and effect parameters. These estimates are obtained through maximum-likelihood methods implemented with the EM algorithm. Extensive simulation studies are performed to investigate the statistical properties of our model. In a case study from a forest tree, this model has successfully identified a significant QTL affecting wood density. Also, the probability of the linkage phase between this QTL and its flanking markers is estimated. The implications of our model and its extension to more general circumstances are discussed.

Chromosome Mapping↗

A likelihood-based method of identifying contaminated lots of blood product.

BACKGROUND: In 1994 a small cluster of hepatitis-C cases in Rhesus-negative women in Ireland prompted a nationwide screening programme for hepatitis-C antibodies in all anti-D recipients. A total of 55 386 women presented for screening and a history of exposure to anti-D was sought from all those testing positive and a sample of those testing negative. The resulting data comprised 620 antibody-positive and 1708 antibody-negative women with known exposure history, and interest was focused on using these data to estimate the infectivity of anti-D in the period 1970-1993. METHODS: Any exposure to anti-D provides an opportunity for infection, but the infection status at each exposure time is not observed. Instead, the available data from antibody testing only indicate whether at least one of the exposures resulted in infection. Using a simple Bernoulli model to describe the risk of infection in each year, the absence of information regarding which exposure(s) led to infection fits neatly into the framework of 'incomplete data'. Hence the expectation-maximization (EM) algorithm provides estimates of the infectiousness of anti-D in each of the 24 years studied. RESULTS: The analysis highlighted the 1977 anti-D as a source of infection, a fact which was confirmed by laboratory investigation. Other suspect batches were also identified, helping to direct the efforts of laboratory investigators. CONCLUSIONS: We have presented a method to estimate the risk of infection at each exposure time from multiple exposure data. The method can also be used to estimate transmission rates and the risk associated with different sources of infection in a range of infectious disease applications.

Algorithms↗

A study on effectiveness of screening mammograms.

BACKGROUND: So far, no randomized controlled trials with a mean mammographic screening interval of > or = 2 years has demonstrated statistically significant mortality reduction for women younger than age 50. The issue of screening frequency is vital in detection of primary breast cancer. METHODS: The study group consisted of cancers diagnosed in women who participated in a serial screening programme with a mean screening interval of 2 years. To study the effectiveness of the screening, a comparison is made between the distribution of age at which the tumour could be detected when biennial mammographic screening is the only detection method, and the distribution of age at which the tumour would be detected by either biennial mammographic screening or the development of symptoms. Some recently developed statistic methods, such as bootstrap, the maximum likelihood distribution estimator for doubly censored data and the EM algorithm, are used in estimation of these distributions. RESULTS: The hypothesis tests and confidence intervals show that the difference between the two distributions was statistically significant for women younger than 50 and 50-70 years old, but not for women over 70 years. CONCLUSIONS: The statistical analysis indicates that for women younger than 50, and 50-70 years of age, a screening mammogram every other year is not frequent enough to detect primary breast cancer, but for women over 70 years, it might be sufficient.

Age Distribution↗

Monte Carlo estimation of variance component models for large complex pedigrees.

Variance component models are widely used in animal and plant breeding. In human genetics, they can be used to identify, among other traits associated with the definition of disease, those that have a significant genetic component in their aetiology. In addition, they can be used in genetic counselling. Most of the methods currently proposed for estimating variance component models often involve repeated inversion of large matrices, resulting in intensive computations, large storage requirements, and numerical instability. Consequently, these methods are restricted to data on nuclear families, to small pedigrees, or to designed pedigrees of simple form. In this paper, the authors propose a method for estimating variance component models for large complex pedigrees using jointly the EM algorithm and the Gibbs sampler. The method can handle variance component models with multiple variance components, without the need for repeated inversion of large matrices even on large complex pedigrees. The method is conceptually simple, numerically stable, and easy to implement.

Algorithms↗

Cancer risks and mortality in heterozygous ATM mutation carriers.

BACKGROUND: Homozygous or compound heterozygous mutations in the ATM gene are the principal cause of ataxia telangiectasia (A-T). Several studies have suggested that heterozygous carriers of ATM mutations are at increased risk of breast cancer and perhaps of other cancers, but the precise risk is uncertain. METHODS: Cancer incidence and mortality information for 1160 relatives of 169 UK A-T patients (including 247 obligate carriers) was obtained through the National Health Service Central Registry. Relative risks (RRs) of cancer in carriers, allowing for genotype uncertainty, were estimated with a maximum-likelihood approach that used the EM algorithm. Maximum-likelihood estimates of cancer risks associated with three groups of mutations were calculated using the pedigree analysis program MENDEL. All statistical tests were two-sided. RESULTS: The overall relative risk of breast cancer in carriers was 2.23 (95% confidence interval [CI] = 1.16 to 4.28) compared with the general population but was 4.94 (95% CI = 1.90 to 12.9) in those younger than age 50 years. The relative risk for all cancers other than breast cancer was 2.05 (95% CI = 1.09 to 3.84) in female carriers and 1.23 (95% CI = 0.76 to 2.00) in male carriers. Breast cancer was the only site for which a clear risk increase was seen, although there was some evidence of excess risks of colorectal cancer (RR = 2.54, 95% CI = 1.06 to 6.09) and stomach cancer (RR = 3.39, 95% CI = 0.86 to 13.4). Carriers of mutations predicted to encode a full-length ATM protein had cancer risks similar to those of people carrying truncating mutations. CONCLUSION: These results confirm a moderate risk of breast cancer in A-T heterozygotes and give some evidence of an excess risk of other cancers but provide no support for large mutation-specific differences in risk.

Adult↗

Haplotypic association of DDAH1 with susceptibility to pre-eclampsia.

Association between pre-eclampsia (PEE1) and the dimethylarginine dimethylaminohydrolase (DDAH) 1 and 2 genes, which play a role in the regulation of nitric oxide synthesis and release, was studied. In a case-control study design single nucleotide polymorphisms (SNPs) were determined at eight sites in the DDAH1 gene and at one site (Pro231Pro) in the DDAH2 gene from 132 women with pre-eclampsia and 112 healthy controls. Three SNPs in the DDAH1 gene were associated with pre-eclampsia, showing complete linkage disequilibrium with each other, but none of the associations in the allele or genotype data reached statistical significance in either of the genes after the correction for multiple testing. Haplotype frequencies were estimated using a population based on a maximum likelihood method (EM algorithm). Four common DDAH1 haplotypes were present and a significant association of haplotypes H2 and H3 with pre-eclampsia (P=0.03) was found. The risk of pre-eclampsia was greatest in individuals (odds ratio: 3.93; 95% confidence interval: 1.54-9.99) who had two copies of the high-risk haplotypes (H2 or H3). The observed haplotypic association provides the first evidence of the importance of DDAH1 polymorphisms in pre-eclampsia susceptibility.

Adult↗

Logistic regression when the outcome is measured with uncertainty.

In epidemiologic research, logistic regression is often used to estimate the odds of some outcome of interest as a function of predictors. However, in some datasets, the outcome of interest is measured with imperfect sensitivity and specificity. It is well known that the misclassification induced by such an imperfect diagnostic test will lead to biased estimates of the odds ratios and their variances. In this paper, the authors show that when the sensitivity and specificity of a diagnostic test are known, it is straightforward to incorporate this information into the fitting of logistic regression models. An EM algorithm that produces unbiased estimates of the odds ratios and their variances is described. The resulting odds ratio estimates tend to be farther from the null but have greater variance than estimates found by ignoring the imperfections of the test. The method can be extended to the situation where the sensitivity and specificity differ for different study subjects, i.e., nondifferential misclassification. The method is useful even when the sensitivity and specificity are not known, as a way to see the degree to which various assumptions about sensitivity and specificity affect one's estimates. The method can also be used to estimate sensitivity and specificity under certain assumptions or when a validation subsample is available. Several examples are provided to compare the results of this method with those obtained by standard logistic regression. A SAS macro that implements the method is available on the World Wide Web at http:@som1.ab.umd.edu/Epidemiology/software.h tml.

Adult↗

Use of two data sources to estimate odds ratios in case-control studies.

Information bias is among the most serious and common problems in epidemiology. Approaches have been developed to reduce information bias by correcting for known amounts of misclassification. Unfortunately, in most studies, the extent of exposure misclassification cannot be easily estimated. We discuss the application to case-control studies of an approach originally proposed by Hui and Walter in 1980 to estimate the sensitivity and specificity of two independent classification schemes (Hui SL, Walter SD. Biometrics 1980;36:167-171). In this paper, we propose using the EM algorithm to provide a simple numeric technique for implementing their method that seems to converge for most real-world data. Our approach allows inclusion of a measure of non-independence of the two classification schemes, and we assess the influence of non-independence on the odds ratio. Finally, we provide a simple variance estimate for the odds ratio based on the delta method and maximum likelihood theory. We exemplify our results and method with data from a case-control study of sudden infant death syndrome in which data on some variables were obtained from both maternal interviews and medical records.

Algorithms↗

Agricultural risk factors for t(14;18) subtypes of non-Hodgkin's lymphoma.

The t(14;18) translocation is a common somatic mutation in non-Hodgkin's lymphoma (NHL) that is associated with bcl-2 activation and inhibition of apoptosis. We hypothesized that some risk factors might act specifically along t(14;18)-dependent pathways, leading to stronger associations with t(14;18)-positive than t(14;18)-negative non-Hodgkin's lymphoma. Archival biopsies from 182 non-Hodgkin's lymphoma cases included in a case-control study of men in Iowa and Minnesota (the Factors Affecting Rural Men, or FARM study) were assayed for t(14;18) using polymerase chain reaction amplification; 68 (37%) were t(14;18)-positive. We estimated adjusted odds ratios (OR) and 95% confidence intervals (CI) for various agricultural risk factors and t(14;18)-positive and -negative cases of non-Hodgkin's lymphoma, based on polytomous logistic regression models fit using the expectation-maximization (EM) algorithm. T(14;18)-positive non-Hodgkin's lymphoma was associated with farming (OR 1.4, 95% CI = 0.9-2.3), dieldrin (OR 3.7, 95% CI = 1.9-7.0), toxaphene (OR 3.0, 95% CI = 1.5-6.1), lindane (OR 2.3, 95% CI = 1.3-3.9), atrazine (OR 1.7, 95% CI = 1.0-2.8), and fungicides (OR 1.8, 95% CI = 0.9-3.6), in marked contrast to null or negative associations for the same self-reported exposures and t(14;18)-negative non-Hodgkin's lymphoma. Causal relations between agricultural exposures and t(14;18)-positive non-Hodgkin's lymphoma are plausible, but associations should be confirmed in a larger study. Results suggest that non-Hodgkin's lymphoma classification based on the t(14;18) translocation is of value in etiologic research.

Adult↗

Diurnal changes in the pharmacokinetic behavior of amikacin.

This retrospective study evaluated possible differences in the pharmacokinetic behavior of amikacin between the morning (AM) and evening (PM). Of 634 patients receiving amikacin therapy, 17 received a dose every 12 hours (an i.v. infusion at 8:00 AM and 8:00 PM) with amikacin serum levels obtained after both the AM and PM infusions. Pharmacokinetic parameter values were estimated by the nonparametric EM algorithm (USC*PACK clinical software) for a one-compartment model. All patient data were analyzed in three ways. The parameter values were estimated by fitting the model first only to the serum levels drawn following the AM dose; second, only to the data following the PM dose; and third, to all serum levels (AM + PM). Parameter values found were (mean, median, SD respectively): AM: Kel = 0.181114 h(-1), 0.224460 h(-1), 0.058820 h(-1); Vol = 23.657507 L; 23.376231 L; 1.353253 L; Cl = 4.326720 L x h(-1), 5.303726 L x h(-1), 1.447731 L x h(-1); PM: Kel = 0.110151 h(-1); 0.121295 h(-1); 0.016860 h(-1); Vol = 28.948043 L; 24.091703 L; 9.266628 L; Cl = 3.081761 L x h(-1), 2.810615 L x h(-1); 0.705874 L x h(-1); AM + PM: Kel = 0.165321 h(-1); 0.131796 h(-1); 0.075425 h(-1); Vol = 25.479043 L; 26.187970 L; 5.367054 L. These findings are in agreement with the known diurnal rhythm of glomerular filtration rate. Because pharmacokinetic parameter values are most often estimated using AM data, this may lead to an overevaluation of these values compared with PM or to values for the entire day. The resulting drug regimens may therefore be overestimated regarding the elimination rate constant and underestimated regarding the volume of distribution.

Adult↗

Risk of small-for-gestational age is associated with common anti-inflammatory cytokine polymorphisms.

BACKGROUND: Anti-inflammatory cytokines play a key role in pregnancy maintenance. Genetic variation in anti-inflammatory cytokines could influence a woman's risk of adverse reproductive outcomes. METHODS: We investigated the relationship of polymorphisms in interleukin 4 (IL4), IL5, IL10, IL13, and transforming growth factor (TGFbeta1) with spontaneous preterm birth and small-for-gestational age (SGA) in a nested case-control study of a prospective pregnancy cohort. Women were recruited between 24 and 29 weeks' gestation at the Wake County and University of North Carolina, Chapel Hill obstetric clinics between February 1996 and June 2000. We inferred haplotypes using the EM algorithm and the Bayesian method, PHASE. Semi-Bayesian hierarchical logistic regression was used to obtain odds ratio (OR) estimates and 95% confidence intervals (CIs) for each polymorphism. RESULTS: African-American mothers who carried the IL4 GCC haplotype had greater risk of spontaneous preterm birth (OR = 2.9; 95% CI = 1.2-7.4). In white mothers, carriers of the "low-producing" IL4 CC and IL10 ATA haplotypes had markedly reduced risk of SGA (for the CC haplotype, 0.2 [0.0-1.2]; for the ATA haplotype, 0.5 [0.3-0.8]), whereas carriers of the "high-producing" IL4(-589)T variant had increased risk of SGA in both African-American and white mothers. CONCLUSIONS: Variants related to decreased anti-inflammatory cytokine production may lower risk of SGA. Furthermore, the same mechanism that protects against SGA might increase risk of spontaneous preterm birth.

Black or African American↗

Risk of spontaneous preterm birth is associated with common proinflammatory cytokine polymorphisms.

BACKGROUND: Preliminary data suggest that common genetic variation in immune response genes can contribute to the risk for spontaneous preterm birth and possibly small-for-gestational age (SGA). METHODS: We investigated the relationship of polymorphisms in 6 cytokine genes associated with inflammation-interleukin (IL)1alpha, IL1beta, IL2, IL6, tumor necrosis factor (TNF), and lymphotoxin alpha (LTA)-with spontaneous preterm and SGA birth in a nested case-control study drawn from a prospective pregnancy cohort. Women were recruited between 24 and 29 weeks' gestation at the Wake County and University of North Carolina, Chapel Hill obstetric clinics between February 1996 and June 2000. We inferred haplotypes using the EM algorithm and the Bayesian method, PHASE. We then compared haplotype frequency distributions and implemented semi-Bayesian hierarchical logistic regression analyses to obtain odds ratio (OR) estimates and 95% confidence intervals (CIs) for each polymorphism. RESULTS: Two haplotypes spanning the TNF/LTA genes were associated with increased risk for spontaneous preterm birth in white subjects (for the AGG haplotype, OR = 1.5 [95% CI=0.8-2.6]; for the GAC haplotype, 1.6 [0.9-2.9]). Additionally, carriers of the GAG haplotype were found to have decreased risk of spontaneous preterm birth (0.6; 0.3-1.0). The TNF(-488)A and LTA(IVS1-82)C variants, constituents of the AGG and GAC haplotypes respectively, were also strongly associated with increased risk of spontaneous preterm birth. CONCLUSIONS: Our results suggest that common genetic variants in proinflammatory cytokine genes could influence the risk for spontaneous preterm birth. Selected TNF/LTA haplotypes were associated with spontaneous preterm birth in both African-American and white subjects. Our data do not support an inflammatory etiology for SGA.

Black or African American↗

Positive associations of polymorphisms in the metabotropic glutamate receptor type 3 gene (GRM3) with schizophrenia.

OBJECTIVES: Glutamatergic dysfunction is one of the major hypotheses of schizophrenia pathophysiology. We have been conducting systematic studies on the association between glutamate receptors and schizophrenia. We focused on the metabotropic glutamate receptor type 3 gene (GRM3) as a candidate for schizophrenia susceptibility. METHODS: We genotyped Japanese schizophrenics (n=100) and controls (n=100) for six single nucleotide polymorphisms (SNPs) located in the GRM3 region at intervals of approximately 50 kb. Statistical differences in genotype, allele and haplotype frequencies between cases and controls were evaluated by the chi2 test and Fisher's exact probability test at a significance level of 0.05. Haplotype frequencies were estimated by the EM algorithm. RESULTS: A case-control association study identified a significant difference in allele frequency distribution of a SNP, rs1468412, between schizophrenics and controls (P=0.011). We also observed significant differences in haplotype frequencies estimated from SNP frequencies between schizophrenics and controls. The haplotype constructed from three SNPs, including rs1468412, showed a significant association with schizophrenia (P=8.30 x 10-4). CONCLUSIONS: Our data indicate that at least one susceptibility locus for schizophrenia is situated within or very close to the GRM3 region in the Japanese patients.

Base Sequence↗

Bilinear dynamical systems.

In this paper, we propose the use of bilinear dynamical systems (BDS)s for model-based deconvolution of fMRI time-series. The importance of this work lies in being able to deconvolve haemodynamic time-series, in an informed way, to disclose the underlying neuronal activity. Being able to estimate neuronal responses in a particular brain region is fundamental for many models of functional integration and connectivity in the brain. BDSs comprise a stochastic bilinear neurodynamical model specified in discrete time, and a set of linear convolution kernels for the haemodynamics. We derive an expectation-maximization (EM) algorithm for parameter estimation, in which fMRI time-series are deconvolved in an E-step and model parameters are updated in an M-Step. We report preliminary results that focus on the assumed stochastic nature of the neurodynamic model and compare the method to Wiener deconvolution.

Algorithms↗

Haplotype and linkage disequilibrium architecture for human cancer-associated genes.

To facilitate association-based linkage studies we have studied the linkage disequilibrium (LD) and haplotype architecture around five genes of interest for cancer risk: ATM, BRCA1, BRCA2, RAD51, and TP53. Single nucleotide polymorphisms (SNPs) were identified and used to construct haplotypes that span 93-200 kb per locus with an average SNP density of 12 kb. These markers were genotyped in four ethnically defined populations that contained 48 each of African Americans, Asian Americans, Hispanic Americans, and European Americans. Haplotypes were inferred using an expectation maximization (EM) algorithm, and the data were analyzed using D', R(2), Fisher's exact P-values, and the four-gamete test for recombination. LD levels varied widely between loci from continuously high LD across 200 kb to a virtual absence of LD across a similar length of genome. LD structure also varied at each gene and between populations studied. This variation indicates that the success of linkage-based studies will require a precise description of LD at each locus and in each population to be studied. One striking consistency between genes was that at each locus a modest number of haplotypes present in each population accounted for a high fraction of the total number of chromosomes. We conclude that each locus has its own genomic profile with regard to LD, and despite this there is the widespread trend of relatively low haplotype diversity. As a result, a low marker density should be adequate to identify haplotypes that represent the common variation at a locus, thereby decreasing costs and increasing efficacy of association studies.

Alleles↗