PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “likelihood”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Likelihood-based inference for the shape parameter of a two parameter Weibull distribution.

Based on a Type 2 censored sample, we use the likelihood-based approach to draw likelihood inference on the shape parameter gamma of a two-parameter Weibull distribution. In particular, we derive the profile, conditional and marginal likelihoods of gamma. Numerical results along with some concluding remarks regarding the use of likelihood-based methods for inference are provided.

Humans↗

On the use of the moments method of estimation to obtain approximate maximum likelihood estimates of linkage between a genetic marker and a quantitative locus.

The approximate maximum likelihood method of Luo & Kearsey (1989) to determine the parameters of a segregating quantitative trait locus (QTL) in an F2 population by linkage to a genetic marker is compared to the 'true' maximum likelihood estimates derived by scanning the seven parameter likelihood space. The two methods did not yield identical results, and diverged markedly for a dominant QTL loosely linked to the genetic marker. Further study is suggested to evaluate other methods that do not simultaneously maximize the likelihood for all parameters.

Genetic Linkage↗

A marginal likelihood model for family-based data.

This paper presents a marginal likelihood model for family-based data based upon the transmission of marker alleles from each heterozygous parent to his/her affected children. The proposed model, extending the maximum-likelihood-binomial (MLB) method and the disequilibrium maximum-likelihood-binomial (DMLB) method (Abel et al. 1998; Abel & Müller-Myhsok, 1998; Huang & Jiang, 1999), is adaptive to linkage disequilibrium (LD) and linkage heterogeneity. Compared with other procedures, the likelihood ratio test (LRT) derived from the proposed model enjoys superior qualities. First, simulations indicate that the power of the LRT is greater than that of the TDT or DMLB in all of our studied scenarios. Second, when we applied the LRT and other tests to a Tourette Syndrome data, the result was data favorable to the use of the LRT. Therefore, we recommend the use of the LRT as an additional linkage test wherever applicable, especially when the amount of LD is uncertain.

Alleles↗

Notes on the maximum likelihood estimation of haplotype frequencies.

The maximum likelihood estimation (MLE) is one of the most popular ways to estimate haplotype frequencies of a population with genotype data whose linkage phases are unknown. The MLE is commonly implemented in the use of the Expectation-Maximization (EM) algorithm. It is known that the EM algorithm carries the risk that an estimator may converge erroneously to one of the local maxima or saddle points of the likelihood surface, resulting in serious errors in the MLE of haplotype frequencies. In this note, by theoretical treatments we present the necessary and sufficient conditions that the local maxima or saddle points on the likelihood surface appear. As a rule of thumb, that the difference between the coupling and repulsive haplotype frequencies in phase known individuals is 3/2 times larger than the frequency of phase ambiguous individuals is the sufficient condition that the likelihood surface is unimodal. Moreover, we present the analytic solution to the biallelic two-locus problem, and construct a general algorithm to obtain the global maximum.

Algorithms↗

Markov chain Monte Carlo without likelihoods.

Many stochastic simulation approaches for generating observations from a posterior distribution depend on knowing a likelihood function. However, for many complex probability models, such likelihoods are either impossible or computationally prohibitive to obtain. Here we present a Markov chain Monte Carlo method for generating observations from a posterior distribution without the use of likelihoods. It can also be used in frequentist applications, in particular for maximum-likelihood estimation. The approach is illustrated by an example of ancestral inference in population genetics. A number of open problems are highlighted in the discussion.

Algorithms↗

Approximate likelihood-ratio test for branches: A fast, accurate, and powerful alternative.

We revisit statistical tests for branches of evolutionary trees reconstructed upon molecular data. A new, fast, approximate likelihood-ratio test (aLRT) for branches is presented here as a competitive alternative to nonparametric bootstrap and Bayesian estimation of branch support. The aLRT is based on the idea of the conventional LRT, with the null hypothesis corresponding to the assumption that the inferred branch has length 0. We show that the LRT statistic is asymptotically distributed as a maximum of three random variables drawn from the chi(0)2 + chi(1)2 distribution. The new aLRT of interior branch uses this distribution for significance testing, but the test statistic is approximated in a slightly conservative but practical way as 2(l1- l2), i.e., double the difference between the maximum log-likelihood values corresponding to the best tree and the second best topological arrangement around the branch of interest. Such a test is fast because the log-likelihood value l2 is computed by optimizing only over the branch of interest and the four adjacent branches, whereas other parameters are fixed at their optimal values corresponding to the best ML tree. The performance of the new test was studied on simulated 4-, 12-, and 100-taxon data sets with sequences of different lengths. The aLRT is shown to be accurate, powerful, and robust to certain violations of model assumptions. The aLRT is implemented within the algorithm used by the recent fast maximum likelihood tree estimation program PHYML (Guindon and Gascuel, 2003).

Classification↗

Phylogenetic analysis and intraspecific variation: performance of parsimony, likelihood, and distance methods.

Intraspecific variation is abundant in all types of systematic characters but is rarely addressed in simulation studies of phylogenetic method performance. We compared the accuracy of 15 phylogenetic methods using simulations to (1) determine the most accurate method(s) for analyzing polymorphic data (under simplified conditions) and (2) test if generalizations about the performance of phylogenetic methods based on previous simulations of fixed (nonpolymorphic) characters are robust to a very different evolutionary model that explicitly includes intraspecific variation. Simulated data sets consisted of allele frequencies that evolved by genetic drift. The phylogenetic methods included eight parsimony coding methods, continuous maximum likelihood, and three distance methods (UPGMA, neighbor joining, and Fitch-Margoliash) applied to two genetic distance measures (Nei's and the modified Cavalli-Sforza and Edwards chord distance). Two sets of simulations were performed. The first examined the effects of different branch lengths, sample sizes (individuals sampled per species), numbers of characters, and numbers of alleles per locus in the eight-taxon case. The second examined more extensively the effects of branch length in the four-taxon, two-allele case. Overall, the most accurate methods were likelihood, the additive distance methods (neighbor joining and Fitch-Margoliash), and the frequency parsimony method. Despite the use of a very different evolutionary model in the present article, many of the results are similar to those from simulations of fixed characters. Similarities include the presence of the "Felsenstein zone," where methods often fail, which suggests that long-branch attraction may occur among closely related species through genetic drift. Differences between the results of fixed and polymorphic data simulations include the following: (1) UPGMA is as accurate or more accurate than nonfrequency parsimony methods across nearly all combinations of branch lengths, and (2) likelihood and the additive distance methods are not positively misled under any combination of branch lengths tested (even when the assumptions of the methods are violated and few characters are sampled). We found that sample size is an important determinant of accuracy and affects the relative success of methods (i.e., distance and likelihood methods outperform parsimony at small sample sizes). Attempts to generalize about the behavior of phylogenetic methods should consider the extreme examples offered by fixed-mutation models of DNA sequence data and genetic-drift models of allele frequencies.

Computer Simulation↗

A fast method for approximating maximum likelihoods of phylogenetic trees from nucleotide sequences.

We have developed a rapid parsimony method for reconstructing ancestral nucleotide states that allows calculation of initial branch lengths that are good approximations to optimal maximum-likelihood estimates under several commonly used substitution models. Use of these approximate branch lengths (rather than fixed arbitrary values) as starting points significantly reduces the time required for iteration to a solution that maximizes the likelihood of a tree. These branch lengths are close enough to the optimal values that they can be used without further iteration to calculate approximate maximum-likelihood scores that are very close to the "exact" scores found by iteration. Several strategies are described for using these approximate scores to substantially reduce times needed for maximum-likelihood tree searches.

Animals↗

Estimating the degree of heterogeneity between event rates using likelihood.

The observed variability between mortality or morbidity rates in epidemiologic studies is partly due to random fluctuations. The same is true for rate ratios relative to reference rates. A method for estimating the distribution of true rate ratios is applied to a data set of perinatal mortality in 515 small areas in the North West Thames Health Region, England, in the period 1986-1990. Combining the random Poisson variability with the assumption that the true rate ratios are drawn from a gamma distribution (a family of positive unimodal distributions) produces a negative binomial log-likelihood for the dispersion parameter of the gamma. The maximum likelihood estimate of this parameter and its confidence interval are then found via direct numerical methods; alternatively, the hypothesis of no heterogeneity is tested by a likelihood ratio. The standardized mortality ratios (SMRs) for the data have an empirical distribution with 5th percentile at 0 and 95th percentile at 1.92, but their true variability, as described by the 5th to 95th percentiles of the fitted gamma distribution, is from 0.72 to 1.32. The likelihood ratio test confirmed this result, rejecting the hypothesis that the true rates are homogeneous (p = 0.015). The method requires only modest computing resources and is useful when assessing the need for more detailed study.

England↗

RAxML-III: a fast program for maximum likelihood-based inference of large phylogenetic trees.

MOTIVATION: The computation of large phylogenetic trees with statistical models such as maximum likelihood or bayesian inference is computationally extremely intensive. It has repeatedly been demonstrated that these models are able to recover the true tree or a tree which is topologically closer to the true tree more frequently than less elaborate methods such as parsimony or neighbor joining. Due to the combinatorial and computational complexity the size of trees which can be computed on a Biologist's PC workstation within reasonable time is limited to trees containing approximately 100 taxa. RESULTS: In this paper we present the latest release of our program RAxML-III for rapid maximum likelihood-based inference of large evolutionary trees which allows for computation of 1.000-taxon trees in less than 24 hours on a single PC processor. We compare RAxML-III to the currently fastest implementations for maximum likelihood and bayesian inference: PHYML and MrBayes. Whereas RAxML-III performs worse than PHYML and MrBayes on synthetic data it clearly outperforms both programs on all real data alignments used in terms of speed and final likelihood values. Availability SUPPLEMENTARY INFORMATION: RAxML-III including all alignments and final trees mentioned in this paper is freely available as open source code at http://wwwbode.cs.tum/~stamatak CONTACT: stamatak@cs.tum.edu.

Algorithms↗

Maximum likelihood estimation of ordered multinomial parameters.

The pool adjacent violator algorithm Ayer et al. (1955, The Annals of Mathematical Statistics, 26, 641-647) has long been known to give the maximum likelihood estimator of a series of ordered binomial parameters, based on an independent observation from each distribution (see Barlow et al., 1972, Statistical Inference under Order Restrictions, Wiley, New York). This result has immediate application to estimation of a survival distribution based on current survival status at a set of monitoring times. This paper considers an extended problem of maximum likelihood estimation of a series of 'ordered' multinomial parameters p(i)= (p(1i),p(2i),.,p(mi)) for 1 <or=i <ro=k, where ordered means that p(j1) <or=p(j2) <or=<or=p(jk) for each j with 1 <or=j <or=m-1. The data consist of k independent observations X(1),., X(k) where X(i) has a multinomial distribution with probability parameter p(i) and known index n(i)\geq 1. By making use of variants of the pool adjacent violator algorithm, we obtain a simple algorithm to compute the maximum likelihood estimator of p(1),., p(k), and demonstrate its convergence. The results are applied to nonparametric maximum likelihood estimation of the sub-distribution functions associated with a survival time random variable with competing risks when only current status data are available (Jewell et al. 2003, Biometrika, 90, 183-197).

Age Factors↗

A likelihood approach to populations samples of microsatellite alleles.

This paper presents a likelihood approach to population samples of microsatellite alleles. A Markov chain recursion method previously published by GRIFFITHS and TAVARE is applied to estimate the likelihood function under different models of microsatellite evolution. The method presented can be applied to estimate a fundamental population genetics parameter theta as well as parameters of the mutational model. The new likelihood estimator provides a better estimator of theta in terms of the mean square error than previous approaches. Furthermore, it is demonstrated how the method may easily be applied to test models of microsatellite evolution. In particular it is shown how to compare a one-step model of microsatellite evolution to a multi-step model by a likelihood ratio test.

Alleles↗

Inference of population history using a likelihood approach.

We introduce an approach to revealing the likelihood of different population histories that utilizes an explicit model of sequence evolution for the DNA segment under study. Based on a phylogenetic tree reconstruction method we show that a Tamura-Nei model with heterogeneous mutation rates is a fair description of the evolutionary process of the hypervariable region I of the mitochondrial DNA from humans. Assuming this complex model still allows the estimation of population history parameters, we suggest a likelihood approach to conducting statistical inference within a class of expansion models. More precisely, the likelihood of the data is based on the mean pairwise differences between DNA sequences and the number of variable sites in a sample. The use of likelihood ratios enables comparison of different hypotheses about population history, such as constant population size during the past or an increase or decrease of population size starting at some point back in time. This method was applied to show that the population of the Basques has expanded, whereas that of the Biaka pygmies is most likely decreasing. The Nuu-Chah-Nulth data are consistent with a model of constant population.

Base Composition↗

Maximum-likelihood estimation of admixture proportions from genetic data.

For an admixed population, an important question is how much genetic contribution comes from each parental population. Several methods have been developed to estimate such admixture proportions, using data on genetic markers sampled from parental and admixed populations. In this study, I propose a likelihood method to estimate jointly the admixture proportions, the genetic drift that occurred to the admixed population and each parental population during the period between the hybridization and sampling events, and the genetic drift in each ancestral population within the interval between their split and hybridization. The results from extensive simulations using various combinations of relevant parameter values show that in general much more accurate and precise estimates of admixture proportions are obtained from the likelihood method than from previous methods. The likelihood method also yields reasonable estimates of genetic drift that occurred to each population, which translate into relative effective sizes (N(e)) or absolute average N(e)'s if the times when the relevant events (such as population split, admixture, and sampling) occurred are known. The proposed likelihood method also has features such as relatively low computational requirement compared with previous ones, flexibility for admixture models, and marker types. In particular, it allows for missing data from a contributing parental population. The method is applied to a human data set and a wolflike canids data set, and the results obtained are discussed in comparison with those from other estimators and from previous studies.

Animals↗

Effect of recombination on the accuracy of the likelihood method for detecting positive selection at amino acid sites.

Maximum-likelihood methods based on models of codon substitution accounting for heterogeneous selective pressures across sites have proved to be powerful in detecting positive selection in protein-coding DNA sequences. Those methods are phylogeny based and do not account for the effects of recombination. When recombination occurs, such as in population data, no unique tree topology can describe the evolutionary history of the whole sequence. This violation of assumptions raises serious concerns about the likelihood method for detecting positive selection. Here we use computer simulation to evaluate the reliability of the likelihood-ratio test (LRT) for positive selection in the presence of recombination. We examine three tests based on different models of variable selective pressures among sites. Sequences are simulated using a coalescent model with recombination and analyzed using codon-based likelihood models ignoring recombination. We find that the LRT is robust to low levels of recombination (with fewer than three recombination events in the history of a sample of 10 sequences). However, at higher levels of recombination, the type I error rate can be as high as 90%, especially when the null model in the LRT is unrealistic, and the test often mistakes recombination as evidence for positive selection. The test that compares the more realistic models M7 (beta) against M8 (beta and omega) is more robust to recombination, where the null model M7 allows the positive selection pressure to vary between 0 and 1 (and so does not account for positive selection), and the alternative model M8 allows an additional discrete class with omega = d(N)/d(S) that could be estimated to be >1 (and thus accounts for positive selection). Identification of sites under positive selection by the empirical Bayes method appears to be less affected than the LRT by recombination.

Bayes Theorem↗

Likelihood, parsimony, and heterogeneous evolution.

Evolutionary rates vary among sites and across the phylogenetic tree (heterotachy). A recent analysis suggested that parsimony can be better than standard likelihood at recovering the true tree given heterotachy. The authors recommended that results from parsimony, which they consider to be nonparametric, be reported alongside likelihood results. They also proposed a mixture model, which was inconsistent but better than either parsimony or standard likelihood under heterotachy. We show that their main conclusion is limited to a special case for the type of model they study. Their mixture model was inconsistent because it was incorrectly implemented. A useful nonparametric model should perform well over a wide range of possible evolutionary models, but parsimony does not have this property. Likelihood-based methods are therefore the best way to deal with heterotachy.

Animals↗

Maximum likelihood Jukes-Cantor triplets: analytic solutions.

Maximum likelihood (ML) is a popular method for inferring a phylogenetic tree of the evolutionary relationship of a set of taxa, from observed homologous aligned genetic sequences of the taxa. Generally, the computation of the ML tree is based on numerical methods, which in a few cases, are known to converge to a local maximum on a tree, which is suboptimal. The extent of this problem is unknown, one approach is to attempt to derive algebraic equations for the likelihood equation and find the maximum points analytically. This approach has so far only been successful in the very simplest cases, of three or four taxa under the Neyman model of evolution of two-state characters. In this paper we extend this approach, for the first time, to four-state characters, the Jukes-Cantor model under a molecular clock, on a tree T on three taxa, a rooted triple. We employ spectral methods (Hadamard conjugation) to express the likelihood function parameterized by the path-length spectrum. Taking partial derivatives, we derive a set of polynomial equations whose simultaneous solution contains all critical points of the likelihood function. Using tools of algebraic geometry (the resultant of two polynomials) in the computer algebra packages (Maple), we are able to find all turning points analytically. We then employ this method on real sequence data and obtain realistic results on the primate-rodents divergence time.

Biological Evolution↗

Reliabilities of parsimony-based and likelihood-based methods for detecting positive selection at single amino acid sites.

The reliabilities of parsimony-based and likelihood-based methods for inferring positive selection at single amino acid sites were studied using the nucleotide sequences of human leukocyte antigen (HLA) genes, in which positive selection is known to be operating at the antigen recognition site. The results indicate that the inference by parsimony-based methods is robust to the use of different evolutionary models and generally more reliable than that by likelihood-based methods. In contrast, the results obtained by likelihood-based methods depend on the models and on the initial parameter values used. It is sometimes difficult to obtain the maximum likelihood estimates of parameters for a given model, and the results obtained may be false negatives or false positives depending on the initial parameter values. It is therefore preferable to use parsimony-based methods as long as the number of sequences is relatively large and the branch lengths of the phylogenetic tree are relatively small.

Amino Acid Substitution↗