PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “likelihood”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Pushing the boundaries of molecular replacement with maximum likelihood.

The molecular-replacement method works well with good models and simple unit cells, but often fails with more difficult problems. Experience with likelihood in other areas of crystallography suggests that it would improve performance significantly. For molecular replacement, the form of the required likelihood function depends on whether there is ambiguity in the relative phases of the contributions from symmetry-related molecules (e.g. rotation versus translation searches). Likelihood functions used in structure refinement are appropriate only for translation (or six-dimensional) searches, where the correct translation will place all of the atoms in the model approximately correctly. A new likelihood function that allows for unknown relative phases is suitable for rotation searches. It is shown that correlations between sequence identity and coordinate error can be used to calibrate parameters for model quality in the likelihood functions. Multiple models of a molecule can be combined in a statistically valid way by setting up the joint probability distribution of the true and model structure factors as a multivariate complex normal distribution, from which the conditional distribution of the true structure factor given the models can be derived. Tests in a new molecular-replacement program, Beast, show that the likelihood-based targets are more sensitive and more accurate than previous targets. The new multiple-model likelihood function has a dramatic impact on success.

Calibration↗

Evaluation of likelihood ratios for complex genetic models.

Although methods for computing likelihoods for simple genetic models on large and complex pedigrees have been known for some time, and although methods for evaluating likelihoods for complex genetic models on small pedigrees have likewise been well known, likelihood evaluation for complex models given data on extended pedigrees has remained an intractable problem. The Gibbs sampler provides a method of Monte Carlo evaluation of likelihood ratios for complex models on extended and/or complex pedigrees. With increasing computer speeds, this approach provides a tractable and efficient approach to many such likelihood evaluation problems in linkage and segregation analysis. In this paper, however, the authors restrict attention to two basic building-blocks of the overall process. The first is the sequential computation of Gaussian likelihoods for multiple random-effects models on extended pedigrees. The second is the use of this in the Monte Carlo evaluation of likelihoods for the classical mixed model of segregation analysis. The implementation of the Gibbs sampler on pedigrees that permits this Monte Carlo evaluation is detailed. An example is then presented, and finally, in the context of this same example, it is also shown how linkage analysis for a quantitative trait falls within this same framework.

Female↗

The robustness of two phylogenetic methods: four-taxon simulations reveal a slight superiority of maximum likelihood over neighbor joining.

The robustness (sensitivity to violation of assumptions) of the maximum-likelihood and neighbor-joining methods was examined using simulation. Maximum likelihood and neighbor joining were implemented with Jukes-Cantor, Kimura, and gamma models of DNA substitution. Simulations were performed in which the assumptions of the methods were violated to varying degrees on three model four-taxon trees. The performance of the methods was evaluated with respect to ability to correctly estimate the unrooted four-taxon tree. Maximum likelihood outperformed neighbor joining in 29 of the 36 cases in which the assumptions of both methods were satisfied. In 133 of 180 of the simulations in which the assumptions of the maximum-likelihood and neighbor-joining methods were violated, maximum likelihood outperformed neighbor joining. These results are consistent with a general superiority of maximum likelihood over neighbor joining under comparable conditions. They extend and clarify an earlier study that found an advantage for neighbor joining over maximum likelihood for gamma-distributed mutation rates.

DNA↗

A note on the asymptotic distribution of likelihood ratio tests to test variance components.

When using maximum likelihood methods to estimate genetic and environmental components of (co)variance, it is common to test hypotheses using likelihood ratio tests, since such tests have desirable asymptotic properties. In particular, the standard likelihood ratio test statistic is assumed asymptotically to follow a chi2 distribution with degrees of freedom equal to the number of parameters tested. Using the relationship between least squares and maximum likelihood estimators for balanced designs, it is shown why the asymptotic distribution of the likelihood ratio test for variance components does not follow a chi2 distribution with degrees of freedom equal to the number of parameters tested when the null hypothesis is true. Instead, the distribution of the likelihood ratio test is a mixture of chi2 distributions with different degrees of freedom. Implications for testing variance components in twin designs and for quantitative trait loci mapping are discussed. The appropriate distribution of the likelihood ratio test statistic should be used in hypothesis testing and model selection.

Chromosome Mapping↗

Factors associated with the likelihood of Giardia spp. and Cryptosporidium spp. in soil from dairy farms.

A study was conducted to identify factors associated with the likelihood of detecting Giardia spp. and Cryptosporidium spp. in the soil of dairy farms in a watershed area. A total of 37 farms were visited, and 782 soil samples were collected from targeted areas on these farms. The samples were analyzed for the presence of Cryptosporidium spp. oocysts, Giardia spp. cysts, percent moisture content, and pH. Logistic regression analysis was used to identify risk factors associated with the likelihood of the presence of these organisms. The use of the land at the sampling site was associated with the likelihood of environmental contamination with Cryptosporidium spp. Barn cleaner equipment area and agricultural fields were associated with increased likelihood of environmental contamination with Cryptosporidium spp. The risk of environmental contamination decreased with the pH of the soil and with the score of the potential likelihood of Cryptosporidium spp. The size of the sampling site, as determined by the sampling design, in square feet, was associated nonlinearly with the risk of detecting Cryptosporidium spp. The likelihood of the Giardia cyst in the soil increased with the prevalence of Giardia spp. in animals (i.e., 18 to 39%). As the size of the farm increased, there was decreased risk of Giardia spp. in the soil, and sampling sites which were covered with brush or bare soil showed a decrease in likelihood of detecting Giardia spp. when compared to land which had managed grass. The number of cattle on the farm less than 6 mo of age was negatively associated with the risk of detecting Giardia spp. in the soil, and the percent moisture content was positively associated with the risk of detecting Giardia spp. Our study showed that these two protozoan exist in dairy farm soil at different rates, and this risk could be modified by manipulating the pH of the soil.

Agriculture↗

A perspective on standardizing the predictive power of noninvasive cardiovascular tests by likelihood ratio computation: 2. Clinical applications.

Likelihood ratio measures may be used as a standard for expressing the predictive power of noninvasive cardiovascular tests, calculated from sensitivity and specificity measures or as ratios of the predictive value odds to pretest odds for positive and negative test results. The positive likelihood ratio, (+)LR, expresses the power of a positive test result to augment an estimate of disease probability independent of the pretest prevalence of disease in a given population; the negative likelihood ratio, (-)LR, expresses the power of a negative test result to augment an estimate of the probability of no disease independent of the pretest prevalence of no disease in the same population. The likelihood ratio principle is applicable to the evaluation of the predictive power of single or combined test results reported for either dichotomous or continuous end points. This part of the perspective exemplifies application of the likelihood ratio principle in a wide variety of testing conditions for coronary artery disease followed by a discussion of the limitations of likelihood ratio computation in test power evaluation. Likelihood ratios provide a more concise and unambiguous standard for calibrating the predictive power of single and combined noninvasive cardiovascular test results than are provided by measures of sensitivity, specificity, and predictive value.

Adrenergic beta-Agonists↗

Likelihood-based confidence intervals for a log-normal mean.

To construct a confidence interval for the mean of a log-normal distribution in small samples, we propose likelihood-based approaches - the signed log-likelihood ratio and modified signed log-likelihood ratio methods. Extensive Monte Carlo simulation results show the advantages of the modified signed log-likelihood ratio method over the signed log-likelihood ratio method and other methods. In particular, the modified signed log-likelihood ratio method produces a confidence interval with a nearly exact coverage probability and highly accurate and symmetric error probabilities even for extremely small sample sizes. We then apply the methods to two sets of real-life data.

Biomedical Research↗

Pseudo-likelihood methods for longitudinal binary data with non-ignorable missing responses and covariates.

In this paper we consider longitudinal studies in which the outcome to be measured over time is binary, and the covariates of interest are categorical. In longitudinal studies it is common for the outcomes and any time-varying covariates to be missing due to missed study visits, resulting in non-monotone patterns of missingness. Moreover, the reasons for missed visits may be related to the specific values of the response and/or covariates that should have been obtained, i.e. missingness is non-ignorable. With non-monotone non-ignorable missing response and covariate data, a full likelihood approach is quite complicated, and maximum likelihood estimation can be computationally prohibitive when there are many occasions of follow-up. Furthermore, the full likelihood must be correctly specified to obtain consistent parameter estimates. We propose a pseudo-likelihood method for jointly estimating the covariate effects on the marginal probabilities of the outcomes and the parameters of the missing data mechanism. The pseudo-likelihood requires specification of the marginal distributions of the missingness indicator, outcome, and possibly missing covariates at each occasions, but avoids making assumptions about the joint distribution of the data at two or more occasions. Thus, the proposed method can be considered semi-parametric. The proposed method is an extension of the pseudo-likelihood approach in Troxel et al. to handle binary responses and possibly missing time-varying covariates. The method is illustrated using data from the Six Cities study, a longitudinal study of the health effects of air pollution.

Air Pollutants↗

Likelihood ratio for trisomy 21 in fetuses with tricuspid regurgitation at the 11 to 13 + 6-week scan.

OBJECTIVE: To determine the likelihood ratio for trisomy 21 in fetuses with tricuspid regurgitation at the 11 to 13 + 6-week scan. METHODS: Fetal echocardiography was carried out by specialist pediatric cardiologists in 742 singleton pregnancies at 11 to 13 + 6 weeks' gestation and pulsed wave Doppler was used to ascertain the presence or absence of tricuspid regurgitation. To avoid confusion with other adjacent signals, a strict definition of tricuspid regurgitation was used, in that it had to occupy at least half of systole and reach a velocity of over 80 cm/s. The fetal crown-rump length (CRL) and the nuchal translucency (NT) thickness were measured and the presence of any congenital heart abnormality noted. Follow-up of the pregnancy was carried out to determine the presence of chromosomal abnormalities. The likelihood ratio for trisomy 21 in fetuses with and without tricuspid regurgitation was determined. RESULTS: The tricuspid valve was successfully examined in 718 (96.8%) cases. Tricuspid regurgitation was present in 39 (8.5%) of the 458 chromosomally normal fetuses, in 82 (65.1%) of the 126 with trisomy 21, in 44 (53.0%) of the 83 with trisomy 18 or 13, and in 11 (21.6%) of the 51 with other chromosomal defects. The prevalence of tricuspid regurgitation was also associated with fetal CRL, delta NT and the presence of cardiac defects. Logistic regression analysis, irrespective of cardiac defects, demonstrated that in the chromosomally normal fetuses significant independent prediction of the likelihood of tricuspid regurgitation was provided by fetal delta NT (odds ratio (OR), 1.26; 95% CI, 1.34-1.41; P < 0.0001), while in trisomy 21 fetuses prediction was provided by CRL (OR, 0.94; 95% CI, 0.89-0.99; P = 0.021). The likelihood ratio for trisomy 21 for tricuspid regurgitation was derived by dividing the likelihood in trisomy 21 by that in normal fetuses. In the chromosomally normal fetuses, the prevalence of tricuspid regurgitation in those with cardiac defects was 46.9% and 5.6% in those without cardiac defects, and the likelihood ratio of tricuspid regurgitation for cardiac defects was 8.4. CONCLUSION: At 11 to 13 + 6 weeks' gestation, there is a high association between tricuspid regurgitation and trisomy 21, as well as other chromosomal defects. The prevalence of tricuspid regurgitation increases with fetal NT thickness and is substantially higher in those with, than those without, a cardiac defect.

Crown-Rump Length↗

The self-reported likelihood of patient delay in breast cancer: new thoughts for early detection.

BACKGROUND: Delayed presentation of self-discovered breast symptoms influences stage of cancer at diagnosis and decreases breast cancer survival. METHODS: A total of 699 asymptomatic women (black, white, and Latino), recruited in community settings and stratified by age, income, and educational level, were surveyed for their likelihood to delay (J-Delay scale) in the event of a breast symptom discovery. Models of likelihood were tested with logistic regression analyses. RESULTS: A total of 166 women (23.7%) reported likelihood to delay. Lower income, lower educational level, self identification as Latino or black, experienced prejudice in care delivery, perceived lack of access to health care, fatalism about breast cancer, poor health care utilization habits, self-care behavior, spouse/partner and employer perceived constraints, problem-solving style, and a lack of knowledge of breast cancer's presenting symptoms were associated with likelihood to delay. A combined sample multiple logistic regression model correctly predicted 40.6% of women reporting a likelihood to delay, 94.9% of those not likely to delay, and 82.4% (551 of 669) of cases overall. CONCLUSIONS: Self-reported likelihood of patient delay is measurable in advance of symptom occurrence, and this measure is consistent with behavioral and knowledge variables previously linked with advanced breast cancer at diagnosis.

Adult↗

Phylogenetic analysis using parsimony and likelihood methods.

The assumptions underlying the maximum-parsimony (MP) method of phylogenetic tree reconstruction were intuitively examined by studying the way the method works. Computer simulations were performed to corroborate the intuitive examination. Parsimony appears to involve very stringent assumptions concerning the process of sequence evolution, such as constancy of substitution rates between nucleotides, constancy of rates across nucleotide sites, and equal branch lengths in the tree. For practical data analysis, the requirement of equal branch lengths means similar substitution rates among lineages (the existence of an approximate molecular clock), relatively long interior branches, and also few species in the data. However, a small amount of evolution is neither a necessary nor a sufficient requirement of the method. The difficulties involved in the application of current statistical estimation theory to tree reconstruction were discussed, and it was suggested that the approach proposed by Felsenstein (1981, J. Mol. Evol. 17: 368-376) for topology estimation, as well as its many variations and extensions, differs fundamentally from the maximum likelihood estimation of a conventional statistical parameter. Evidence was presented showing that the Felsenstein approach does not share the asymptotic efficiency of the maximum likelihood estimator of a statistical parameter. Computer simulations were performed to study the probability that MP recovers the true tree under a hierarchy of models of nucleotide substitution; its performance relative to the likelihood method was especially noted. The results appeared to support the intuitive examination of the assumptions underlying MP. When a simple model of nucleotide substitution was assumed to generate data, the probability that MP recovers the true topology could be as high as, or even higher than, that for the likelihood method. When the assumed model became more complex and realistic, e.g., when substitution rates were allowed to differ between nucleotides or across sites, the probability that MP recovers the true topology, and especially its performance relative to that of the likelihood method, generally deteriorates. As the complexity of the process of nucleotide substitution in real sequences is well recognized, the likelihood method appears preferable to parsimony. However, the development of a statistical methodology for the efficient estimation of the tree topology remains a difficult open problem.

Animals↗

On the bootstrap and monotone likelihood in the cox proportional hazards regression model.

Recent literature has provided encouragement for using the bootstrap for inference on regression parameters in the Cox proportional hazards (PH) model. However, generating and performing the necessary partial likelihood computations on multitudinous bootstrap samples greatly increases the chances of incurring problems with monotone likelihood at some point in the analysis. The only symptom of monotone likelihood may be a failure to converge in the numerical maximization procedure, and so the problem might naïvely be dismissed by deleting the offending data set and replacing it with a new one. This strategy is shown to lead to potentially high selection biases in the subsequent summary statistics. This note discusses the importance of keeping track of these monotone likelihood cases and provides recommendations for their use in interpreting bootstrap findings, and for avoiding unwanted biases that may result from high rates of occurrence. In many cases, high monotone likelihood rates indicate that a more highly-specified model may be preferred. Special consideration is given to the problem of high monotone likelihood incidence in Monte Carlo studies of the bootstrap.

Biometry↗

Ascertainment-adjusted maximum likelihood estimation for the additive genetic gamma frailty model.

The additive genetic gamma frailty model has been proposed for genetic linkage analysis for complex diseases to account for variable age of onset and possible covariates effects. To avoid ascertainment biases in parameter estimates, retrospective likelihood ratio tests are often used, which may result in loss of efficiency due to conditioning. This paper considers when the sibships are ascertained by having at least two affected sibs with the disease before a given age and provides two approaches for estimating the parameters in the additive gamma frailty model. One approach is based on the likelihood function conditioning on the ascertainment event, the other is based on maximizing a full ascertainment-adjusted likelihood. Explicit forms for these likelihood functions are derived. Simulation studies indicate that when the baseline hazard function can be correctly pre-specified, both approaches give accurate estimates of the model parameters. However, when the baseline hazard function has to be estimated simultaneously, only the ascertainment-adjusted likelihood method gives an unbiased estimate of the parameters. These results imply that the ascertainment-adjusted likelihood ratio test in the context of the additive genetic gamma frailty may be used for genetic linkage analysis.

Age of Onset↗

Upper bounds on maximum likelihood for phylogenetic trees.

We introduce a mechanism for analytically deriving upper bounds on the maximum likelihood for genetic sequence data on sets of phylogenies. A simple 'partition' bound is introduced for general models. Tighter bounds are developed for the simplest model of evolution, the two state symmetric model of nucleotide substitution under the molecular clock. This follows earlier theoretical work which has been restricted to this model by analytic complexity. A weakness of current numerical computation is that reported 'maximum likelihood' results cannot be guaranteed, both for a specified tree (because of the possibility of multiple maxima) or over the full tree space (as the computation is intractable for large sets of trees). The bounds we develop here can be used to conclusively eliminate large proportions of tree space in the search for the maximum likelihood tree. This is vital in the development of a branch and bound search strategy for identifying the maximum likelihood tree. We report the results from a simulation study of approximately 10(6) data sets generated on clock-like trees of five leaves. In each trial a likelihood value of one specific instance of a parameterised tree is compared to the bound determined for each of the 105 possible rooted binary trees. The proportion of trees that are eliminated from the search for the maximum likelihood tree ranged from 92% to almost 98%, indicating a computational speed-up factor of between 12 and 44.

Algorithms↗

A genetic algorithm for maximum-likelihood phylogeny inference using nucleotide sequence data.

Phylogeny reconstruction is a difficult computational problem, because the number of possible solutions increases with the number of included taxa. For example, for only 14 taxa, there are more than seven trillion possible unrooted phylogenetic trees. For this reason, phylogenetic inference methods commonly use clustering algorithms (e.g., the neighbor-joining method) or heuristic search strategies to minimize the amount of time spent evaluating nonoptimal trees. Even heuristic searches can be painfully slow, especially when computationally intensive optimality criteria such as maximum likelihood are used. I describe here a different approach to heuristic searching (using a genetic algorithm) that can tremendously reduce the time required for maximum-likelihood phylogenetic inference, especially for data sets involving large numbers of taxa. Genetic algorithms are simulations of natural selection in which individuals are encoded solutions to the problem of interest. Here, labeled phylogenetic trees are the individuals, and differential reproduction is effected by allowing the number of offspring produced by each individual to be proportional to that individual's rank likelihood score. Natural selection increases the average likelihood in the evolving population of phylogenetic trees, and the genetic algorithm is allowed to proceed until the likelihood of the best individual ceases to improve over time. An example is presented involving rbcL sequence data for 55 taxa of green plants. The genetic algorithm described here required only 6% of the computational effort required by a conventional heuristic search using tree bisection/reconnection (TBR) branch swapping to obtain the same maximum-likelihood topology.

Algorithms↗

Likelihood analysis of phylogenetic networks using directed graphical models.

A method for computing the likelihood of a set of sequences assuming a phylogenetic network as an evolutionary hypothesis is presented. The approach applies directed graphical models to sequence evolution on networks and is a natural generalization of earlier work by Felsenstein on evolutionary trees, including it as a special case. The likelihood computation involves several steps. First, the phylogenetic network is rooted to form a directed acyclic graph (DAG). Then, applying standard models for nucleotide/amino acid substitution, the DAG is converted into a Bayesian network from which the joint probability distribution involving all nodes of the network can be directly read. The joint probability is explicitly dependent on branch lengths and on recombination parameters (prior probability of a parent sequence). The likelihood of the data assuming no knowledge of hidden nodes is obtained by marginalization, i.e., by summing over all combinations of unknown states. As the number of terms increases exponentially with the number of hidden nodes, a Markov chain Monte Carlo procedure (Gibbs sampling) is used to accurately approximate the likelihood by summing over the most important states only. Investigating a human T-cell lymphotropic virus (HTLV) data set and optimizing both branch lengths and recombination parameters, we find that the likelihood of a corresponding phylogenetic network outperforms a set of competing evolutionary trees. In general, except for the case of a tree, the likelihood of a network will be dependent on the choice of the root, even if a reversible model of substitution is applied. Thus, the method also provides a way in which to root a phylogenetic network by choosing a node that produces a most likely network.

Computer Graphics↗

Semiparametric maximum likelihood for measurement error model regression.

This paper presents an EM algorithm for semiparametric likelihood analysis of linear, generalized linear, and nonlinear regression models with measurement errors in explanatory variables. A structural model is used in which probability distributions are specified for (a) the response and (b) the measurement error. A distribution is also assumed for the true explanatory variable but is left unspecified and is estimated by nonparametric maximum likelihood. For various types of extra information about the measurement error distribution, the proposed algorithm makes use of available routines that would be appropriate for likelihood analysis of (a) and (b) if the true x were available. Simulations suggest that the semiparametric maximum likelihood estimator retains a high degree of efficiency relative to the structural maximum likelihood estimator based on correct distributional assumptions and can outperform maximum likelihood based on an incorrect distributional assumption. The approach is illustrated on three examples with a variety of structures and types of extra information about the measurement error distribution.

Algorithms↗

Penalized partial likelihood regression for right-censored data with bootstrap selection of the penalty parameter.

The Cox proportional hazards model is often used for estimating the association between covariates and a potentially censored failure time, and the corresponding partial likelihood estimators are used for the estimation and prediction of relative risk of failure. However, partial likelihood estimators are unstable and have large variance when collinearity exists among the explanatory variables or when the number of failures is not much greater than the number of covariates of interest. A penalized (log) partial likelihood is proposed to give more accurate relative risk estimators. We show that asymptotically there always exists a penalty parameter for the penalized partial likelihood that reduces mean squared estimation error for log relative risk, and we propose a resampling method to choose the penalty parameter. Simulations and an example show that the bootstrap-selected penalized partial likelihood estimators can, in some instances, have smaller bias than the partial likelihood estimators and have smaller mean squared estimation and prediction errors of log relative risk. These methods are illustrated with a data set in multiple myeloma from the Eastern Cooperative Oncology Group.

Antineoplastic Agents↗