PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Likelihood Functions”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Localization-based sensor validation using the Kullback-Leibler divergence.

A sensor validation criteria based on the sensor's object localization accuracy is proposed. Assuming that the true probability distribution of an object or event in space f(x) is known and a spatial likelihood function (SLF) psi(x) for the same object or event in space is obtained from a sensor, then the expected value of the SLF E[psi(x)] is proposed as a suitable validity metric for the sensor, where the expectation is performed over the distribution f(x). It is shown that for the class of increasing linear log likelihood SLFs, the proposed validity metric is equivalent to the Kullback-Leibler distance between f(x) and the unknown sensor-based distribution g(x) where the SLF psi(x) is an observable increasing function of the unobservable g(x). The proposed technique is illustrated through several simulated and experimental examples.

Journal Article↗

A comment on Heathcote, Brown, and Mewhort's QMLE method for response time distributions.

Heathcote, Brown, and Mewhort (2002) have introduced a new, robust method of estimating response time distributions. Their method may have practical advantages over conventional maximum likelihood estimation. The basic idea is that the likelihood of parameters is maximized given a few quantiles from the data. We show that Heathcote et al.'s likelihood function is not correct and provide the appropriate correction. However, although our correction stands on firmer theoretical ground than Heathcote et al.'s, it appears to yield worse parameter estimates. This result further indicates that, at least for some distributions and situations, quantile maximum likelihood estimation may have better nonasymptotic properties than a more theoretically justified approach.

Humans↗

An E-M algorithm and testing strategy for multiple-locus haplotypes.

This paper gives an expectation maximization (EM) algorithm to obtain allele frequencies, haplotype frequencies, and gametic disequilibrium coefficients for multiple-locus systems. It permits high polymorphism and null alleles at all loci. This approach effectively deals with the primary estimation problems associated with such systems; that is, there is not a one-to-one correspondence between phenotypic and genotypic categories, and sample sizes tend to be much smaller than the number of phenotypic categories. The EM method provides maximum-likelihood estimates and therefore allows hypothesis tests using likelihood ratio statistics that have chi 2 distributions with large sample sizes. We also suggest a data resampling approach to estimate test statistic sampling distributions. The resampling approach is more computer intensive, but it is applicable to all sample sizes. A strategy to test hypotheses about aggregate groups of gametic disequilibrium coefficients is recommended. This strategy minimizes the number of necessary hypothesis tests while at the same time describing the structure of disequilibrium. These methods are applied to three unlinked dinucleotide repeat loci in Navajo Indians and to three linked HLA loci in Gila River (Pima) Indians. The likelihood functions of both data sets are shown to be maximized by the EM estimates, and the testing strategy provides a useful description of the structure of gametic disequilibrium. Following these applications, a number of simulation experiments are performed to test how well the likelihood-ratio statistic distributions are approximated by chi 2 distributions. In most circumstances the chi 2 grossly underestimated the probability of type I errors. However, at times they also overestimated the type 1 error probability. Accordingly, we recommended hypothesis tests that use the resampling method.

Algorithms↗

Introduction to Bayesian methods I: measuring the strength of evidence.

Bayesian inference is a formal method to combine evidence external to a study, represented by a prior probability curve, with the evidence generated by the study, represented by a likelihood function. Because Bayes theorem provides a proper way to measure and to combine study evidence, Bayesian methods can be viewed as a calculus of evidence, not just belief. In this introduction, we explore the properties and consequences of using the Bayesian measure of evidence, the Bayes factor (in its simplest form, the likelihood ratio). The Bayes factor compares the relative support given to two hypotheses by the data, in contrast to the P-value, which is calculated with reference only to the null hypothesis. This comparative property of the Bayes factor, combined with the need to explicitly predefine the alternative hypothesis, produces a different assessment of the strength of evidence against the null hypothesis than does the P-value, and it gives Bayesian procedures attractive frequency properties. However, the most important contribution of Bayesian methods is the way in which they affect both who participates in a scientific dialogue, and what is discussed. With the emphasis moved from "error rates" to evidence, content experts have an opportunity for their input to be meaningfully incorporated, making it easier for regulatory decisions to be made correctly.

Bayes Theorem↗

Optimal sampling for pedigree analysis: sequential schemes for sibships.

Methods for inferring the mode of inheritance of a trait from familial data are becoming widely used. It is therefore important to assess alternative procedures for the collection of the relevant data. Cannings and Thompson (1977, Clinical Genetics 12, 208-212) and Thompson and Cannings (1979. In Genetic Analysis of Common Diseases, 363-382, New York: Liss) have advocated sequential schemes on the grounds that these admit simple methods for the correction of the likelihood function for ascertainment bias. Here it is shown that sequential procedures may also greatly increase efficiency, as measured by (statistical) information gained per individual sampled. Although attention is restricted to sampling sibships, and to a simple genetic model, the measures introduced are more widely applicable. A practical procedure for the construction of schemes, via a relationship between expected log likelihood and entropy, is also presented. This too is more widely applicable. A numerical example demonstrates the gains which can be achieved in practice, relative to alternative hypotheses which have been considered in several medical-genetic studies.

Data Collection↗

A new class of parametric IRT models for dichotomous item scores.

A new class of parametric IRT models for dichotomously scored items is presented. The new class of models is a subclass of both the class of models defined by the four-parameter logistic item response function and the nonparametric Double Monotonicity (DM) model. Three special cases of this new class of models are discussed. One of these special cases is shown to be the one-parameter logistic Rasch model. Both specific objectivity at the interval level of measurement and the sufficiency of the total score for the latent trait are shown to be measurement properties of the whole new class of models. For maximum likelihood estimation of the model parameters, both a joint and a conditional likelihood function are proposed.

Humans↗

Limits of quantal analysis reliability: quantal and unimodal constraints and setting of confidence intervals for quantal size.

An accurate objective method for determining the reliability of estimates of quantal size (Q) at central synapses was developed. To do this, distributions of amplitudes of postsynaptic responses were simulated by convolving a number of discrete amplitudes separated by equal increments Q with gaussian noise, after which the value of Q was estimated by the maximum likelihood method under different constraints on the discrete distribution. It was shown that the likelihood function (LF) had several local maxima under the quantal constraint, and, if the value of the ratio between Q and the standard deviation of the noise (sigma) was less than 3, the global maximum of the LF corresponded to a biased estimate of Q lying in a range of values less than 1.5 sigma. The best estimates of Q were obtained when unimodal discrete distributions of amplitudes resulting from the maximum likelihood method were selected. However, this method also gave biased estimates when Q/sigma was less than 1.5-2.3. The limit of reliability depended on the number of discrete components and the sample size. To calculate confidence intervals for the quantal size, different numbers and weights of components were used to simulate amplitude histograms with different values of Q/sigma. Three data sets were used to illustrate the procedure.

Animals↗

Weighted neighbor joining: a likelihood-based approach to distance-based phylogeny reconstruction.

We introduce a distance-based phylogeny reconstruction method called "weighted neighbor joining," or "Weighbor" for short. As in neighbor joining, two taxa are joined in each iteration; however, the Weighbor criterion for choosing a pair of taxa to join takes into account that errors in distance estimates are exponentially larger for longer distances. The criterion embodies a likelihood function on the distances, which are modeled as correlated Gaussian random variables with different means and variances, computed under a probabilistic model for sequence evolution. The Weighbor criterion consists of two terms, an additivity term and a positivity term, that quantify the implications of joining the pair. The first term evaluates deviations from additivity of the implied external branches, while the second term evaluates confidence that the implied internal branch has a positive branch length. Compared with maximum-likelihood phylogeny reconstruction, Weighbor is much faster, while building trees that are qualitatively and quantitatively similar. Weighbor appears to be relatively immune to the "long branches attract" and "long branch distracts" drawbacks observed with neighbor joining, BIONJ, and parsimony.

Animals↗

A comparison of the logistic risk function and the proportional hazards model in prospective epidemiologic studies.

The logistic regression and proportional hazards models are each currently being used in the analysis of prospective epidemiologic studies examining risk factors in chronic disease applications. The advantages and disadvantages of each are yet to be fully described. However, a theoretical relationship between the two models has been documented. In this paper the conditions under which results from the two models approximate one another are described. It is shown that where the follow-up period is short and the disease is generally rare, the regression coefficients of the logistic model approximate those of the proportional hazards model with a constant underlying hazard rate. Since under the same conditions the likelihood functions approximate one another, the regression coefficients have similar estimated standard errors. Further, estimation of relative risk with these models is contrasted. These results are illustrated utilizing a previously published data set on metastatic cancer of the breast. With increasing follow-up time, the logistic regression coefficients become uncertain and less reliable.

Epidemiologic Methods↗

Refinement of macromolecular structures by the maximum-likelihood method.

This paper reviews the mathematical basis of maximum likelihood. The likelihood function for macromolecular structures is extended to include prior phase information and experimental standard uncertainties. The assumption that different parts of a structure might have different errors is considered. A method for estimating sigma(A) using 'free' reflections is described and its effects analysed. The derived equations have been implemented in the program REFMAC. This has been tested on several proteins at different stages of refinement (bacterial alpha-amylase, cytochrome c', cross-linked insulin and oligopeptide binding protein). The results derived using the maximum-likelihood residual are consistently better than those obtained from least-squares refinement.

Journal Article↗

Estimation of variance components including competitive effects of Large White growing gilts.

Records of on-test ADG of Large White gilts were analyzed to estimate variance components of direct and associative genetic effects. Models included the effects of contemporary group (farm-barn-batch), birth litter, pen group, and direct and associative additive genetic effects. The area of each pen was 14 m2. The additive genetic variance was a function of the number of competitors in a group, the additive relationships between the animal performing the record and its pen mates, and the additive relationships between pen mates. To partially account for differences in the number of pen mates, a covariable (qi = 1, 1/n, or 1/n(1/2)) was added to the associative genetic effect. There were 4,946 records from 2,409 litters and 362 pen groups. Pen group size ranged from 12 to 16 gilts. Analyses by REML converged very slowly. A grid search showed that the likelihood function was almost flat when the additive genetic associative effect was fitted. Estimates of direct and associative heritability were 0.15 and 0.03, respectively. Within the BLUPF90 family of programs, the mixed-model equations can be set up directly. For variance component estimation, simple programs (REMLF90 and GIBBSF90) worked without modifications, but more optimized programs did not. Estimates obtained using the three values of qi were similar. With the data structure available for this study and under an environment with relative low competition among animals, accurate estimation of associative genetic effects was not possible. Estimation of competitive effects with large pen size is difficult. The magnitude of competition effects may be larger in commercial populations, where housing is denser and food is limited.

Analysis of Variance↗

Modeling lactation curves and estimation of genetic parameters for first lactation test-day records of French Holstein cows.

Several functions were used to model the fixed part of the lactation curve and genetic parameters of milk test-day records to estimate using French Holstein data. Parametric curves (Legendre polynomials, Ali-Schaeffer curve, Wilmink curve), fixed classes curves (5-d classes), and regression splines were tested. The latter were appealing because they adjusted the data well, were relatively insensitive to outliers, were flexible, and resulted in smooth curves without requiring the estimation of a large number of parameters. Genetic parameters were estimated with an Average Information REML algorithm where the average information matrix and the first derivatives of the likelihood functions were pooled over 10 samples. This approach made it possible to handle larger data sets. The residual variance was modeled as a quadratic function of days in milk. Quartic Legendre polynomials were used to estimate (co)variances of random effects. The estimates were within the range of most other studies. The greatest genetic variance was in the middle of the lactation while residual and permanent environmental variances mostly decreased during the lactation. The resulting heritability ranged from 0.15 to 0.40. The genetic correlation between the extreme parts of the lactation was 0.35 but genetic correlations were higher than 0.90 for a large part of the lactation. The use of the pooling approach resulted in smaller standard errors for the genetic parameters when compared to those obtained with a single sample.

Animals↗

Analysis of the relationship between type traits, inbreeding, and functional survival in Jersey cattle using a Weibull proportional hazards model.

A Weibull proportional hazards model was used to analyze the effects of 13 linear type traits, final score, and inbreeding on the functional survival of 268,008 US Jersey cows in 2416 herds with first calving from 1981 to 2000. Functional survival was defined as the number of days from first calving until involuntary culling or censoring. The statistical model included the time-dependent effects of herd-year-season of calving, parity by stage of lactation interaction, and within-herd-year quintile for mature equivalent milk yield, as well as the time-independent effects of inbreeding, age at first calving, and linear type traits or final score (analyzed one at a time). Each type trait was divided into 10 classes, and the relative risk of involuntary culling was calculated for animals in each class after accounting for the aforementioned management factors. Type traits with the greatest contribution to the likelihood function were udder depth, fore udder attachment, front teat placement, and udder support. Cows with low scores for these traits had a risk of culling that was 1.3 to 1.8 times that of cows with intermediate scores. Cows with high scores for udder depth and udder support had a risk of culling only 0.7 to 0.85 as great as that of cows with intermediate scores. Intermediate scores were desirable for rear leg set, dairy form, and strength, but stature, rump angle, and rump width had negligible effects on survival. Cows with low final scores had a risk of culling that was 1.35 times that of cows with intermediate scores, whereas cows with high final scores had a risk of culling that was 0.8 times that of cows with intermediate scores. Animals with inbreeding coefficients greater than 10% had a slightly higher risk of culling than animals with inbreeding coefficients less than 5%.

Animals↗

Predicting the outcome of radiotherapy for prostate carcinoma: a model-building strategy.

BACKGROUND: Clinical research of prostate carcinoma could be enhanced by models that allow early and reliable prediction of outcome. In this study, the authors describe a model-building strategy and compare different models. METHODS: The sample population was comprised of 158 patients treated definitively with radiotherapy. Univariate and multivariate logistic regression analyses were conducted to identify prognostic factors and select the best predictive model. Variables included age, race, method of diagnosis (needle biopsy vs. transurethral resection of the prostate), stage, grade, pretreatment prostate specific antigen (PSA), in-treatment PSA (PSA(tx)), posttreatment PSA (PSA(post)), and nadir PSA. The following indices were used to compare discriminatory power: log-likelihood function, Akaike information criterion, the generalized coefficient of determination, and the area under the receiver operating characteristic curve. RESULTS: At last follow-up, 49 patients (31%) had recurrence of carcinoma. By univariate analysis, the failure rate was significantly higher in patients with advanced stage, higher grade, higher pretherapy PSA, and nadir PSA > 1 ng/mL (P < 0.0001). Pretherapy PSA was associated significantly with stage, age, and nadir PSA (P = 0.001, P = 0.001, and P = 0.001, respectively). All PSA measurements were significantly interrelated. Nadir PSA was the most predictive variable. Significant gains (P = 0.01) in predictive power were derived from inclusion of PSA(tx), but not PSA (post). Age, race, stage, grade, and method of diagnosis contributed predictive power in addition to that derived from PSA levels (P = 0.01, log-likelihood test). The authors' model of choice predicts outcome with an overall correctness, sensitivity, specificity, and false-negative rate of 81.8%, 87.2%, 79.6%, and 12.8%, respectively. CONCLUSIONS: Applying the strategy described, a model was selected that allowed accurate prediction of failure shortly after the completion of therapy.

Aged↗

Relationship between type traits and longevity in Canadian Jerseys and Ayrshires using a Weibull proportional hazards model.

The aim of this study was to use a Weibull proportional hazards model to explore the impact of type traits on the functional survival of Canadian Jersey and Ayrshire cows. The data set consisted of 49,791 registered Jersey cows from 900 herds calving from 1985 to 2003. The corresponding figures for Ayrshire were 77,109 cows and 921 herds. Functional survival was defined as the number of days from first calving to culling, death, or censoring. Type information consisted of phenotypic type scores for 8 composite traits and 19 linear descriptive traits. The statistical model included the effects of stage of lactation; season of production; annual change in herd size; type of milk recording supervision; age at first calving; effects of milk, fat, and protein yields calculated as within herd-year-parity deviations; herd-year-season of calving; each type trait; and the animal's sire. Analysis was done one trait at a time for each of 27 type traits in each breed. The relative culling risk was calculated for animals in each class after accounting for the previously mentioned effects. Among the composite type traits with the greatest contribution to the likelihood function was final score followed by mammary system for Jersey breed, while in Ayrshire breed feet and legs was the second most important trait next to final score. Cows classified as Poor for final score in both breeds were >5 times more likely to be culled compared with the cows classified as Good Plus. In both breeds, cows classified as Poor for feet and legs were 5 times more likely to be culled than were cows classified as Excellent, and cows classified as Excellent for mammary system were >9 times more likely to survive than were cows classified as Poor.

Animals↗

Construction of conditional lod tables from multiple-locus linkage data.

Although multipoint linkage data are becoming quite common, economical and efficient methods for presenting these data are not yet in use. Tables giving the full likelihood function would be very voluminous, whereas reduction to standard lods destroys part of the information. We suggest a special lod table representation which preserves nearly all the information about both gene order and genetic map distance at a great reduction in data presented. Conditional lods, coined "c-lods" to distinguish them clearly from ordinary lods, are calculated for distances between each adjacent pair of loci, with all distances among other loci in the system conditionally optimized. Thus, in an n-locus system, conditional lods are presented for the n - 1 adjacent pairs of loci only. The principal reason for reporting lod scores is the potential for combining data from separate studies to reach conclusive evidence for linkage and gene order. Because estimates of recombination are different in each set of data, the crux of the problem is to present scores that provide a close approximation to the true likelihood away from maximum likelihood (ML). The sum of conditional lods closely approximates the likelihood throughout the domain. Therefore, conditional lods contribute appropriately when mapping from several sources of data, so contradictory estimates can be reconciled efficiently.

Biometry↗

Applications of segmented regression models for biomedical studies.

In many biological models, a relationship between variables may be modeled as a linear or polynomial function that changes abruptly when an independent variable obtains a threshold level. Usually, the transition point is unknown, and a major objective of the analysis is its estimation. This type of model is known as a segmented regression model. We present two methods, Gallant and Fuller's (J Am. Stat. Assoc. 68: 144-147, 1973) method and Tishler and Zang's (J. Am. Stat. Assoc. 76: 980-987, 1981) method, using nonlinear least-squares techniques for estimating the transition point. We give the following three examples: a hypoglycemia study, a testosterone study, and an estimate of age-cortisol relationship. Simulation techniques are used to compare the two methods. We conclude that these models provide useful information and that the two methods studied produce essentially equivalent results. We recommend that both methods be used to analyze a data set if possible to avoid problems due to local minima and that if the results do not agree, then evaluation of the likelihood function in the range of the estimates be used to determine the best estimate.

Aging↗

Advantages of terminating Zippy Estimation by Sequential Testing (ZEST) with dynamic criteria for white-on-white perimetry.

PURPOSE: A number of automated perimeters use the Zippy Estimation by Sequential Testing (ZEST) algorithm, which is an adaptive Bayesian method, for determining sensitivity measures. There are two popular rules for deciding when to terminate Bayesian procedures: (1) after a fixed number of presentations; or (2) when the probability density function (pdf) over all thresholds modified by the procedure becomes sufficiently narrow (a dynamic termination criterion). It has recently been argued that fixed termination criteria perform equally as well as dynamic criteria when applied in a fashion typical of laboratory-based visual psychophysics. Perimetry, however, has specific requirements; the tests must be very short, there is a wide range of possible sensitivities, and erroneous responses from the patient must be tolerated. This study used computer simulation to compare fixed and dynamic termination criteria for the ZEST algorithm using conditions typical of white-on-white perimetry. METHODS: Eight ZEST procedures were compared using the following termination criteria: fixed termination after 4, 5, 6, 7, and 8 presentations; dynamic termination when the standard deviation of the pdf was 1 dB, 1.5 dB, and 2 dB. Four patient error models were used: ideal, typical false-positive, typical false-negative, and unreliable patients. We also ran a version of ZEST that set the likelihood function exactly equal to the patient's frequency of seeing curve. RESULTS: The mean absolute error and standard deviation of error in threshold measurement was higher for the fixed termination criteria than for dynamic termination criteria of the same average number of presentations. CONCLUSIONS: The results of our simulations indicate that dynamic procedures have some distinct benefits over fixed termination procedures when a minimum of presentations are required and response errors are made as in a white-on-white perimetric setting. Dynamic termination criteria are at least partially successful in expending more presentations when required to enhance test precision.

Algorithms↗