PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Likelihood Functions”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Power and related statistical properties of conditional likelihood score tests for association studies in nuclear families with parental genotypes.

Both population based and family based case control studies are used to test whether particular genotypes are associated with disease. While population based studies have more power, cryptic population stratification can produce false-positive results. Family-based methods have been introduced to control for this problem. This paper presents the full likelihood function for family-based association studies for nuclear families ascertained on the basis of their number of affected and unaffected children. The likelihood of a family factors into the probability of parental mating type, conditional on offspring phenotypes, times the probability of offspring genotypes given their phenotypes and the parental mating type. The first factor can be influenced by population stratification, whereas the latter factor, called the conditional likelihood, is not. The conditional likelihood is used to obtain score tests with proper size in the presence of population stratification (see also Clayton (1999) and Whittemore & Tu (2000)). Under either the additive or multiplicative model, the TDT is known to be the optimal score test when the family has only one affected child. Thus, the class of score tests explored can be considered as a general family of TDT-like procedures. The relative informativeness of the various mating types is assessed using the Fisher information, which depends on the number of affected and unaffected offspring and the penetrances. When the additive model is true, families with parental mating type Aa x Aa are most informative. Under the dominant (recessive) model, however, a family with mating type Aa x aa(AA x Aa) is more informative than a family with doubly heterozygous (Aa x Aa) parents. Because we derive explicit formulae for all components of the likelihood, we are able to present tables giving required sample sizes for dominant, additive and recessive inheritance models.

Case-Control Studies↗

Estimation of population parameters and recombination rates from single nucleotide polymorphisms.

Some general likelihood and Bayesian methods for analyzing single nucleotide polymorphisms (SNPs) are presented. First, an efficient method for estimating demographic parameters from SNPs in linkage equilibrium is derived. The method is applied in the estimation of growth rates of a human population based on 37 SNP loci. It is demonstrated how ascertainment biases, due to biased sampling of loci, can be avoided, at least in some cases, by appropriate conditioning when calculating the likelihood function. Second, a Markov chain Monte Carlo (MCMC) method for analyzing linked SNPs is developed. This method can be used for Bayesian and likelihood inference on linked SNPs. The utility of the method is illustrated by estimating recombination rates in a human data set containing 17 SNPs and 60 individuals. Both methods are based on assumptions of low mutation rates.

Genetic Linkage↗

Non-parametric estimation for baseline hazards function and covariate effects with time-dependent covariates.

Often in many biomedical and epidemiologic studies, estimating hazards function is of interest. The Breslow's estimator is commonly used for estimating the integrated baseline hazard, but this estimator requires the functional form of covariate effects to be correctly specified. It is generally difficult to identify the true functional form of covariate effects in the presence of time-dependent covariates. To provide a complementary method to the traditional proportional hazard model, we propose a tree-type method which enables simultaneously estimating both baseline hazards function and the effects of time-dependent covariates. Our interest will be focused on exploring the potential data structures rather than formal hypothesis testing. The proposed method approximates the baseline hazards and covariate effects with step-functions. The jump points in time and in covariate space are searched via an algorithm based on the improvement of the full log-likelihood function. In contrast to most other estimating methods, the proposed method estimates the hazards function rather than integrated hazards. The method is applied to model the risk of withdrawal in a clinical trial that evaluates the anti-depression treatment in preventing the development of clinical depression. Finally, the performance of the method is evaluated by several simulation studies.

Algorithms↗

Estimation of the parameters of a binary Markov random field on a graph with application to fibre type distributions in a muscle cross-section.

Methods are discussed for the estimation of the parameters of a binary Markov random field (BMRF) defined on a graph. The standard method is maximum pseudo-likelihood (MPL) estimation. Maximum likelihood (ML) estimation has been hampered in the past by the intractability of the likelihood function. Recently Markov chain Monte Carlo (MCMC) methods have been introduced for ML estimation. In this paper a new method for Monte Carlo maximum likelihood is described. It is used for the estimation of the parameters of a simple model (the Ising model of statistical physics). As an application the distribution of fibre types in a cross-section of human muscle is analysed.

Animals↗

Image recognition in the presence of non-Gaussian noise with unknown statistics.

We design receivers to detect a known pattern or a reference signal in the presence of very general and non-Gaussian types of noise. Three sources of input-noise degradation are considered: additive, multiplicative, and disjoint background. The detection process involves two steps: (1) estimation of the relevant noise parameters within the framework of hypothesis testing and (2) maximizing a certain metric that measures the likelihood of the target being at a given location. The parameter estimation portion is carried out by moment-matching techniques. Because of the number of unknown parameters and the fact that various types of input-noise processes are non-Gaussian, the methods that are used to estimate these parameters differ from the standard methods of maximizing the likelihood function. To verify the existence of the target at a certain location, we use l(p)-norm metric for p > or = 0 to measure the likelihood of the target being present at the location of interest. Computer simulations are used to show that for the images tested here, the receivers designed herein perform better than some existing receivers.

Journal Article↗

The contributions of Jerome Cornfield to the theory of statistics.

This paper is a review of the contributions of Jerome Cornfield to the theory of statistics. It discusses several highlights of his theoretical work as well as describing his philosophy relating theory to application. The three areas discussed are: linear programming, urn sampling and its generalizations to the analysis of variance, and Bayesian inference. It is not widely known that Jerome Cornfield was perhaps the first to formulate and approximately solve the linear programming problem in 1941. His formulation was made for the famous "Diet Problem". An early publication introduced the method of indicator random variables in the context of urn sampling. This simple method allowed straightforward calculations of the low order moments for estimates arising from sampling finite populations and was later generalized to the two-way analysis of variance. The application of the urn sampling model to the analysis of variance served to illuminate how one chooses proper error terms for making tests in the analysis of variance table. Jerome Cornfield's philosophy on applications of statistics was dominated by a Bayesian outlook. His theoretical contributions in the past two decades were mainly concerned with the development of Bayesian ideas and methods. A brief survey is made of his main contributions to this area. A particularly noteworthy result was his demonstration that for the two-sample slippage problem of location, the likelihood function under a permutation setting is uninformative for the slippage parameter. However, the posterior distribution differs from the prior distribution despite the fact that the likelihood is uninformative.

Bayes Theorem↗

Score test for mapping quantitative-trait loci with sibships of arbitrary size when the dominance effect is not negligible.

In the linakge analysis of quantitative traits, an additive model that assumes no dominance effect is often adopted. Intuitively, when the no-dominance-effect assumption does not hold, such a practice does not make efficient use of the data, and its power to detect linkage can be improved. Here we introduce a score statistic for detecting quantitative trait loci when the dominance effect is not neglible or the dominance effect is a concern. This statistic is derived from a normal likelihood function for sibships of arbitrary size. In the derivation, the inherent genetic constraints on model parameters are fully taken into consideration. This score statistic is asymptotically equivalent to the corresponding likelihood ratio statistic, but it is much easier to compute. The asymptotic distribution of this statistic is derived, which is a mixture of chi(0) (2), chi(1) (2), and chi(2) (2). Weights for distribution components are functions of the informativeness of the marker data. The type I error rate and the power of the proposed statistic in finite sample are evaluated via simulations.

Chromosome Mapping↗

Multivariate scan statistics for disease surveillance.

In disease surveillance, there are often many different data sets or data groupings for which we wish to do surveillance. If each data set is analysed separately rather than combined, the statistical power to detect an outbreak that is present in all data sets may suffer due to low numbers in each. On the other hand, if the data sets are added by taking the sum of the counts, then a signal that is primarily present in one data set may be hidden due to random noise in the other data sets. In this paper, we present an extension of the spatial and space-time scan statistic that simultaneously incorporates multiple data sets into a single likelihood function, so that a signal is generated whether it occurs in only one or in multiple data sets. This is done by defining the combined log likelihood as the sum of the individual log likelihoods for those data sets for which the observed case count is more than the expected. We also present another extension, where the concept of combining likelihoods from different data sets is used to adjust for covariates. Using data from the National Bioterrorism Syndromic Surveillance Demonstration Project, we illustrate the new method using physician telephone calls, regular physician visits and urgent care visits by Harvard Pilgrim Health Care members cared for by Harvard Vanguard Medical Associates, a large multi-specialty group practice in Massachusetts. For upper and lower gastrointestinal (GI) illness, there were on average 20 telephone calls, nine urgent care visits and 22 regular physician visits per day. The strongest signal was generated by a single data set and due to a familial outbreak of pinworm disease. The second and third strongest signals were generated by the combined strength of two of the three data sets.

Boston↗

The role of parametric assumptions in adaptive Bayesian estimation.

Variants of adaptive Bayesian procedures for estimating the 5% point on a psychometric function were studied by simulation. Bias and standard error were the criteria to evaluate performance. The results indicated a superiority of (a) uniform priors, (b) model likelihood functions that are odd symmetric about threshold and that have parameter values larger than their counterparts in the psychometric function, (c) stimulus placement at the prior mean, and (d) estimates defined as the posterior mean. Unbiasedness arises in only 10 trials, and 20 trials ensure constant standard errors. The standard error of the estimates equals 0.617 times the inverse of the square root of the number of trials. Other variants yielded bias and larger standard errors.

Bayes Theorem↗

Identifying space-time disease clusters.

A cluster of cases of disease that are close both in space and in time is suggestive of an infectious aetiology. We present statistical tests for space-time clusters of disease for the two situations where the population at risk is either known or unknown as a function of space and time. The tests are derived using standard statistical methodology from a simple mathematical model of disease spread, i.e. they are derived as score tests from a likelihood function in which the infection process is modelled as a point process whose intensity becomes greater near an infector. A problem for such tests is that, when investigating whether or not a disease may be of infectious origin, the space and time distances characterising closeness to an infection are very likely to be unknown. The proposed methodology copes with this difficulty in a statistically acceptable way, without requiring multiple tests whose interpretation would be doubtful. When the underlying population size is unknown, the test reduces to a modification of the Knox test. An example of its use is given as epidemiology, risk, space-time cluster, likelihood and Knox test.

Animals↗

Efficient Huber-Markov edge-preserving image restoration.

The regularization of the least-squares criterion is an effective approach in image restoration to reduce noise amplification. To avoid the smoothing of edges, edge-preserving regularization using a Gaussian Markov random field (GMRF) model is often used to allow realistic edge modeling and provide stable maximum a posteriori (MAP) solutions. However, this approach is computationally demanding because the introduction of a non-Gaussian image prior makes the restoration problem shift-variant. In this case, a direct solution using fast Fourier transforms (FFTs) is not possible, even when the blurring is shift-invariant. We consider a class of edge-preserving GMRF functions that are convex and have nonquadratic regions that impose less smoothing on edges. We propose a decomposition-enabled edge-preserving image restoration algorithm for maximizing the likelihood function. By decomposing the problem into two subproblems, with one shift-invariant and the other shift-variant, our algorithm exploits the sparsity of edges to define an FFT-based iteration that requires few iterations and is guaranteed to converge to the MAP estimate.

Algorithms↗

A random-effects model for analysis of infectious disease final-state data.

Ball (1986, Advances in Applied Probability 18, 289-310) presented an extension to the "General Epidemic Model" in which an individual's (random) infectious period could have any distribution whose Laplace transform could be specified. This paper describes the fitting of Ball's model to data on the final state of infection within households, and gives an intuitive mathematical derivation of the corresponding likelihood function. We extend the model in several ways, including an extension to allow for random-effects heterogeneity in disease transmission rate between individuals. We give an algorithm for the efficient numerical computation of maximum likelihood estimators of the transmission rates, and describe the assessment of goodness of model fit. The methodology is illustrated with recent survey data on outbreaks of Shigella sonnei in 102 households in Manchester, UK. The results are consistent with previous anecdotal evidence of the infectiousness and susceptibility of individuals within households as a function of age and sex.

Adult↗

On the asymptotic behavior of the estimate of the recombination fraction under the null hypothesis of no linkage when the model is misspecified.

We show that under the null hypothesis of no linkage the maximum likelihood estimator of the recombination fraction converges to 1/2 even when the trait-related parameter values in the likelihood function are misspecified. Furthermore, we show that under the null hypothesis of no linkage, but with misspecified trait-related parameter values, the negative of twice the natural logarithm of the likelihood ratio statistic still has a limiting chi-square distribution with 1 degree of freedom.

Chi-Square Distribution↗

Modeling association among demographic parameters in analysis of open population capture-recapture data.

We present a hierarchical extension of the Cormack-Jolly-Seber (CJS) model for open population capture-recapture data. In addition to recaptures of marked animals, we model first captures of animals and losses on capture. The parameter set includes capture probabilities, survival rates, and birth rates. The survival rates and birth rates are treated as a random sample from a bivariate distribution, thus the model explicitly incorporates correlation in these demographic rates. A key feature of the model is that the likelihood function, which includes a CJS model factor, is expressed entirely in terms of identifiable parameters; losses on capture can be factored out of the model. Since the computational complexity of classical likelihood methods is prohibitive, we use Markov chain Monte Carlo in a Bayesian analysis. We describe an efficient candidate-generation scheme for Metropolis-Hastings sampling of CJS models and extensions. The procedure is illustrated using mark-recapture data for the moth Gonodontis bidentata.

Animals↗

Physicians and medical innovation.

Previous attempts to model some aspects of physician behaviour include those of Evans, Sloan and Feldman and Wolfson. It is suggested that the introduction of knowledge as a distinct element in a microeconomic model of physician behaviour is preferable to the inclusion of a variable called 'discretionary influence' or 'quality of care' in the physician's utility function. This is because the properties of functions containing either of these variables appear to be indeterminate. By comparison the properties of the knowledge constraints can be specified with some confidence. The factors affecting a physician's demand for treatment on behalf of patients are identified as (1) the physician's objective function, (2) his knowledge and (3) the availability of medical resources. Furthermore, the knowledge element can be sub-divided into two parts: the set of prior probabilities and the set of likelihood functions. The former may be identified with the physician's local knowledge, whereas the latter may be associated with the physician's medical training. A significant fraction of the growing demand for hospital care has been attributed to changes in medical technology. During the late fifties and afterwards 'more cases became treatable' and physicians, it is argued, cannot resist the 'technological imperative'. The paper shows that the model may be used to generate testable hypothesis regarding the adoption by physicians of both process and product innovations. The discussion of the physician's medical knowledge is fundamental to the inducement mechanism. The policy instruments available to achieve an optimal diffusion of innovations are reviewed.

Communication↗

A result on a 2 x 2 survival experiment.

Lifetime data classified according to categorical variables under the proportionality of the hazard functions of response variables for various treatment combinations is assumed. The proposed model is a combination of Cox's proportional hazards model and ANOVA model. The existence of a solution to the marginal likelihood function is examined for the case of 2 x 2 two-way classification. We provide an easily verifiable condition for the existence of a unique estimate.

Humans↗

Short communication: Optimal random regression models for milk production in dairy cattle.

Legendre polynomials of orders 3 to 8 in random regression models (RRM) for first-lactation milk production in Canadian Holsteins were compared statistically to determine the best model. Twenty-six RRM were compared using LP of order 5 for the phenotypic age-season groupings. Variance components of RRM were estimated using Bayesian estimation via Gibbs sampling. Several statistical criteria for model comparison were used including the total residual variance, the log likelihood function, Akaike's information criterion, the Bayesian information criterion, Bayes factors, an information-theoretic measure of model complexity, and the percentage relative reduction in complexity. The residual variance always picks the model with the most parameters. The log likelihood and information-theoretic measure picked the model with order 5 for additive genetic effects and order 7 for permanent environmental effects. The currently used model in Canada (order 5 for both additive and permanent environmental effects) was not the best for any single criterion, but was optimal when considering all criteria.

Animals↗

Models of varying parametric form in case-referent studies.

The analysis of case-referent data can be based on a consideration of case and referent counts in various exposure categories as the realization of a set of binomial processes. After appropriate modeling of the binomial parameters, a joint likelihood function can be formed and maximized to obtain estimates of the parameters constituting the model elements. The procedure has been applied to the problem of additive and multiplicative models of disease incidence rates, as encountered in case-referent studies. Likelihood ratios can be used to compare models with equal numbers of parameters. These ratios do not lead to significance tests, but to estimate of the relative degree of corroboration of different hypotheses by the data at hand.

Body Weight↗