PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Likelihood Functions”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11Linked to original sources

Reconstruction of extended cortical sources for EEG and MEG based on a Monte-Carlo-Markov-chain estimator.

A new procedure to model extended cortical sources from EEG and MEG recordings based on a probabilistic approach is presented. The method (SPMECS) was implemented within the framework of maximum likelihood estimators. Neuronal activity generating EEG or MEG signals was characterized by the number of sources and their location and extension. Based on the noise distribution of the measured data, source configurations were associated with the according value of the likelihood function. To find the most likely source, i.e., the maximum likelihood estimator, and its level of confidence, a stochastic solver (Metropolis algorithm) was applied. The method presented supports the incorporation of virtually any constraint, e.g., based on physiological and anatomical a priori knowledge. Thus, ambiguity of the ill-posed inverse problem was reduced considerably by confining sources to the cortical surface extracted from individual MR images. The influence of different levels and types of noise on the outcome was investigated by means of simulations. Somatosensory evoked magnetic fields analyzed by the method presented suggest that larger extended cortical areas are involved in the processing of combined finger stimulation as compared to single finger stimulation.

Algorithms↗

Comparisons of test statistics arising from marginal analyses of multivariate survival data.

We investigate the properties of several statistical tests for comparing treatment groups with respect to multivariate survival data, based on the marginal analysis approach introduced by Wei, Lin and Weissfeld ["Regression Analysis of multivariate incomplete failure time data by modelling marginal distributions," JASA vol. 84 pp. 1065-1073]. We consider two types of directional tests, based on a constrained maximization and on linear combinations of the unconstrained maximizer of the working likelihood function, and the omnibus test arising from the same working likelihood. The directional tests are members of a larger class of tests, from which an asymptotically optimal test can be found. We compare the asymptotic powers of the tests under general contiguous alternatives for a variety of settings, and also consider the choice of the number of survival times to include in the multivariate outcome. We illustrate the results with simulations and with the results from a clinical trial examining recurring opportunistic infections in persons with HIV.

AIDS-Related Opportunistic Infections↗

A zero-inflated Poisson mixed model to analyze diagnosis related groups with majority of same-day hospital stays.

With increasing trend of same-day procedures and operations performed for hospital admissions, it is important to analyze those Diagnosis Related Groups (DRGs) consisting of mainly same-day separations. A zero-inflated Poisson (ZIP) mixed model is presented to identify health- and patient-related characteristics associated with length of stay (LOS) and to model variations in LOS within such DRGs. Random effects are introduced to account for inter-hospital variations and the dependence of clustered LOS observations via the generalized linear mixed models (GLMM) approach. Parameter estimation is achieved by maximizing an appropriate log-likelihood function using the EM algorithm to obtain approximate residual maximum likelihood (REML) estimates. An S-Plus macro is developed to provide a unified ZIP modeling approach. The determination of pertinent factors would benefit hospital administrators and clinicians to manage LOS and expenditures efficiently.

Algorithms↗

A Bayesian approach to joint feature selection and classifier design.

This paper adopts a Bayesian approach to simultaneously learn both an optimal nonlinear classifier and a subset of predictor variables (or features) that are most relevant to the classification task. The approach uses heavy-tailed priors to promote sparsity in the utilization of both basis functions and features; these priors act as regularizers for the likelihood function that rewards good classification on the training data. We derive an expectation-maximization (EM) algorithm to efficiently compute a maximum a posteriori (MAP) point estimate of the various parameters. The algorithm is an extension of recent state-of-the-art sparse Bayesian classifiers, which in turn can be seen as Bayesian counterparts of support vector machines. Experimental comparisons using kernel classifiers demonstrate both parsimonious feature selection and excellent classification accuracy on a range of synthetic and benchmark data sets.

Algorithms↗

Finding regulatory elements using joint likelihoods for sequence and expression profile data.

A recent, popular method of finding promoter sequences is to look for conserved motifs upstream of genes clustered on the basis of expression data. This method presupposes that the clustering is correct. Theoretically, one should be better able to find promoter sequences and create more relevant gene clusters by taking a unified approach to these two problems. We present a likelihood function for a "sequence-expression" model giving a joint likelihood for a promoter sequence and its corresponding expression levels. An algorithm to estimate sequence-expression model parameters using Gibbs sampling and Expectation/Maximization is described. A program, called kimono, that implements this algorithm has been developed: the source code is freely available on the Internet.

Algorithms↗

Variance-Components QTL linkage analysis of selected and non-normal samples: conditioning on trait values.

Standard variance-components quantitative trait loci (QTL) linkage analysis can produce an elevated rate of type 1 errors when applied to selected samples and non-normal data. Here we describe an adjustment of the log-likelihood function based on conditioning on trait values. This leads to a likelihood ratio test that is valid in selected samples and non-normal data, and equal in power to alternative methods for analyzing selected samples that require knowledge of the ascertainment procedure or the trait values of non-selected individuals.

Analysis of Variance↗

The statistical analysis of truncated data: application to the Sverdlovsk anthrax outbreak.

An outbreak of anthrax occurred in the city of Sverdlovsk in Russia in the spring of 1979. The outbreak was due to the inhalation of spores that were accidentally released from a military microbiology facility. In response to the outbreak a public health intervention was mounted that included distribution of antibiotics and vaccine. The objective of this paper is to develop and apply statistical methodology to analyse the Sverdlovsk outbreak, and in particular to estimate the incubation period of inhalational anthrax and the number of deaths that may have been prevented by the public health intervention. The data available for analysis from this common source epidemic are the incubation periods of reported deaths. The statistical problem is that incubation periods are truncated because some individuals may have had their deaths prevented by the public health interventions and thus are not included in the data. However, it is not known how many persons received the intervention or how efficacious was the intervention. A likelihood function is formulated that accounts for the effects of truncation. The likelihood is decomposed into a binomial likelihood with unknown sample size and a conditional likelihood for the incubation periods. The methods are extended to allow for a phase-in of the intervention over time. Assuming a lognormal model for the incubation period distribution, the median and mean incubation periods were estimated to be 11.0 and 14.2 days respectively. These estimates are longer than have been previously reported in the literature. The death toll from the Sverdlovsk anthrax outbreak could have been about 14% larger had there not been a public health intervention; however, the confidence intervals are wide (95% CI 0-61%). The sensitivity of the results to model assumptions and the parametric model for the incubation period distribution are investigated. The results are useful for determining how long antibiotic therapy should be continued in suspected anthrax cases and also for estimating the ultimate number of deaths in a new outbreak in the absence of any public health interventions.

Journal Article↗

Parameter estimation in a Gompertzian stochastic model for tumor growth.

The problem of estimating parameters in the drift coefficient when a diffusion process is observed continuously requires some specific assumptions. In this paper, we consider a stochastic version of the Gompertzian model that describes in vivo tumor growth and its sensitivity to treatment with antiangiogenic drugs. An explicit likelihood function is obtained, and we discuss some properties of the maximum likelihood estimator for the intrinsic growth rate of the stochastic Gompertzian model. Furthermore, we show some simulation results on the behavior of the corresponding discrete estimator. Finally, an application is given to illustrate the estimate of the model parameters using real data.

Angiogenesis Inhibitors↗

A general likelihood approach to trait-based multipoint linkage analysis in large groups of half-sibs and super sisters.

The idea of trait-based linkage analysis in half-sibs is extended by comparing the frequency of parental marker haplotypes in animals with different phenotypes. This article first presents the likelihood of observing different classes of paternal haplotypes in a half-sib family, where only family members of a certain phenotype (e.g., affected) are genotyped and are fully informative. The likelihood function is then generalized to multiple phenotypic categories. A linear predictor allows for discontinuous as well as for continuous phenotypes and other explanatory variables. Finally, how to incorporate not fully informative offspring and how to analyze super sister families are shown. Maximum-likelihood estimates of all parameters can be found by a Newton-Raphson algorithm, which mimics an iteratively weighted least-squares procedure. The method allows for any multilocus feasible mapping function and, among others, for situations with selective or nonselective genotyping, single or multiple traits, and continuous or categorical traits. No parameters are required to describe the mode of inheritance and the method copes with virtually any family size. Fields of applications are therefore mapping experiments in species with a high reproductive capacity, such as cattle, pigs, horses, honey bees, trees, and fish.

Animals↗

Estimation in regression models for longitudinal binary data with outcome-dependent follow-up.

In many observational studies, individuals are measured repeatedly over time, although not necessarily at a set of pre-specified occasions. Instead, individuals may be measured at irregular intervals, with those having a history of poorer health outcomes being measured with somewhat greater frequency and regularity. In this paper, we consider likelihood-based estimation of the regression parameters in marginal models for longitudinal binary data when the follow-up times are not fixed by design, but can depend on previous outcomes. In particular, we consider assumptions regarding the follow-up time process that result in the likelihood function separating into two components: one for the follow-up time process, the other for the outcome measurement process. The practical implication of this separation is that the follow-up time process can be ignored when making likelihood-based inferences about the marginal regression model parameters. That is, maximum likelihood (ML) estimation of the regression parameters relating the probability of success at a given time to covariates does not require that a model for the distribution of follow-up times be specified. However, to obtain consistent parameter estimates, the multinomial distribution for the vector of repeated binary outcomes must be correctly specified. In general, ML estimation requires specification of all higher-order moments and the likelihood for a marginal model can be intractable except in cases where the number of repeated measurements is relatively small. To circumvent these difficulties, we propose a pseudolikelihood for estimation of the marginal model parameters. The pseudolikelihood uses a linear approximation for the conditional distribution of the response at any occasion, given the history of previous responses. The appeal of this approximation is that the conditional distributions are functions of the first two moments of the binary responses only. When the follow-up times depend only on the previous outcome, the pseudolikelihood requires correct specification of the conditional distribution of the current outcome given the outcome at the previous occasion only. Results from a simulation study and a study of asymptotic bias are presented. Finally, we illustrate the main results using data from a longitudinal observational study that explored the cardiotoxic effects of doxorubicin chemotherapy for the treatment of acute lymphoblastic leukemia in children.

Adolescent↗

Testing hypotheses in case-control studies--equivalence of Mantel-Haenszel statistics and logit score tests.

The two approaches in common use for the analysis of case-control studies are cross-classification by confounding variables, and modeling the logarithm of the odds ratio as a function of exposure and confounding variables. We show here that score statistics derived from the likelihood function in the latter approach are identical to the Mantel-Haenszel test statistics appropriate for the former approach. This identity holds in the most general situation considered, testing for marginal homogeneity in mK tables. This equivalence is demonstrated by a permutational argument which leads to a general likelihood expression in which the exposure variable may be a vector of discrete and/or continuous variables and in which more than two comparison groups may be considered. This likelihood can be used in analyzing studies in which there are multiple controls for each case or in which several disease categories are being compared. The possibility of including continuous variables makes this likelihood useful in situations that cannot be treated using the Mantel-Haenszel cross-classification approach.

Epidemiologic Methods↗

Estimating haplotype-disease associations with pooled genotype data.

The genetic dissection of complex human diseases requires large-scale association studies which explore the population associations between genetic variants and disease phenotypes. DNA pooling can substantially reduce the cost of genotyping assays in these studies, and thus enables one to examine a large number of genetic variants on a large number of subjects. The availability of pooled genotype data instead of individual data poses considerable challenges in the statistical inference, especially in the haplotype-based analysis because of increased phase uncertainty. Here we present a general likelihood-based approach to making inferences about haplotype-disease associations based on possibly pooled DNA data. We consider cohort and case-control studies of unrelated subjects, and allow arbitrary and unequal pool sizes. The phenotype can be discrete or continuous, univariate or multivariate. The effects of haplotypes on disease phenotypes are formulated through flexible regression models, which allow a variety of genetic hypotheses and gene-environment interactions. We construct appropriate likelihood functions for various designs and phenotypes, accommodating Hardy-Weinberg disequilibrium. The corresponding maximum likelihood estimators are approximately unbiased, normally distributed, and statistically efficient. We develop simple and efficient numerical algorithms for calculating the maximum likelihood estimators and their variances, and implement these algorithms in a freely available computer program. We assess the performance of the proposed methods through simulation studies, and provide an application to the Finland-United States Investigation of NIDDM Genetics Study. The results show that DNA pooling is highly efficient in studying haplotype-disease associations. As a by-product, this work provides valid and efficient methods for estimating haplotype-disease associations with unpooled DNA samples.

Algorithms↗

Parallel simulated annealing for emission tomography.

A method for implementing simulated annealing in parallel to speed up the execution of emission tomography (ET) image reconstruction is presented. A high degree of parallelism can be attained by using a parallel-acceptance partitioning strategy, in which perturbations to subsets of the estimate are evaluated in parallel. However because the point spread function in ET imaging systems is globally dependent, processors cannot update the current estimate independently. Consequently, processors must be synchronized each time a perturbation is accepted to avoid introducing error. This can produce excessive communications overhead, especially when the acceptance rate is high. In this paper an energy function is constructed to reduce the synchronization requirements by using a reformulation of the log-likelihood function from the expectation maximization (EM) algorithm. The approach is to change the global dependence in the energy function from the current estimate to the estimate generated during the last iteration. The synchronization requirements for guaranteed convergence are then significantly reduced from once per acceptance to once per iteration. This parallel implementation on 54 Inmos T800 transputers connected in a ring topology resulted in execution times that were almost 50 times faster than on a VAX 8600.

Algorithms↗

Semiparametric estimation of marginal hazard function from case-control family studies.

Estimating marginal hazard function from the correlated failure time data arising from case-control family studies is complicated by noncohort study design and risk heterogeneity due to unmeasured, shared risk factors among the family members. Accounting for both factors in this article, we propose a two-stage estimation procedure. At the first stage, we estimate the dependence parameter in the distribution for the risk heterogeneity without obtaining the marginal distribution first or simultaneously. Assuming that the dependence parameter is known, at the second stage we estimate the marginal hazard function by iterating between estimation of the risk heterogeneity (frailty) for each family and maximization of the partial likelihood function with an offset to account for the risk heterogeneity. We also propose an iterative procedure to improve the efficiency of the dependence parameter estimate. The simulation study shows that both methods perform well under finite sample sizes. We illustrate the method with a case-control family study of early onset breast cancer.

Adult↗

Emission image reconstruction for randoms-precorrected PET allowing negative sinogram values.

Most positron emission tomography (PET) emission scans are corrected for accidental coincidence (AC) events by real-time subtraction of delayed-window coincidences, leaving only the randoms-precorrected data available for image reconstruction. The real-time randoms precorrection compensates in mean for AC events but destroys the Poisson statistics. The exact log-likelihood for randoms-precorrected data is inconvenient, so practical approximations are needed for maximum likelihood or penalized-likelihood image reconstruction. Conventional approximations involve setting negative sinogram values to zero, which can induce positive systematic biases, particularly for scans with low counts per ray. We propose new likelihood approximations that allow negative sinogram values without requiring zero-thresholding. With negative sinogram values, the log-likelihood functions can be nonconcave, complicating maximization; nevertheless, we develop monotonic algorithms for the new models by modifying the separable paraboloidal surrogates and the maximum-likelihood expectation-maximization (ML-EM) methods. These algorithms ascend to local maximizers of the objective function. Analysis and simulation results show that the new shifted Poisson (SP) model is nearly free of systematic bias yet keeps low variance. Despite its simpler implementation, the new SP performs comparably to the saddle-point model which has shown the best performance (as to systematic bias and variance) in randoms-precorrected PET emission reconstruction.

Algorithms↗

Maximum likelihood for genome phylogeny on gene content.

With the rapid growth of entire genome data, reconstructing the phylogenetic relationship among different genomes has become a hot topic in comparative genomics. Maximum likelihood approach is one of the various approaches, and has been very successful. However, there is no reported study for any applications in the genome tree-making mainly due to the lack of an analytical form of a probability model and/or the complicated calculation burden. In this paper we studied the mathematical structure of the stochastic model of genome evolution, and then developed a simplified likelihood function for observing a specific phylogenetic pattern under four genome situation using gene content information. We use the maximum likelihood approach to identify phylogenetic trees. Simulation results indicate that the proposed method works well and can identify trees with a high correction rate. Real data application provides satisfied results. The approach developed in this paper can serve as the basis for reconstructing phylogenies of more than four genomes.

Journal Article↗

Parametric inference for epidemic models.

The likelihood function corresponding to epidemic data is often very complicated. We illustrate that the EM algorithm can sometimes help to simplify likelihood inferences. Difficulties with likelihood inferences about parameters of epidemic models have established a role for martingale methods. These are methods of statistical inference based on estimating equations derived from the rich theory of martingales, and they have produced simple methods of inference in a number of important applications to epidemic data. We contrast likelihood methods with martingale methods and determine which specific assumptions cause changes in inferences about the infection potential of a disease. It is found that the martingale-based estimate of the infection potential remains unaltered under a variety of commonly used model specifications but that the precision of this estimate changes as model assumptions are altered.

Algorithms↗

A likelihood-based approach to capture-recapture estimation of demographic parameters under the robust design.

The Jolly-Seber method has been the traditional approach to the estimation of demographic parameters in long-term capture-recapture studies of wildlife and fish species. This method involves restrictive assumptions about capture probabilities that can lead to biased estimates, especially of population size and recruitment. Pollock (1982, Journal of Wildlife Management 46, 752-757) proposed a sampling scheme in which a series of closely spaced samples were separated by longer intervals such as a year. For this "robust design," Pollock suggested a flexible ad hoc approach that combines the Jolly-Seber estimators with closed population estimators, to reduce bias caused by unequal catchability, and to provide estimates for parameters that are unidentifiable by the Jolly-Seber method alone. In this paper we provide a formal modelling framework for analysis of data obtained using the robust design. We develop likelihood functions for the complete data structure under a variety of models and examine the relationship among the models. We compute maximum likelihood estimates for the parameters by applying a conditional argument, and compare their performance against those of ad hoc and Jolly-Seber approaches using simulation.

Animals↗