PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “penalized likelihood maximization”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Gene networks inference using dynamic Bayesian networks.

This article deals with the identification of gene regulatory networks from experimental data using a statistical machine learning approach. A stochastic model of gene interactions capable of handling missing variables is proposed. It can be described as a dynamic Bayesian network particularly well suited to tackle the stochastic nature of gene regulation and gene expression measurement. Parameters of the model are learned through a penalized likelihood maximization implemented through an extended version of EM algorithm. Our approach is tested against experimental data relative to the S.O.S. DNA Repair network of the Escherichia coli bacterium. It appears to be able to extract the main regulations between the genes involved in this network. An added missing variable is found to model the main protein of the network. Good prediction abilities on unlearned data are observed. These first results are very promising: they show the power of the learning algorithm and the ability of the model to capture gene interactions.

Algorithms↗

A novel method with improved power to detect recombination hotspots from polymorphism data reveals multiple hotspots in human genes.

We introduce a new method for detection of recombination hotspots from population genetic data. This method is based on (a) defining an (approximate) penalized likelihood for how recombination rate varies with physical position and (b) maximizing this penalized likelihood over possible sets of recombination hotspots. Simulation results suggest that this is a more powerful method for detection of hotspots than are existing methods. We apply the method to data from 89 genes sequenced in African American and European American populations. We find many genes with multiple hotspots, and some hotspots show evidence of being population-specific. Our results suggest that hotspots are randomly positioned within genes and could be as frequent as one per 30 kb.

Black People↗

PHMPL: a computer program for hazard estimation using a penalized likelihood method with interval-censored and left-truncated data.

The Cox model is the model of choice when analyzing right-censored and possibly left-truncated survival data. The present paper proposes a program to estimate the hazard function in a proportional hazards model and also to treat more complex observation schemes involving general censored and left-truncated data. The hazard function estimator is defined non-parametrically as the function which maximizes a penalized likelihood, and the solution is approximated using splines. The smoothing parameter is chosen using approximate cross-validation. Confidence bands for the estimator are given. As an illustration, the age-specific incidence of dementia is estimated and one of its risk factors is studied.

Age Factors↗

A penalized likelihood approach for an illness-death model with interval-censored data: application to age-specific incidence of dementia.

We consider the problem of estimating the intensity functions for a continuous time 'illness-death' model with intermittently observed data. In such a case, it may happen that a subject becomes diseased between two visits and dies without being observed. Consequently, there is an uncertainty about the precise number of transitions. Estimating the intensity of transition from health to illness by survival analysis (treating death as censoring) is biased downwards. Furthermore, the dates of transitions between states are not known exactly. We propose to estimate the intensity functions by maximizing a penalized likelihood. The method yields smooth estimates without parametric assumptions. This is illustrated using data from a large cohort study on cerebral ageing. The age-specific incidence of dementia is estimated using an illness-death approach and a survival approach.

Journal Article↗

Penalized likelihood optimization for censored missing value imputation in proteomics.

Label-free bottom-up proteomics using mass spectrometry and liquid chromatography has long been established as one of the most popular high-throughput analysis workflows for proteome characterization. However, it produces data hindered by complex and heterogeneous missing values, which imputation has long remained problematic. To cope with this, we introduce Pirat, an algorithm that harnesses this challenge using an original likelihood maximization strategy. Notably, it models the instrument limit by learning a global censoring mechanism from the data available. Moreover, it estimates the covariance matrix between enzymatic cleavage products (ie peptides or precursor ions), while offering a natural way to integrate complementary transcriptomic information when multi-omic assays are available. Our benchmarking on several datasets covering a variety of experimental designs (number of samples, acquisition mode, missingness patterns, etc.) and using a variety of metrics (differential analysis ground truth or imputation errors) shows that Pirat outperforms all pre-existing imputation methods. Beyond the interest of Pirat as an imputation tool, these results pinpoint the need for a paradigm change in proteomics imputation, as most pre-existing strategies could be boosted by incorporating similar models to account for the instrument censorship or for the correlation structures, either grounded to the analytical pipeline or arising from a multi-omic approach.

Proteomics↗

Warping two-dimensional electrophoresis gel images to correct for geometric distortions of the spot pattern.

A crucial step in two-dimensional gel based protein expression analysis is to match spots in different gel images that correspond to the same protein. It still requires extensive and time-consuming manual interference, although several semiautomatic techniques exist. Geometric distortion of the protein patterns inherent to the electrophoresis procedure is one of the main causes of these difficulties. An image warping method to reduce this problem is presented. A warping is a function that deforms images by mapping between image domains. The method proceeds in two steps. Firstly, a simple physicochemical model is formulated and applied for warping of each gel image to correct for what might be one of the main causes of the distortions: current leakage across the sides during the second-dimensional electrophoresis. Secondly, the images are automatically aligned by maximizing a penalized likelihood criterion. The method is applied to a set of ten gel images showing the radioactively labeled proteome of yeast Saccharomyces cerevisiae during normal and steady-state saline growth. The improvement in matching when given the warped images instead of the original ones is exemplified by a comparison within a commercially available software.

Algorithms↗

Smooth centile curves for skew and kurtotic data modelled using the Box-Cox power exponential distribution.

The Box-Cox power exponential (BCPE) distribution, developed in this paper, provides a model for a dependent variable Y exhibiting both skewness and kurtosis (leptokurtosis or platykurtosis). The distribution is defined by a power transformation Y(nu) having a shifted and scaled (truncated) standard power exponential distribution with parameter tau. The distribution has four parameters and is denoted BCPE (mu,sigma,nu,tau). The parameters, mu, sigma, nu and tau, may be interpreted as relating to location (median), scale (approximate coefficient of variation), skewness (transformation to symmetry) and kurtosis (power exponential parameter), respectively. Smooth centile curves are obtained by modelling each of the four parameters of the distribution as a smooth non-parametric function of an explanatory variable. A Fisher scoring algorithm is used to fit the non-parametric model by maximizing a penalized likelihood. The first and expected second and cross derivatives of the likelihood, with respect to mu, sigma, nu and tau, required for the algorithm, are provided. The centiles of the BCPE distribution are easy to calculate, so it is highly suited to centile estimation. This application of the BCPE distribution to smooth centile estimation provides a generalization of the LMS method of the centile estimation to data exhibiting kurtosis (as well as skewness) different from that of a normal distribution and is named here the LMSP method of centile estimation. The LMSP method of centile estimation is applied to modelling the body mass index of Dutch males against age.

Adolescent↗

Hazard regression for interval-censored data with penalized spline.

This article introduces a new approach for estimating the hazard function for possibly interval- and right-censored survival data. We weakly parameterize the log-hazard function with a piecewise-linear spline and provide a smoothed estimate of the hazard function by maximizing the penalized likelihood through a mixed model-based approach. We also provide a method to estimate the amount of smoothing from the data. We illustrate our approach with two well-known interval-censored data sets. Extensive numerical studies are conducted to evaluate the efficacy of the new procedure.

Biometry↗

Uses of the EM algorithm in the analysis of data on HIV/AIDS and other infectious diseases.

The analysis of data on infectious diseases is a natural setting for applications of the EM algorithm, because the infection process is only partially observable. Difficulties in determining the expectation at the E step have been side-stepped by adopting pragmatic models which reflect only part of the mechanism that generates the data. In the HIV/AIDS context the EM algorithm has helped in the reconstruction of the unobserved HIV infection curve, the so-called backprojection problem, as well as in the estimation of the distribution for the incubation period until AIDS, in estimating the infectivity of HIV in partnerships and in estimating parameters describing the decline in the immune system. There is a need for smooth estimates of functions in these applications, suggesting the use of the EMS algorithm or use of the EM algorithm to maximize a penalized likelihood. For data on other infectious diseases the application of the EM algorithm has so far been restricted to analyses of data on the size of outbreaks in a sample of households.

Acquired Immunodeficiency Syndrome↗

Penalized-likelihood sinogram smoothing for low-dose CT.

We have developed a sinogram smoothing approach for low-dose computed tomography (CT) that seeks to estimate the line integrals needed for reconstruction from the noisy measurements by maximizing a penalized-likelihood objective function. The maximization is performed by an algorithm derived by use of the separable paraboloidal surrogates framework. The approach overcomes some of the computational limitations of a previously proposed spline-based penalized-likelihood sinogram smoothing approach, and it is found to yield better resolution-variance tradeoffs than this spline-based approach as well an existing adaptive filtering approach. Such sinogram smoothing approaches could be valuable when applied to the low-dose data acquired in CT screening exams, such as those being considered for lung-nodule detection.

Algorithms↗

Penalized likelihood in Cox regression.

In a Cox regression model, instability of the estimated regression coefficients can be reduced by maximizing a penalized partial log-likelihood, where a penalty function of the regression coefficients is substracted from the partial log-likelihood. In this paper, we choose the optimal weight of the penalty function by maximizing the predictive value of the model, as measured by the crossvalidated partial log-likelihood. Our methods are illustrated by a study of ovarian cancer survival and by a study of centre-effects in kidney graft survival.

Clinical Trials as Topic↗

Penalized likelihood approach to estimate a smooth mean curve on longitudinal data.

This paper aims to propose a penalized likelihood approach to estimate a smooth mean curve for the evolution with time of a Gaussian variable taking into account the correlation structure of longitudinal data. The model is an extension of the mixed effects linear model including an unspecified function of time f(t). The estimator (circumflex)f(t) is defined as the solution of the maximization of the penalized likelihood and is approximated on a basis of cubic M-spline with a reduced number of knots. We present modifications of four criteria (cross-validation, generalized cross-validation, T of Rice, Akaike's criterion) to estimate the smoothing parameter when data are correlated; these four criteria gave very similar results in the simulation study. The simulation study showed also the superiority of the Bayesian confidence bands of the mean curve over the frequentist ones. We develop empirical Bayes estimates of subject-specific deviations. This approach was applied to study the progression of CD4+ lymphocyte counts in a cohort of HIV patients treated with protease inhibitors.

Acquired Immunodeficiency Syndrome↗

A method for assessing age-time disease incidence using serial prevalence data.

This paper considers nonparametric estimation of age- and time-specific trends in disease incidence using serial prevalence data collected from multiple cross-sectional samples of a population over time. The methodology accounts for differential selection of diseased and undiseased individuals resulting, for example, from differences in mortality. It is shown that when a log-linear incidence odds model is adopted, an EM algorithm provides a convenient method for carrying out maximum likelihood estimation, primarily using existing generalized linear models software. The procedure is quite general, allowing a range of age-time incidence models to be fitted under the same framework. Furthermore, by making use of existing software for fitting generalized additive models, the procedure can be generalized with virtually no extra complexity to allow maximization of a penalized likelihood for smooth nonparametric estimation. Automatic choice of smoothing level for the penalized likelihood estimates is discussed, using generalized cross-validation. The method is applied to a data set on serial toxoplasmosis prevalence, which has previously been analyzed under the assumption of nondifferential selection. A variety of age-time incidence models are fitted, and the sensitivity to plausible differential selection patterns is considered. It is found that nonmultiplicative models are unnecessary and that qualitative incidence trends are fairly robust to differential selection.

Algorithms↗

Empirical Bayes versus fully Bayesian analysis of geographical variation in disease risk.

This paper reviews methods for mapping geographical variation in disease incidence and mortality. Recent results in Bayesian hierarchical modelling of relative risk are discussed. Two approaches to relative risk estimation, along with the related computational procedures, are described and compared. The first is an empirical Bayes approach that uses a technique of penalized log-likelihood maximization; the second approach is fully Bayesian, and uses an innovative stochastic simulation technique called the Gibbs sampler. We chose to map geographical variation in breast cancer and Hodgkin's disease mortality as observed in all the health care districts of Sardinia, to illustrate relevant problems, methods and techniques.

Bayes Theorem↗

Smooth random effects distribution in a linear mixed model.

A linear mixed model with a smooth random effects density is proposed. A similar approach to P-spline smoothing of Eilers and Marx (1996, Statistical Science 11, 89-121) is applied to yield a more flexible estimate of the random effects density. Our approach differs from theirs in that the B-spline basis functions are replaced by approximating Gaussian densities. Fitting the model involves maximizing a penalized marginal likelihood. The best penalty parameters minimize Akaike's Information Criterion employing Gray's (1992, Journal of the American Statistical Association 87, 942-951) results. Although our method is applicable to any dimensions of the random effects structure, in this article the two-dimensional case is explored. Our methodology is conceptually simple, and it is relatively easy to fit in practice and is applied to the cholesterol data first analyzed by Zhang and Davidian (2001, Biometrics 57, 795-802). A simulation study shows that our approach yields almost unbiased estimates of the regression and the smoothing parameters in small sample settings. Consistency of the estimates is shown in a particular case.

Biometry↗

A competing risks analysis of presenting AIDS diagnoses trends.

The proportions of gay men presenting with various AIDS diagnoses display temporal trends. In particular, the proportion of initial diagnoses reported as Kaposi's sarcoma (KS) has declined over time. Epidemiologists have hypothesized that (a) KS may require a cofactor, whose prevalence has declined over time, or (b) KS may have a shorter incubation period than other presenting diagnoses. We examine whether this latter hypothesis, considered in a competing risks framework, could account for the observed decline in KS. We nonparametrically estimate the relevant cause-specific hazard functions from the doubly-censored data of the San Francisco City Clinic Cohort by maximizing a roughness penalized likelihood using an EM algorithm. These estimates suggest that differences in the underlying cause-specific hazard functions account for a substantial portion of the observed diagnoses trends.

Acquired Immunodeficiency Syndrome↗

Relaxed ordered-subset algorithm for penalized-likelihood image restoration.

The expectation-maximization (EM) algorithm for maximum-likelihood image recovery is guaranteed to converge, but it converges slowly. Its ordered-subset version (OS-EM) is used widely in tomographic image reconstruction because of its order-of-magnitude acceleration compared with the EM algorithm, but it does not guarantee convergence. Recently the ordered-subset, separable-paraboloidal-surrogate (OS-SPS) algorithm with relaxation has been shown to converge to the optimal point while providing fast convergence. We adapt the relaxed OS-SPS algorithm to the problem of image restoration. Because data acquisition in image restoration is different from that in tomography, we employ a different strategy for choosing subsets, using pixel locations rather than projection angles. Simulation results show that the relaxed OS-SPS algorithm can provide an order-of-magnitude acceleration over the EM algorithm for image restoration. This new algorithm now provides the speed and guaranteed convergence necessary for efficient image restoration.

Journal Article↗

Estimation and inference for a spline-enhanced population pharmacokinetic model.

This article is motivated by an application where subjects were dosed three times with the same drug and the drug concentration profiles appeared to be the lowest after the third dose. One possible explanation is that the pharmacokinetic (PK) parameters vary over time. Therefore, we consider population PK models with time-varying PK parameters. These time-varying PK parameters are modeled by natural cubic spline functions in the ordinary differential equations. Mean parameters, variance components, and smoothing parameters are jointly estimated by maximizing the double penalized log likelihood. Mean functions and their derivatives are obtained by the numerical solution of ordinary differential equations. The interpretation of PK parameters in the model and its flexibility are discussed. The proposed methods are illustrated by application to the data that motivated this article. The model's performance is evaluated through simulation.

Biometry↗