PubMed HealthSearch

SEARCH · PubMed Health

Results for “Likelihood Functions”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 recordsLinked to original sources

Likelihood-based optimization enables accurate copy number estimation for paralogous genes using exome data.

MOTIVATION: Exome sequencing is widely used for genetic studies; however, accurate detection of copy number variants (CNV) in paralogous genes is challenging due to short-read mapping ambiguity and extensive copy-number variation. The human genome contains several hundred paralogous genes, many of which are known to harbor disease-associated CNVs. Existing exome CNV callers are primarily designed for rare CNV detection in uniquely mappable regions and are not well-suited for paralogous genes. METHODS: We describe a computational method (EdgeCopy) for copy number profiling of paralogous genes using whole-exome sequence data. EdgeCopy aggregates reads mapped to all copies of paralogous genes and relates observed read depth to copy number for multiple exome samples using an approximate composite likelihood function. The likelihood function is optimized using numerical optimization to obtain gene-level fractional copy number estimates that are discretized and refined using a Hidden Markov Model to obtain exon-level copy number estimates. RESULTS: Benchmarking of Edgecopy using experimental copy number data showed high concordance (mean = 0.973) for six disease-associated paralogous genes. We evaluated performance using whole-exome data from approximately 2400 samples across five continental populations from the 1000 Genomes Project. EdgeCopy shows robust concordance with whole-genome sequencing based estimates (0.974-0.982) across populations and 130 paralogous genes spanning a wide range of copy-number variation. In comparison, copy number analysis using a state-of-the-art exome CNV caller failed to estimate copy number for paralogous genes with very high mapping ambiguity and showed much lower concordance (0.565) for CNV events compared to EdgeCopy (0.908). AVAILABILITY: EdgeCopy is freely available at https://github.com/vibansal-lab/edgecopy.

Humans

Derivative-free restricted maximum likelihood estimation in animal models with a sparse matrix solver.

Estimation of (co)variance components by derivative-free REML requires repeated evaluation of the log-likelihood function of the data. Gaussian elimination of the augmented mixed model coefficient matrix is often used to evaluate the likelihood function, but it can be costly for animal models with large coefficient matrices. This study investigated the use of a direct sparse matrix solver to obtain the log-likelihood function. The sparse matrix package SPARSPAK was used to reorder the mixed model equations once and then repeatedly to solve the equations by Cholesky factorization to generate the terms required to calculate the likelihood. The animal model used for comparison contained 19 fixed levels, 470 maternal permanent environmental effects, and 1586 direct and 1586 maternal genetic effects, resulting in a coefficient matrix of order 3661 with .3% nonzero elements after including numerator relationships. Compared with estimation via Gaussian elimination of the unordered system, utilization of SPARSPAK required 605 and 240 times less central processing unit time on mainframes and personal computers, respectively. The SPARSPAK package also required less memory and provided solutions for all effects in the model.

Algorithms

Distribution-free regression analysis of grouped survival data.

Methods based on regression models for logarithmic hazard functions, Cox models, are given for analysis of grouped and censored survival data. By making an approximation it is possible to obtain explicitly a maximum likelihood function involving only the regression parameters. This likelihood function is a convenient analog to Cox's partial likelihood for ungrouped data. The method is applied to data from a toxicological experiment.

Animals

Accounting for contact tracing in epidemiological birth-death models.

Phylodynamics bridges the gap between classical epidemiology and pathogen genome sequence data by estimating epidemiological parameters from time-scaled pathogen phylogenetic trees. The models used in phylodynamics typically assume that the sampling procedure is independent between infected individuals. However, this assumption does not hold for many epidemics, in particular for such sexually transmitted infections as HIV-1, for which contact tracing schemes are included in health policies of many countries. We extended phylodynamic multi-type birth-death (MTBD) models with contact tracing (CT), and developed a simulator to generate trees under MTBD and MTBD-CT models. We proposed a non-parametric test for detecting contact tracing in pathogen phylogenetic trees. Its application to simulated data showed that it is both highly specific and sensitive. For the simplest representative of the MTBD-CT family, the BD-CT(1) model, where only the last contact can be notified, we solved the differential equations and proposed a closed form solution for the likelihood function. We implemented a maximum-likelihood program, which estimates the BD-CT(1) model parameters and their confidence intervals from phylogenetic trees. It performed accurate parameter inference on BD and BD-CT(1) simulated data, and detected contact tracing in HIV-1 B epidemics in Zurich and the UK. Importantly, we showed that not accounting for contact tracing when it is present, leads to bias in parameter estimation with the BD model (overestimation of the becoming-non-infectious rate). This bias is also present, but greatly reduced, when the BD-CT(1) model is used on data where multiple contacts can be notified. Our CT test, MTBD-CT tree simulator and BD-CT(1) parameter estimator are freely available at GitHub (evolbioinfo/treesimulator and evolbioinfo/bdct).

Contact Tracing

Maximum likelihood estimation and testing of a poisson regression model.

A Poisson regression model is proposed for the analysis of incidence rates presented in a two-way table classified by two categorical variables. It is shown that the likelihood function is the same as that using Glasser's exponential covariate model. An algorithm is given to solve the maximum likelihood estimates of the regression parameters. The model is evaluated via deviance and the method is illustrated with an example. Some extensions of the model are discussed.

Adult

Testing hypotheses in case-control studies--equivalence of Mantel-Haenszel statistics and logit score tests.

The two approaches in common use for the analysis of case-control studies are cross-classification by confounding variables, and modeling the logarithm of the odds ratio as a function of exposure and confounding variables. We show here that score statistics derived from the likelihood function in the latter approach are identical to the Mantel-Haenszel test statistics appropriate for the former approach. This identity holds in the most general situation considered, testing for marginal homogeneity in mK tables. This equivalence is demonstrated by a permutational argument which leads to a general likelihood expression in which the exposure variable may be a vector of discrete and/or continuous variables and in which more than two comparison groups may be considered. This likelihood can be used in analyzing studies in which there are multiple controls for each case or in which several disease categories are being compared. The possibility of including continuous variables makes this likelihood useful in situations that cannot be treated using the Mantel-Haenszel cross-classification approach.

Epidemiologic Methods

Parallel simulated annealing for emission tomography.

A method for implementing simulated annealing in parallel to speed up the execution of emission tomography (ET) image reconstruction is presented. A high degree of parallelism can be attained by using a parallel-acceptance partitioning strategy, in which perturbations to subsets of the estimate are evaluated in parallel. However because the point spread function in ET imaging systems is globally dependent, processors cannot update the current estimate independently. Consequently, processors must be synchronized each time a perturbation is accepted to avoid introducing error. This can produce excessive communications overhead, especially when the acceptance rate is high. In this paper an energy function is constructed to reduce the synchronization requirements by using a reformulation of the log-likelihood function from the expectation maximization (EM) algorithm. The approach is to change the global dependence in the energy function from the current estimate to the estimate generated during the last iteration. The synchronization requirements for guaranteed convergence are then significantly reduced from once per acceptance to once per iteration. This parallel implementation on 54 Inmos T800 transputers connected in a ring topology resulted in execution times that were almost 50 times faster than on a VAX 8600.

Algorithms

Familial aggregation of lipids and lipoproteins in families ascertained through random and nonrandom probands in the Stanford Lipid Research Clinics Family Study.

We examined the familial aggregation of lipids [total cholesterol (CH) and triglyceride (TG)] and lipoproteins [high-density lipoprotein cholesterol (HDL) and low-density lipoprotein cholesterol (LDL)] in families ascertained through random and nonrandom probands in the Stanford Lipid Research Clinics Family Study. Nonrandom probands were selected because their lipid levels at a prior screening visit exceeded a certain prespecified threshold. The statistical method is based on selection through indirect truncation on a correlated trait (in which the likelihood function is conditioned on the actual event that the proband's value is beyond the threshold). This method allows for estimation of the path model parameters in randomly and nonrandomly ascertained families jointly and separately, thus enabling tests of heterogeneity between the two types of samples. The results suggest that the multifactorial transmission is homogeneous in the random and hyperlipidemic samples for CH. However, the evidence for heterogeneity is moderate for LDL, marked HDL, and mixed for TG. The general pattern of observed results is for somewhat higher genetic heritabilities in the random than nonrandom samples, which is compatible with a higher prevalence in the random sample of certain dyslipoproteinemias associated with nonelevated lipids. Substantial genetic heritability is found for CH, HDL, and LDL, with somewhat lower estimates for TG. Cultural heritability is low but significant for all four traits. Little or no spouse resemblance or nontransmitted shared sibship effects are seen. In contrast to the findings from previous studies, little or no parental cultural transmission is seen.

Adult

Using an autoregressive model to detect departures from steady states in unequally spaced tumour biomarker data.

A new method, based on a continuous time autoregressive [CAR(1)] model of time series data, is provided for detecting departures of tumour markers from steady states in breast cancer patients following surgery. A Kalman filter recursive algorithm is used to calculate the likelihood function arising from the CAR(1) model and to calculate recursive residuals, which are monitored by a Shewhart-Cusum scheme. This approach can be used to monitor the serial marker data of large numbers of patients even when the series are short and the data are serially correlated and unequally spaced. Further, the methodology can be used to recommend appropriate testing intervals.

Algorithms

Myxoma virus expresses a secreted protein with homology to the tumor necrosis factor receptor gene family that contributes to viral virulence.

Poxviruses are known to contain a large number of open reading frames, particularly near the termini of the viral genome, that are not required for growth in tissue culture. However, many of these gene products are believed to play important roles in determining the virulence of the virus by modulating the host immune response to the infection. Recently it has been shown that Shope fibroma virus encodes, within the terminal inverted repeats, a protein (T2) related to the cellular tumor necrosis factor receptor (TNFR) and which specifically binds both TNF alpha and TNF beta. We have sequenced the terminal regions of two other Leporipoxviruses (myxoma virus and malignant rabbit fibroma virus) that are extremely invasive and capable of inducing extensive immunosuppression in rabbits and demonstrate that they also encode a closely related T2 homolog with all the structural motifs predicted for a secreted TNF binding protein. To investigate the biological role of the T2 protein, we have inactivated the myxoma virus T2 gene within each copy of the viral TIR by the insertion of a dominant selectable marker (Escherichia coli guanosine phosphoribosyltransferase) and selection of the recombinant virus in the presence of mycophenolic acid. The success of the inactivation of both copies of T2 was confirmed by the loss a broad protein band (52-56 kDa) of the predicted size for T2 from the profile of proteins secreted from mutant virus-infected BGMK cells at early times after infection. Although the T2-minus recombinant myxoma virus grew normally in tissue culture, upon infection of susceptible rabbits the viral disease was observed to be significantly attenuated. The majority of infected rabbits were able to mount an effective immune response to the infection and completely recovered. These survivor rabbits became immune to subsequent challenge with wild type myxoma virus. We conclude that the T2 viral protein is an important secreted virulence factor and that it in all likelihood functions by compromising the antiviral effects of TNF. We propose the term "viroceptor" to describe viral-encoded homologs of cellular lymphokine receptors whose function is to intercept the activity of the cognate lymphokine in order to short circuit the host immune response to the viral infection.

Amino Acid Sequence

An optimization strategy for a biokinetic model of inhaled radionuclides.

Models for material disposition and dosimetry involve predictions of the biokinetics of the material among compartments representing organs and tissues in the body. Because of a lack of human data for most toxicants, many of the basic data are derived by modeling the results obtained from studies using laboratory animals. Such a biomathematical model is usually developed by adjusting the model parameters to make the model predictions match the measured retention and excretion data visually. The fitting process can be very time-consuming for a complicated model, and visual model selections may be subjective and easily biased by the scale or the data used. Due to the development of computerized optimization methods, manual fitting could benefit from an automated process. However, for a complicated model, an automated process without an optimization strategy will not be efficient, and may not produce fruitful results. In this paper, procedures for, and implementation of, an optimization strategy for a complicated mathematical model is demonstrated by optimizing a biokinetic model for 144Ce in fused aluminosilicate particles inhaled by beagle dogs. The optimized results using SimuSolv were compared to manual fitting results obtained previously using the model simulation software GASP. Also, statistical criteria provided by SimuSolv, such as likelihood function values, were used to help or verify visual model selections.

Administration, Inhalation

Bayesian image processing in magnetic resonance imaging.

In the past several years, image processing techniques based on Bayesian models have received considerable attention. In our earlier work, we developed a novel Bayesian approach which was primarily aimed at the processing and reconstruction of images in positron emission tomography. In this paper, we describe how the technique has been adopted to process magnetic resonance images in order to reduce noise and artifacts, thereby improving image quality. In this framework, the image is assumed to be a statistical variable whose posterior probability density conditional on the observed image is modeled by the product of the likelihood function of the observed data with a prior density based our prior knowledge. A Gibbs random field incorporating local continuity information and with edge-detection capability is used as the prior model. Based on the formalism of the posterior density, we can compute an estimate of the image using an iterative technique. We have implemented this technique and applied it to phantom and clinical images. Our results indicate that the approach works reasonably well for reducing noise, enhancing edges, and removing ringing artifact.

Algorithms

Point and interval estimation in the combination of bioassay results.

A procedure for combining evidence from different biological assays is shown to be equivalent both to generalized least-squares and to maximum-likelihood estimation. By appropriate nesting of hypotheses, the likelihood function can be used to test the agreement between the assays and to obtain probability limits for the combined estimate of potency. The properties of these limits are examined, with particular reference to the situation, unusual but not impossible in practice, in which the values of relative potency that they define consist of several disjoint segments instead of a single interval. The connection with general theory of estimating linear functional relations is pointed out.

Biological Assay

AIDS in Ireland: the reporting delay distribution and the implementation of integral equation models.

This paper deals with two basic aspects concerning the modelling of AIDS incidence in the context of Irish data. We describe initially the adjustment of the number of AIDS cases (Xij) to allow for reporting delays, where a simple form of the likelihood function for the Xij is supported by GLIM. Subsequently, we consider the accessibility of numerical solution (through a NAG routine) of the integral equation models generated by the back-projection method for the adjusted AIDS cases. Results for the Irish data are summarized for various choices of the incidence distribution.

Acquired Immunodeficiency Syndrome

Sparse Logistic Regression on Genomic Data for Prediction of Tumour Pathological Subtype.

The correct prediction of tumour subtype is critical for the treatment of cancer patients to maximise the chance of survival. The patients' genomic information, such as copy number alterations (CNA) profile, has increasingly become an important factor in the prediction to supplement the traditional pathological subtyping. The incorporation of the CNA information in a prediction model, such as logistic regression, faces two major statistical challenges: first, how to estimate the model parameters in the thousands and, second, how to deal with the correlation of CNA between genomic regions. To address them, we propose a sparse logistic regression model with random effects where some of its parameters are estimated to zero while the other parameters are non-zero. In effect, a variable selection is embedded in the modelling. To deal with the correlation of CNA across genomic regions, we extend further the model to incorporate an additional penalty in the corresponding likelihood function in the logistic regression. The results show that we can identify selected genomic regions that are informative to distinguish different tumour subtypes, while giving a good prediction ability. We illustrate the methodology using CNA dataset from a lung cancer cohort.

Journal Article

Ancestral inference. I. The problem and the method.

A method for inferring the ancestral genotypes for the founders of a population is developed. This method uses the algorithms for the computation of probabilities on pedigrees of arbitrary complexity, developed by Cannings et al. (1978) and implemented by Thompson (1977b). When characteristics are simply determined by underlying genotypes the inference problem is simplified, and larger and more complex pedigrees may therefore be analysed. The problem of estimating the allele frequencies to be used in computing prior genotype probabilities for those founders on whom a likelihood function is not required is discussed. The same method allows us to compute extinction probabilities for any combination of original founder genes; these probabilities are interesting parameters of pedigree structure, which, since they relate to the actual genes present in a population, help to provide a clearer understanding of observed distributions of autosomal traits.

Gene Frequency

Familial aggregation of lipids and lipoproteins in families ascertained through random and nonrandom probands in the Iowa Lipid Research Clinics family study.

The aggregation of lipids [total cholesterol (CH) and triglyceride (TG)] and lipoproteins [high-density lipoprotein cholesterol (HDL) and low-density lipoprotein cholesterol (LDL)] in families ascertained through random and nonrandom probands in the Iowa Lipid Research Clinics family study was examined. Nonrandom probands were selected because their lipid levels (at a prior screening visit) exceeded a certain pre-specified threshold. The statistical method conditions the likelihood function on the actual event that the proband's value is beyond the threshold. This method allows for estimation of the path model parameters in randomly and nonrandomly ascertained families jointly and separately, thus enabling tests of heterogeneity between the two types of samples. Marked heterogeneity between the random and the hyperlipidemic samples is detected in the multifactorial transmission for TG and HDL, and moderate heterogeneity is detected for CH and LDL, with a pattern of higher genetic heritability estimates in the random than nonrandom samples. The observed pattern of heterogeneity is compatible with a higher prevalence in the random sample of certain dyslipoproteinemias that are associated with nonelevated lipids. For the random samples, genetic heritabilities are higher for CH and HDL (about 60%) than for TG and LDL (about 50%). For the nonrandom samples those estimates are about 45, 40, 35 and 30% for HDL, CH, LDL and TG, respectively. Little to no cultural (familial environmental) heritability is evident for CH and LDL, although 10-20% of the phenotypic variance is due to cultural factors for TG and HDL. These results suggest that the etiologies for lipids and lipoproteins may be quite different in random versus hyperlipidemic samples.

Adult