PubMed Health⌕ Search

Biomedical subjects

Asger Hobolth

Publications and source records attributed to Asger Hobolth.

6 recordsLinked to original sources

CpG + CpNpG analysis of protein-coding sequences from tomato.

We develop codon-based models for simultaneously inferring the mutational effects of CpG and CpNpG methylation in coding regions. In a data set of 369 tomato genes, we show that there is very little effect of CpNpG methylation but a strong effect of CpG methylation affecting almost all genes. We further show that the CpNpG and CpG effects are largely uncorrelated. Our results suggest different roles of CpG and CpNpG methylation, with CpNpG methylation possibly playing a specialized role in defense against transposons and RNA viruses.

Base Composition↗

Statistical inference in evolutionary models of DNA sequences via the EM algorithm.

We describe statistical inference in continuous time Markov processes of DNA sequences related by a phylogenetic tree. The maximum likelihood estimator can be found by the expectation maximization (EM) algorithm and an expression for the information matrix is also derived. We provide explicit analytical solutions for the EM algorithm and information matrix.

Journal Article↗

Comparative analysis of protein coding sequences from human, mouse and the domesticated pig.

BACKGROUND: The availability of abundant sequence data from key model organisms has made large scale studies of molecular evolution an exciting possibility. Here we use full length cDNA alignments comprising more than 700,000 nucleotides from human, mouse, pig and the Japanese pufferfish Fugu rubrices in order to investigate 1) the relationships between three major lineages of mammals: rodents, artiodactyls and primates, and 2) the rate of evolution and the occurrence of positive Darwinian selection using codon based models of sequence evolution. RESULTS: We provide evidence that the evolutionary splits among primates, rodents and artiodactyls happened shortly after each other, with most gene trees favouring a topology with rodents as outgroup to primates and artiodactyls. Using an unrooted topology of the three mammalian species we show that since their diversification, the pig and mouse lineages have on average experienced 1.44 and 2.86 times as many synonymous substitutions as humans, respectively, whereas the rates of non-synonymous substitutions are more similar. The analysis shows the highest average dN/dS ratio in the human lineage, followed by the pig and then the mouse lineages. Using codon based models we detect signals of positive Darwinian selection in approximately 5.3%, 4.9% and 6.0% of the genes on the human, pig and mouse lineages respectively. Approximately 16.8% of all the genes studied here are not currently annotated as functional genes in humans. Our analyses indicate that a large fraction of these genes may have lost their function quite recently or may still be functional genes in some or all of the three mammalian species. CONCLUSIONS: We present a comparative analysis of protein coding genes from three major mammalian lineages. Our study demonstrates the usefulness of codon-based likelihood models in detecting selection and it illustrates the value of sequencing organisms at different phylogenetic distances for comparative studies.

Animals↗

Pseudo-likelihood analysis of codon substitution models with neighbor-dependent rates.

Recently, Markov processes for the evolution of coding DNA with neighbor dependence in the instantaneous substitution rates have been considered. The neighbor dependency makes the models analytically intractable, and previously Markov chain Monte Carlo methods have been used for statistical inference. Using a pseudo-likelihood idea, we introduce in this paper an approximative estimation method which is fast to compute. The pseudo-likelihood estimates are shown to be very accurate, and from analyzing 348 human-mouse coding sequences we conclude that the incorporation of a CpG effect improves the fit of the model considerably.

Algorithms↗

Applications of hidden Markov models for characterization of homologous DNA sequences with a common gene.

Identifying and characterizing the structure in genome sequences is one of the principal challenges in modern molecular biology, and comparative genomics offers a powerful tool. In this paper, we introduce a hidden Markov model that allows a comparative analysis of multiple sequences related by a phylogenetic tree, and we present an efficient method for estimating the parameters of the model. The model integrates structure prediction methods for one sequence, statistical multiple alignment methods, and phylogenetic information. This unified model is particularly useful for a detailed characterization of DNA sequences with a common gene. We illustrate the model on a variety of homologous sequences.

Agrobacterium tumefaciens↗

The spherical deformation model.

Miller et al. (1994) describe a model for representing spatial objects with no obvious landmarks. Each object is represented by a global translation and a normal deformation of a sphere. The normal deformation is defined via the orthonormal spherical-harmonic basis. In this paper we analyse the spherical deformation model in detail and describe how it may be used to summarize the shape of star-shaped three-dimensional objects with few parameters. It is of interest to make statistical inference about the three-dimensional shape parameters from continuous observations of the surface and from a single central section of the object. We use maximum-likelihood-based inference for this purpose and demonstrate the suggested methods on real data.

Analysis of Variance↗