PubMed Health⌕ Search

Biomedical subjects

D Husmeier

Publications and source records attributed to D Husmeier.

6 recordsLinked to original sources

Detection of recombination in DNA multiple alignments with hidden Markov models.

Conventional phylogenetic tree estimation methods assume that all sites in a DNA multiple alignment have the same evolutionary history. This assumption is violated in data sets from certain bacteria and viruses due to recombination, a process that leads to the creation of mosaic sequences from different strains and, if undetected, causes systematic errors in phylogenetic tree estimation. In the current work, a hidden Markov model (HMM) is employed to detect recombination events in multiple alignments of DNA sequences. The emission probabilities in a given state are determined by the branching order (topology) and the branch lengths of the respective phylogenetic tree, while the transition probabilities depend on the global recombination probability. The present study improves on an earlier heuristic parameter optimization scheme and shows how the branch lengths and the recombination probability can be optimized in a maximum likelihood sense by applying the expectation maximization (EM) algorithm. The novel algorithm is tested on a synthetic benchmark problem and is found to clearly outperform the earlier heuristic approach. The paper concludes with an application of this scheme to a DNA sequence alignment of the argF gene from four Neisseria strains, where a likely recombination event is clearly detected.

Algorithms↗

Probabilistic divergence measures for detecting interspecies recombination.

This paper proposes a graphical method for detecting interspecies recombination in multiple alignments of DNA sequences. A fixed-size window is moved along a given DNA sequence alignment. For every position, the marginal posterior probability over tree topologies is determined by means of a Markov chain Monte Carlo simulation. Two probabilistic divergence measures are plotted along the alignment, and are used to identify recombinant regions. The method is compared with established detection methods on a set of synthetic benchmark sequences and two real-world DNA sequence alignments.

Computational Biology↗

Learning non-stationary conditional probability distributions.

While sophisticated neural networks and graphical models have been developed for predicting conditional probabilities in a non-stationary environment, major improvements in the training schemes are still required to make these approaches practically viable.

Models, Neurological↗

The Bayesian evidence scheme for regularizing probability-density estimating neural networks.

Training probability-density estimating neural networks with the expectation-maximization (EM) algorithm aims to maximize the likelihood of the training set and therefore leads to overfitting for sparse data. In this article, a regularization method for mixture models with generalized linear kernel centers is proposed, which adopts the Bayesian evidence approach and optimizes the hyperparameters of the prior by type II maximum likelihood. This includes a marginalization over the parameters, which is done by Laplace approximation and requires the derivation of the Hessian of the log-likelihood function. The incorporation of this approach into the standard training scheme leads to a modified form of the EM algorithm, which includes a regularization term and adapts the hyperparameters on-line after each EM cycle. The article presents applications of this scheme to classification problems, the prediction of stochastic time series, and latent space models.

Acquired Immunodeficiency Syndrome↗

An empirical evaluation of Bayesian sampling with hybrid Monte Carlo for training neural network classifiers.

This article gives a concise overview of Bayesian sampling for neural networks, and then presents an extensive evaluation on a set of various benchmark classification problems. The main objective is to study the sensitivity of this scheme to changes in the prior distribution of the parameters and hyperparameters, and to evaluate the efficiency of the so-called automatic relevance determination (ARD) method. The article concludes with a comparison of the achieved classification results with those obtained with (i) the evidence scheme and (ii) with non-Bayesian methods.

Journal Article↗

Structural fluctuations and conformational entropy in proteins: entropy balance in an intramolecular reaction in methemoglobin.

The reversible intramolecular binding of the distal histidine side chain to the heme iron in methemoglobin is of special interest due to the very large negative reaction entropy which overcompensates the large reaction enthalpy. It may be considered as a prominent example of the ability of proteins (including enzymes) to provide global entropy in a local process. In this work new experiments and model calculations are reported which aim at finding the structural elements contributing to the reaction entropy. Geometrical studies prove the implication of the 20 residue E-helix being shifted by more than 2 A. Vibrational entropies are calculated by a procedure derived from the method of Karplus and Kushik. It turns out that neither the histidine alone nor the complete E-helix contribute more than 15 per cent of the required entropy.

Amino Acid Sequence↗