PubMed Health⌕ Search

Biomedical subjects

Tom Heskes

Publications and source records attributed to Tom Heskes.

5 recordsLinked to original sources

Predicting carcinoid heart disease with the noisy-threshold classifier.

OBJECTIVE: To predict the development of carcinoid heart disease (CHD), which is a life-threatening complication of certain neuroendocrine tumors. To this end, a novel type of Bayesian classifier, known as the noisy-threshold classifier, is applied. MATERIALS AND METHODS: Fifty-four cases of patients that suffered from a low-grade midgut carcinoid tumor, of which 22 patients developed CHD, were obtained from the Netherlands Cancer Institute (NKI). Eleven attributes that are known at admission have been used to classify whether the patient develops CHD. Classification accuracy and area under the receiver operating characteristics (ROC) curve of the noisy-threshold classifier are compared with those of the naive-Bayes classifier, logistic regression, the decision-tree learning algorithm C4.5, and a decision rule, as formulated by an expert physician. RESULTS: The noisy-threshold classifier showed the best classification accuracy of 72% correctly classified cases, although differences were significant only for logistic regression and C4.5. An area under the ROC curve of 0.66 was attained for the noisy-threshold classifier, and equaled that of the physician's decision-rule. CONCLUSIONS: The noisy-threshold classifier performed favorably to other state-of-the-art classification algorithms, and equally well as a decision-rule that was formulated by the physician. Furthermore, the semantics of the noisy-threshold classifier make it a useful machine learning technique in domains where multiple causes influence a common effect.

Algorithms↗

Improving Cox survival analysis with a neural-Bayesian approach.

In this article we show that traditional Cox survival analysis can be improved upon when supplemented with sensible priors and analysed within a neural Bayesian framework. We demonstrate that the Bayesian method gives more reliable predictions, in particular for relatively small data sets. The obtained posterior (the probability distribution of network parameters given the data) which in itself is intractable, can be made accessible by several approximations. We review approximations by Hybrid Markov Chain Monte Carlo sampling, a variational method and the Laplace approximation. We argue that although each Bayesian approach circumvents the shortcomings of the original Cox analysis, and therefore yields better predictive results, in practice the use of variational methods or Laplace is preferable. Since Cox survival analysis is infamous for its poor results with (too) many inputs, we use the Bayesian posterior to estimate p-values on the inputs and to formulate an algorithm for backward elimination. We show that after removal of irrelevant inputs Bayesian methods still achieve significantly better results than classical Cox.

Antineoplastic Agents↗

On the uniqueness of loopy belief propagation fixed points.

We derive sufficient conditions for the uniqueness of loopy belief propagation fixed points. These conditions depend on both the structure of the graph and the strength of the potentials and naturally extend those for convexity of the Bethe free energy. We compare them with (a strengthened version of) conditions derived elsewhere for pairwise potentials. We discuss possible implications for convergent algorithms, as well as for other approximate free energies.

Journal Article↗

Clustering ensembles of neural network models.

We show that large ensembles of (neural network) models, obtained e.g. in bootstrapping or sampling from (Bayesian) probability distributions, can be effectively summarized by a relatively small number of representative models. In some cases this summary may even yield better function estimates. We present a method to find representative models through clustering based on the models' outputs on a data set. We apply the method on an ensemble of neural network models obtained from bootstrapping on the Boston housing data, and use the results to discuss bootstrapping in terms of bias and variance. A parallel application is the prediction of newspaper sales, where we learn a series of parallel tasks. The results indicate that it is not necessary to store all samples in the ensembles: a small number of representative models generally matches, or even surpasses, the performance of the full ensemble. The clustered representation of the ensemble obtained thus is much better suitable for qualitative analysis, and will be shown to yield new insights into the data.

Algorithms↗

Task-Dependent Learning of Attention.

In this article, we propose a neural network model for selective covert visual attention. This model can learn to focus its attention on important features depending on the task to be fulfilled by gating the flow of information from the lower to the higher levels of the visual system. The model is kept as simple as possible, but it is still capable of reproducing attentional behavior observed in psychological experiments. Computer simulations demonstrate that (1) it can learn categories to reduce reaction time without a decrease in performance, (2) the model reveals a performance similar to that of humans in feature and conjunction search, and (3) its learning dynamics are comparable with those of humans. Copyright 1997 Elsevier Science Ltd.

Journal Article↗