PubMed Health⌕ Search

Biomedical subjects

Motoaki Kawanabe

Publications and source records attributed to Motoaki Kawanabe.

5 recordsLinked to original sources

Asymptotic properties of the Fisher kernel.

This letter analyzes the Fisher kernel from a statistical point of view. The Fisher kernel is a particularly interesting method for constructing a model of the posterior probability that makes intelligent use of unlabeled data (i.e., of the underlying data density). It is important to analyze and ultimately understand the statistical properties of the Fisher kernel. To this end, we first establish sufficient conditions that the constructed posterior model is realizable (i.e., it contains the true distribution). Realizability immediately leads to consistency results. Subsequently, we focus on an asymptotic analysis of the generalization error, which elucidates the learning curves of the Fisher kernel and how unlabeled data contribute to learning. We also point out that the squared or log loss is theoretically preferable-because both yield consistent estimators-to other losses such as the exponential loss, when a linear classifier is used together with the Fisher kernel. Therefore, this letter underlines that the Fisher kernel should be viewed not as a heuristics but as a powerful statistical tool with well-controlled statistical properties.

Algorithms↗

Trading variance reduction with unbiasedness: the regularized subspace information criterion for robust model selection in kernel regression.

A well-known result by Stein (1956) shows that in particular situations, biased estimators can yield better parameter estimates than their generally preferred unbiased counterparts. This letter follows the same spirit, as we will stabilize the unbiased generalization error estimates by regularization and finally obtain more robust model selection criteria for learning. We trade a small bias against a larger variance reduction, which has the beneficial effect of being more precise on a single training set. We focus on the subspace information criterion (SIC), which is an unbiased estimator of the expected generalization error measured by the reproducing kernel Hilbert space norm. SIC can be applied to the kernel regression, and it was shown in earlier experiments that a small regularization of SIC has a stabilization effect. However, it remained open how to appropriately determine the degree of regularization in SIC. In this article, we derive an unbiased estimator of the expected squared error, between SIC and the expected generalization error and propose determining the degree of regularization of SIC such that the estimator of the expected squared error is minimized. Computer simulations with artificial and real data sets illustrate that the proposed method works effectively for improving the precision of SIC, especially in the high-noise-level cases. We furthermore compare the proposed method to the original SIC, the cross-validation, and an empirical Bayesian method in ridge parameter selection, with good results.

Bayes Theorem↗

A resampling approach to estimate the stability of one-dimensional or multidimensional independent components.

When applying unsupervised learning techniques in biomedical data analysis, a key question is whether the estimated parameters of the studied system are reliable. In other words, can we assess the quality of the result produced by our learning technique? We propose resampling methods to tackle this question and illustrate their usefulness for blind-source separation (BSS). We demonstrate that our proposed reliability estimation can be used to discover stable one-dimensional or multidimensional independent components, to choose the appropriate BSS-model, to enhance significantly the separation performance, and, most importantly, to flag components that carry physical meaning. Application to different biomedical testbed data sets (magnetoencephalography (MEG)/electrocardiography (ECG)-recordings) underline the usefulness of our approach.

Algorithms↗

A new discriminative kernel from probabilistic models.

Recently, Jaakkola and Haussler (1999) proposed a method for constructing kernel functions from probabilistic models. Their so-called Fisher kernel has been combined with discriminative classifiers such as support vector machines and applied successfully in, for example, DNA and protein analysis. Whereas the Fisher kernel is calculated from the marginal log-likelihood, we propose the TOP kernel derived; from tangent vectors of posterior log-odds. Furthermore, we develop a theoretical framework on feature extractors from probabilistic models and use it for analyzing the TOP kernel. In experiments, our new discriminative TOP kernel compares favorably to the Fisher kernel.

Models, Neurological↗

On-line learning in changing environments with applications in supervised and unsupervised learning.

An adaptive on-line algorithm extending the learning of learning idea is proposed and theoretically motivated. Relying only on gradient flow information it can be applied to learning continuous functions or distributions, even when no explicit loss function is given and the Hessian is not available. The framework is applied for unsupervised and supervised learning. Its efficiency is demonstrated for drifting and switching non-stationary blind separation tasks of acoustic signals. Furthermore applications to classification (US postal service data set) and time-series prediction in changing environments are presented.

Algorithms↗