PubMed Health⌕ Search

PubMed · 15070510

Are loss functions all the same?

Abstract

In this letter, we investigate the impact of choosing different loss functions from the viewpoint of statistical learning theory. We introduce a convexity assumption, which is met by all loss functions commonly used in the literature, and study how the bound on the estimation error changes with the loss. We also derive a general result on the minimizer of the expected risk for a convex loss function in the case of classification. The main outcome of our analysis is that for classification, the hinge loss appears to be the loss of choice. Other things being equal, the hinge loss leads to a convergence rate practically indistinguishable from the logistic loss rate and much better than the square loss rate. Furthermore, if the hypothesis space is sufficiently rich, the bounds obtained for the hinge loss are not loosened by the thresholding stage.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Lorenzo Rosasco, Ernesto De Vito, Andrea Caponnetto, Michele Piana, Alessandro Verri. 2004. Are loss functions all the same?. https://doi.org/10.1162/089976604773135104

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

A statistical property of multiagent learning based on Markov decision process.

We exhibit an important property called the asymptotic equipartition property (AEP) on empirical sequences in an ergodic multiagent Markov decision process (MDP). Using the AEP which facilitates the analysis of multiagent learning, we give a statistical property of multiagent learning, such as reinforcement learning (RL), near the end of the learning process. We examine the effect of the conditions among the agents on the achievement of a cooperative policy in three different cases: blind, visible, and communicable. Also, we derive a bound on the speed with which the empirical sequence converges to the best sequence in probability, so that the multiagent learning yields the best cooperative result.

Learning↗

Second order neurons and learning in Cohen-Grossberg networks.

The well known Cohen-Grossberg network is modified to include second order neural interconnections and also to have a learning component. Sufficient conditions are obtained for the existence of a globally exponentially stable equilibrium. The model provides a two-fold generalization of the Cohen-Grossberg network in the sense if one removes the learning component, then one gets a network with second order synaptic interactions; if both the learning component and the second order interactions are removed, then the model reduces to the standard Cohen-Grossberg network.

Learning↗