PubMed Health⌕ Search

PubMed · 8888616

How dependencies between successive examples affect on-line learning.

Abstract

We study the dynamics of on-line learning for a large class of neural networks and learning rules, including backpropagation for multilayer perceptrons. In this paper, we focus on the case where successive examples are dependent, and we analyze how these dependencies affect the learning process. We define the representation error and the prediction error. The representation error measures how well the environment is represented by the network after learning. The prediction error is the average error that a continually learning network makes on the next example. In the neighborhood of a local minimum of the error surface, we calculate these errors. We find that the more predictable the example presentation, the higher the representation error, i.e., the less accurate the asymptotic representation of the whole environment. Furthermore we study the learning process in the presence of a plateau. Plateaus are flat spots on the error surface, which can severely slow down the learning process. In particular, they are notorious in applications with multilayer perceptrons. Our results, which are confirmed by simulations of a multilayer perceptron learning a chaotic time series using backpropagation, explain how dependencies between examples can help the learning process to escape from a plateau.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

W Wiegerinck, T Heskes. 1996-11-15. How dependencies between successive examples affect on-line learning.. https://doi.org/10.1162/neco.1996.8.8.1743

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Cascade associative memory storing hierarchically correlated patterns with various correlations.

In conventional models for storing hierarchically correlated patterns, correlations between ancestors (first-level patterns) and their descendants (second-level ones) are assumed to be uniform, so that the descendants are distributed around their ancestors with equal distances. However, this assumption might be unnatural. We believe that objects are encoded into patterns by preserving the similarity between them. In this case, descendants are distributed around their ancestors with various distances, so that the assumption is invalid and the conventional models become inapplicable. To overcome this, we propose a model CASM3 for storing hierarchically correlated patterns with various correlations. In CASM3, critical load levels vary with the descendants, and become higher with increasing correlations. Increase in load level successively destroys the memories of the descendants in descending order of their correlations. The size of the basins of attraction depends on the range of the correlations, and becomes larger as the correlation range is shifted toward lower levels.

Learning↗

Bounds on error expectation for support vector machines.

We introduce the concept of span of support vectors (SV) and show that the generalization ability of support vector machines (SVM) depends on this new geometrical concept. We prove that the value of the span is always smaller (and can be much smaller) than the diameter of the smallest sphere containing the support vectors, used in previous bounds (Vapnik, 1998). We also demonstrate experimentally that the prediction of the test error given by the span is very accurate and has direct application in model selection (choice of the optimal parameters of the SVM).

Learning↗