PubMed Health⌕ Search

PubMed · 11674845

Predictability, complexity, and learning.

Abstract

We define predictive information I(pred)(T) as the mutual information between the past and the future of a time series. Three qualitatively different behaviors are found in the limit of large observation times T:I(pred)(T) can remain finite, grow logarithmically, or grow as a fractional power law. If the time series allows us to learn a model with a finite number of parameters, then I(pred)(T) grows logarithmically with a coefficient that counts the dimensionality of the model space. In contrast, power-law growth is associated, for example, with the learning of infinite parameter (or nonparametric) models such as continuous functions with smoothness constraints. There are connections between the predictive information and measures of complexity that have been defined both in learning theory and the analysis of physical systems through statistical mechanics and dynamical systems theory. Furthermore, in the same way that entropy provides the unique measure of available information consistent with some simple and plausible conditions, we argue that the divergent part of I(pred)(T) provides the unique measure for the complexity of dynamics underlying a time series. Finally, we discuss how these ideas may be useful in problems in physics, statistics, and biology.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

W Bialek, I Nemenman, N Tishby. 2001. Predictability, complexity, and learning.. https://doi.org/10.1162/089976601753195969

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related citations

Probability of nitrate contamination of recently recharged groundwaters in the conterminous United States.

A new logistic regression (LR) model was used to predict the probability of nitrate contamination exceeding 4 mg/L in predominantly shallow, recently recharged groundwaters of the United States. The new model contains variables representing (1) N fertilizer loading (p < 0.001), (2) percent cropland-pasture (p < 0.001), (3) natural log of human population density (p < 0.001), (4) percent well-drained soils (p < 0.001), (5) depth to the seasonally high water table (p < 0.001), and (6) presence or absence of unconsolidated sand and gravel aquifers (p = 0.002). Observed and average predicted probabilities associated with deciles of risk are well correlated (r2 = 0.875), indicating that the LR model fits the data well. The likelihood of nitrate contamination is greater in areas with high N loading and well-drained surficial soils over unconsolidated sand and gravels. The LR model correctly predicted the status of nitrate contamination in 75% of wells in a validation data set. Considering all wells used in both calibration and validation, observed median nitrate concentration increased from 0.24 to 8.30 mg/L as the mapped probability of nitrate exceeding 4 mg/L increased from < or =0.17 to >0.83.

Forecasting↗

Use of the most likely failure point method for risk estimation and risk uncertainty analysis.

The most likely failure point (MLFP) method, developed within the field of structural reliability analysis (where it is known as the FORM/SORM method) is a technique for estimating the risk (probability) that a calculated quantity Q exceeds a set limit Q(lim) when some or all of the inputs to the calculation are uncertain. It can be used as an efficient stand-alone method for this type of risk calculation. However, for application within the field of toxic hazards, it is proposed as a means for performing sensitivity analyses, possibly in parallel with a risk calculation carried out by conventional methods. The basis of the method is outlined and its use is demonstrated by means of an example calculation of the risk arising from an installation containing chlorine. The calculation uses, as a consequence model, commercial software for the prediction of dense gas transport. The risk estimate is shown to be acceptably close to that obtained by the Monte Carlo method. The use of a proposed screening procedure utilising the sensitivity formulas that the method provides, in order to identify the most significant uncertainties, is demonstrated. The identification of a single set of input values containing sufficient information to summarise (at least approximately) the entire risk analysis is considered to be an important feature of the method and is proposed as the basis of a means for assessing the validity of the consequence model.

Forecasting↗