Effects of delay of reinforcement on probability learning by aphasic subjects.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
We exhibit an important property called the asymptotic equipartition property (AEP) on empirical sequences in an ergodic multiagent Markov decision process (MDP). Using the AEP which facilitates the analysis of multiagent learning, we give a statistical property of multiagent learning, such as reinforcement learning (RL), near the end of the learning process. We examine the effect of the conditions among the agents on the achievement of a cooperative policy in three different cases: blind, visible, and communicable. Also, we derive a bound on the speed with which the empirical sequence converges to the best sequence in probability, so that the multiagent learning yields the best cooperative result.
To select appropriate behaviors leading to rewards, the brain needs to learn associations among sensory stimuli, selected behaviors, and rewards. Recent imaging and neural-recording studies have revealed that the dorsal striatum plays an important role in learning such stimulus-action-reward associations. However, the putamen and caudate nucleus are embedded in distinct cortico-striatal loop circuits, predominantly connected to motor-related cerebral cortical areas and frontal association areas, respectively. This difference in their cortical connections suggests that the putamen and caudate nucleus are engaged in different functional aspects of stimulus-action-reward association learning. To determine whether this is the case, we conducted an event-related and computational model-based functional MRI (fMRI) study with a stochastic decision-making task in which a stimulus-action-reward association must be learned. A simple reinforcement learning model not only reproduced the subject's action selections reasonably well but also allowed us to quantitatively estimate each subject's temporal profiles of stimulus-action-reward association and reward-prediction error during learning trials. These two internal representations were used in the fMRI correlation analysis. The results revealed that neural correlates of the stimulus-action-reward association reside in the putamen, whereas a correlation with reward-prediction error was found largely in the caudate nucleus and ventral striatum. These nonuniform spatiotemporal distributions of neural correlates within the dorsal striatum were maintained consistently at various levels of task difficulty, suggesting a functional difference in the dorsal striatum between the putamen and caudate nucleus during stimulus-action-reward association learning.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
Many real-life decision-making problems incorporate higher-order structure, involving interdependencies between different stimuli, actions, and subsequent rewards. It is not known whether brain regions implicated in decision making, such as the ventromedial prefrontal cortex (vmPFC), use a stored model of the task structure to guide choice (model-based decision making) or merely learn action or state values without assuming higher-order structure as in standard reinforcement learning. To discriminate between these possibilities, we scanned human subjects with functional magnetic resonance imaging while they performed a simple decision-making task with higher-order structure, probabilistic reversal learning. We found that neural activity in a key decision-making region, the vmPFC, was more consistent with a computational model that exploits higher-order structure than with simple reinforcement learning. These results suggest that brain regions, such as the vmPFC, use an abstract model of task structure to guide behavioral choice, computations that may underlie the human capacity for complex social interactions and abstract strategizing.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
The notion of prediction error has established itself at the heart of formal models of animal learning and current hypotheses of dopamine function. Several interpretations of prediction error have been offered, including the model-free reinforcement learning method known as temporal difference learning (TD), and the important Rescorla-Wagner (RW) learning rule. Here, we present a model-based adaptation of these ideas that provides a good account of empirical data pertaining to dopamine neuron firing patterns and associative learning paradigms such as latent inhibition, Kamin blocking and overshadowing. Our departure from model-free reinforcement learning also offers: 1) a parsimonious distinction between tonic and phasic dopamine functions; 2) a potential generalization of the role of phasic dopamine from valence-dependent "reward" processing to valence-independent "salience" processing; 3) an explanation for the selectivity of certain dopamine manipulations on motivation for distal rewards; and 4) a plausible link between formal notions of prediction error and accounts of disturbances of thought in schizophrenia (in which dopamine dysfunction is strongly implicated). The model distinguishes itself from existing accounts by offering novel predictions pertaining to the firing of dopamine neurons in various untested behavioral scenarios.
This report summarizes 4 experiments which deal with the effects of surgical removal of the rat's telencephalic forebrain structures on performance of an inhibitory avoidance response, which was acquired either before or after the lesion was made. Two experiments provide evidence that inhibitory avoidance learning (using the single-trial up-hill avoidance task) is still possible after removal of all of the forebrain structures except for the hypothalamus. A third study using this preparation dealt with the question of whether this conditioned avoidance response can be eliminated as a consequence of it being punished. In a further experiment the conditioned avoidance response was established prior to ablation of the telencephalon plus thalamus. Recall of the conditioned response survived the lesion, suggesting that the avoidance response is also stored at a subtelencephalic level in the brain-intact animal.
Hungry fruit flies can be trained by exposing them to two chemical odorants, one paired with the opportunity to feed on 1 M sucrose. On later testing, when given a choice between odorants the flies migrate specifically toward the sucrose-paired odor. This appetitively reinforced learning by the flies is similar in strength and character to previously demonstrated negatively reinforced learning, but it differs in several properties. Both memory consolidation and memory decay proceed relatively slowly after training with sucrose reward. Consolidation of learned information into anesthesia-resistant long-term memory requires about 100 min after training with sucrose compared to about 30 min after training with electric shock. Memory in wild-type flies persists for 24 hr after training with sucrose compared to 4-6 hr after training with electric shock. Memory in amnesiac mutants appears to be similarly lengthened, from 1 hr to 6 hr, by substituting sucrose reward for shock punishment. Two other mutants, dunce and rutabaga, which were isolated because they failed to learn the shock-avoidance task, learn normally in response to sucrose reward but forget rapidly afterward. One mutant, turnip, does not learn in either paradigm. Reward and punishment can be combined in olfactory discrimination training by pairing one odor to sucrose and the other to electric shock. In this situation, the expression of learning is approximately the sum of that obtained by using either reinforcement alone. After such training, memory decays at two distinct rates, each characteristic of one type of reinforcement.
We consider the problem of stabilizing unstable equilibria by discrete controls (the controls take discrete values at discrete moments of time). We prove that discrete control typically creates a chaotic attractor in the vicinity of an equilibrium. Artificial neural networks with reinforcement learning are known to be able to learn such a control scheme. We consider examples of such systems, discuss some details of implementing the reinforcement learning to controlling unstable equilibria, and show that the arising dynamics is characterized by positive Lyapunov exponents, and hence is chaotic. This chaos can be observed both in the controlled system and in the activity patterns of the controller.
Rats were trained on a 3-dimensional, 4-arm radial maze. In Experiment 1, Ss trained to climb to the single goal platform chose fewer novel routes to the goal than Ss trained to climb to the 4 spatially distinct platforms. In Experiment 2 a reinforcement contingency was imposed, requiring a novel route choice on each trial to receive reinforcement. Learning to associate route choice with reinforcement outcome was much more difficult for Ss tested with the single goal than for Ss tested with the 4 distinct goals. In Experiment 3 a partitioned central platform group learned the reinforcement contingency as quickly as the Ss given 4 spatially distinct platforms. In Experiment 4, distinctive floor inserts did not affect performance relative to no inserts.
The role of external reinforcement is an issue of much debate and uncertainty in perceptual learning research. Although it is commonly acknowledged that external reinforcement, such as performance feedback, can aid in perceptual learning (M. H. Herzog & M. Fahle, 1997), there are many examples in which it is not required (K. Ball & R. Sekuler, 1987; M. Fahle, S. Edelman, & T. Poggio, 1995; A. Karni & D. Sagi, 1991; S. P. McKee & G. Westheimer, 1978; L. P. Shiu & H. Pashler, 1992). Additionally, learning without external reinforcement can occur even for stimuli that are irrelevant to the subject's task (A. R. Seitz & T. Watanabe, 2003). It has been thus hypothesized that internal reinforcement can serve a similar role as external reinforcement in learning (M. H. Herzog & M. Fahle, 1998; A. Seitz & T. Watanabe, 2005). This idea suggests that perceptual learning should occur in the absence of external reinforcement provided that easy exemplars are utilized as a basis for the subject to generate internal reinforcement. Here, we report results from two studies that show that this is not always the case. In the first study, subjects participated in two sessions of a motion direction discrimination task with low-contrast dots moving in directions separated by 90 degrees. In the second study, subjects participated in 12 orientation-discrimination sessions using oriented bars (oriented either 70 degrees or 110 degrees) that were masked by spatial noise. Trials of different signal levels (yielding psychometric functions ranging from chance to ceiling) were randomly interleaved. In both studies, subjects experiencing external reinforcement showed significant learning, whereas subjects receiving no external reinforcement failed to show learning. We conclude that while internal reinforcement is an important learning signal, the presence of easy exemplars is not sufficient to generate reinforcement signals.
Most reinforcement learning models of animal conditioning operate under the convenient, though fictive, assumption that Pavlovian conditioning concerns prediction learning whereas instrumental conditioning concerns action learning. However, it is only through Pavlovian responses that Pavlovian prediction learning is evident, and these responses can act against the instrumental interests of the subjects. This can be seen in both experimental and natural circumstances. In this paper we study the consequences of importing this competition into a reinforcement learning context, and demonstrate the resulting effects in an omission schedule and a maze navigation task. The misbehavior created by Pavlovian values can be quite debilitating; we discuss how it may be disciplined.
In a multi-agent environment, where the outcomes of one's actions change dynamically because they are related to the behavior of other beings, it becomes difficult to make an optimal decision about how to act. Although game theory provides normative solutions for decision making in groups, how such decision-making strategies are altered by experience is poorly understood. These adaptive processes might resemble reinforcement learning algorithms, which provide a general framework for finding optimal strategies in a dynamic environment. Here we investigated the role of prefrontal cortex (PFC) in dynamic decision making in monkeys. As in reinforcement learning, the animal's choice during a competitive game was biased by its choice and reward history, as well as by the strategies of its opponent. Furthermore, neurons in the dorsolateral prefrontal cortex (DLPFC) encoded the animal's past decisions and payoffs, as well as the conjunction between the two, providing signals necessary to update the estimates of expected reward. Thus, PFC might have a key role in optimizing decision-making strategies.
The effects of lesions, receptor blocking, electrical self-stimulation, and drugs of abuse suggest that midbrain dopamine systems are involved in processing reward information and learning approach behavior. Most dopamine neurons show phasic activations after primary liquid and food rewards and conditioned, reward-predicting visual and auditory stimuli. They show biphasic, activation-depression responses after stimuli that resemble reward-predicting stimuli or are novel or particularly salient. However, only few phasic activations follow aversive stimuli. Thus dopamine neurons label environmental stimuli with appetitive value, predict and detect rewards and signal alerting and motivating events. By failing to discriminate between different rewards, dopamine neurons appear to emit an alerting message about the surprising presence or absence of rewards. All responses to rewards and reward-predicting stimuli depend on event predictability. Dopamine neurons are activated by rewarding events that are better than predicted, remain uninfluenced by events that are as good as predicted, and are depressed by events that are worse than predicted. By signaling rewards according to a prediction error, dopamine responses have the formal characteristics of a teaching signal postulated by reinforcement learning theories. Dopamine responses transfer during learning from primary rewards to reward-predicting stimuli. This may contribute to neuronal mechanisms underlying the retrograde action of rewards, one of the main puzzles in reinforcement learning. The impulse response releases a short pulse of dopamine onto many dendrites, thus broadcasting a rather global reinforcement signal to postsynaptic neurons. This signal may improve approach behavior by providing advance reward information before the behavior occurs, and may contribute to learning by modifying synaptic transmission. The dopamine reward signal is supplemented by activity in neurons in striatum, frontal cortex, and amygdala, which process specific reward information but do not emit a global reward prediction error signal. A cooperation between the different reward signals may assure the use of specific rewards for selectively reinforcing behaviors. Among the other projection systems, noradrenaline neurons predominantly serve attentional mechanisms and nucleus basalis neurons code rewards heterogeneously. Cerebellar climbing fibers signal errors in motor performance or errors in the prediction of aversive events to cerebellar Purkinje cells. Most deficits following dopamine-depleting lesions are not easily explained by a defective reward signal but may reflect the absence of a general enabling function of tonic levels of extracellular dopamine. Thus dopamine systems may have two functions, the phasic transmission of reward information and the tonic enabling of postsynaptic neurons.
Explore the source record for details and available documents.