PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12Linked to original sources

Neuronal activity in the rodent dorsal striatum in sequential navigation: separation of spatial and reward responses on the multiple T task.

The striatum plays an important role in "habitual" learning and memory and has been hypothesized to implement a reinforcement-learning algorithm to select actions to perform given the current sensory input. Many experimental approaches to striatal activity have made use of temporally structured tasks, which imply that the striatal representation is temporal. To test this assumption, we recorded neurons in the dorsal striatum of rats running a sequential navigation task: the multiple T maze. Rats navigated a sequence of four T maze turns to receive food rewards delivered in two locations. The responses of neurons that fired phasically were examined. Task-responsive phasic neurons were active as rats ran on the maze (maze-responsive) or during reward receipt (reward-responsive). Neither maze- nor reward-responsive neurons encoded simple motor commands: maze-responses were not well correlated with the shape of the rat's path and most reward-responsive neurons did not fire at similar rates at both food-delivery sites. Maze-responsive neurons were active at one or more locations on the maze, but these responses did not cluster at spatial landmarks such as turns. Across sessions the activity of maze-responsive neurons was highly correlated when rats ran the same maze. Maze-responses encoded the location of the rat on the maze and imply a spatial representation in the striatum in a task with prominent spatial demands. Maze-responsive and reward-responsive neurons were two separate populations, suggesting a divergence in striatal information processing of navigation and reward.

Algorithms↗

Learning obstacle avoidance with an operant behavior model.

Artificial intelligence researchers have been attracted by the idea of having robots learn how to accomplish a task, rather than being told explicitly. Reinforcement learning has been proposed as an appealing framework to be used in controlling mobile agents. Robot learning research, as well as research in biological systems, face many similar problems in order to display high flexibility in performing a variety of tasks. In this work, the controlling of a vehicle in an avoidance task by a previously developed operant learning model (a form of animal learning) is studied. An environment in which a mobile robot with proximity sensors has to minimize the punishment for colliding against obstacles is simulated. The results were compared with the Q-Learning algorithm, and the proposed model had better performance. In this way a new artificial intelligence agent inspired by neurobiology, psychology, and ethology research is proposed.

Avoidance Learning↗

Hippocampal lesions facilitate instrumental learning with delayed reinforcement but induce impulsive choice in rats.

BACKGROUND: Animals must frequently act to influence the world even when the reinforcing outcomes of their actions are delayed. Learning with action-outcome delays is a complex problem, and little is known of the neural mechanisms that bridge such delays. When outcomes are delayed, they may be attributed to (or associated with) the action that caused them, or mistakenly attributed to other stimuli, such as the environmental context. Consequently, animals that are poor at forming context-outcome associations might learn action-outcome associations better with delayed reinforcement than normal animals. The hippocampus contributes to the representation of environmental context, being required for aspects of contextual conditioning. We therefore hypothesized that animals with hippocampal lesions would be better than normal animals at learning to act on the basis of delayed reinforcement. We tested the ability of hippocampal-lesioned rats to learn a free-operant instrumental response using delayed reinforcement, and what is potentially a related ability -- the ability to exhibit self-controlled choice, or to sacrifice an immediate, small reward in order to obtain a delayed but larger reward. RESULTS: Rats with sham or excitotoxic hippocampal lesions acquired an instrumental response with different delays (0, 10, or 20 s) between the response and reinforcer delivery. These delays retarded learning in normal rats. Hippocampal-lesioned rats responded slightly less than sham-operated controls in the absence of delays, but they became better at learning (relative to shams) as the delays increased; delays impaired learning less in hippocampal-lesioned rats than in shams. In contrast, lesioned rats exhibited impulsive choice, preferring an immediate, small reward to a delayed, larger reward, even though they preferred the large reward when it was not delayed. CONCLUSION: These results support the view that the hippocampus hinders action-outcome learning with delayed outcomes, perhaps because it promotes the formation of context-outcome associations instead. However, although lesioned rats were better at learning with delayed reinforcement, they were worse at choosing it, suggesting that self-controlled choice and learning with delayed reinforcement tax different psychological processes.

Animals↗

Crossmodal interactions between olfactory and visual learning in Drosophila.

Different modalities of sensation interact in a synergistic or antagonistic manner during sensory perception, but whether there is also interaction during memory acquisition is largely unknown. In Drosophila reinforcement learning, we found that conditioning with concurrent visual and olfactory cues reduced the threshold for unimodal memory retrieval. Furthermore, bimodal preconditioning followed by unimodal conditioning with either a visual or olfactory cue led to crossmodal memory transfer. Crossmodal memory acquisition in Drosophila may contribute significantly to learning in a natural environment.

Animals↗

Learning movement sequences with a delayed reward signal in a hierarchical model of motor function.

A key problem in reinforcement learning is how an animal is able to learn a sequence of movements when the reward signal only occurs at the end of the sequence. We describe how a hierarchical dynamical model of motor function is able to solve the problem of delayed reward in learning movement sequences using associative (Hebbian) learning. At the lowest level, the motor system encodes simple movements or primitives, while at higher levels the system encodes sequences of primitives. During training, the network is able to learn a high level motor program composed of a specific temporal sequence of motor primitives. The network is able to achieve this despite the fact that the reward signal, which indicates whether or not the desired motor program has been performed correctly, is received only at the end of each trial during learning. Use of a continuous attractor network in the architecture enables the network to generate the motor outputs required to produce the continuous movements necessary to implement the motor sequence.

Animals↗

Positive and negative recency effects in retirement savings decisions.

Retirement savings decisions can be influenced by the fund composition of the retirement savings plan. In 2 experiments, strong composition effects were observed, with a larger percentage of resources being invested in stock funds when more stock than bond funds were offered. Although participants changed their allocations repeatedly, the opportunity to learn did not alter the composition effects. Learning processes led to positive and negative recency effects as well, providing evidence that allocations were strongly influenced by the recent performance of the different allocation options. Two learning models were tested to explain these learning processes. The first, a local adaptation learning model, assumes that people change their behavior on the basis of recent experience, whereas the second, a reinforcement learning model, assumes that decisions are made on the basis of the totality of accumulated experience. The local adaptation model was more accurate in predicting allocation decisions, in explaining positive and negative recency effects, and in showing why composition effects are not overcome by learning.

Adult↗

Modeling perceptual learning with multiple interacting elements: a neural network model describing early visual perceptual learning.

We introduce a neural network model of an early visual cortical area, in order to understand better results of psychophysical experiments concerning perceptual learning during odd element (pop-out) detection tasks (Ahissar and Hochstein, 1993, 1994a). The model describes a network, composed of orientation selective units, arranged in a hypercolumn structure, with receptive field properties modeled from real monkey neurons. Odd element detection is a final pattern of activity with one (or a few) salient units active. The learning algorithm used was the Associative reward-penalty (Ar-p) algorithm of reinforcement learning (Barto and Anandan, 1985), following physiological data indicating the role of supervision in cortical plasticity. Simulations show that network performance improves dramatically as the weights of inter-unit connections reach a balance between lateral iso-orientation inhibition, and facilitation from neighboring neurons with different preferred orientations. The network is able to learn even from chance performance, and in the presence of a large amount of noise in the response function. As additional tests of the model, we conducted experiments with human subjects in order to examine learning strategy and test model predictions.

Humans↗

Prediction of the main cortical areas and connections involved in the tactile function of the visual cortex by network analysis.

We explored the cortical pathways from the primary somatosensory cortex to the primary visual cortex (V1) by analysing connectional data in the macaque monkey using graph-theoretical tools. Cluster analysis revealed the close relationship of the dorsal visual stream and the sensorimotor cortex. It was shown that prefrontal area 46 and parietal areas VIP and 7a occupy a central position between the different clusters in the visuo-tactile network. Among these structures all the shortest paths from primary somatosensory cortex (3a, 1 and 2) to V1 pass through VIP and then reach V1 via MT, V3 and PO. Comparison of the input and output fields suggested a larger specificity for the 3a/1-VIP-MT/V3-V1 pathways among the alternative routes. A reinforcement learning algorithm was used to evaluate the importance of the aforementioned pathways. The results suggest a higher role for V3 in relaying more direct sensorimotor information to V1. Analysing cliques, which identify areas with the strongest coupling in the network, supported the role of VIP, MT and V3 in visuo-tactile integration. These findings indicate that areas 3a, 1, VIP, MT and V3 play a major role in shaping the tactile information reaching V1 in both sighted and blind subjects. Our observations greatly support the findings of the experimental studies and provide a deeper insight into the network architecture underlying visuo-tactile integration in the primate cerebral cortex.

Algorithms↗

The Schmidt Model for program planning.

The purpose of this program planning model is to provide a guide for program planners that will result in effective training programs. This program planning model emphasizes the importance of identifying learners' stressors, decreasing stress to a therapeutic level, developing a trust relationship between the programmer and the learners, and positively reinforcing learned behavior in the classroom and clinical settings.

Adaptation, Psychological↗

Learned predictions of error likelihood in the anterior cingulate cortex.

The anterior cingulate cortex (ACC) and the related medial wall play a critical role in recruiting cognitive control. Although ACC exhibits selective error and conflict responses, it has been unclear how these develop and become context-specific. With use of a modified stop-signal task, we show from integrated computational neural modeling and neuroimaging studies that ACC learns to predict error likelihood in a given context, even for trials in which there is no error or response conflict. These results support a more general error-likelihood theory of ACC function based on reinforcement learning, of which conflict and error detection are special cases.

Brain Mapping↗

Contributions of an avian basal ganglia-forebrain circuit to real-time modulation of song.

Cortical-basal ganglia circuits have a critical role in motor control and motor learning. In songbirds, the anterior forebrain pathway (AFP) is a basal ganglia-forebrain circuit required for song learning and adult vocal plasticity but not for production of learned song. Here, we investigate functional contributions of this circuit to the control of song, a complex, learned motor skill. We test the hypothesis that neural activity in the AFP of adult birds can direct moment-by-moment changes in the primary motor areas responsible for generating song. We show that song-triggered microstimulation in the output nucleus of the AFP induces acute and specific changes in learned parameters of song. Moreover, under both natural and experimental conditions, variability in the pattern of AFP activity is associated with variability in song structure. Finally, lesions of the output nucleus of the AFP prevent naturally occurring modulation of song variability. These findings demonstrate a previously unappreciated capacity of the AFP to direct real-time changes in song. More generally, they suggest that frontal cortical and basal ganglia areas may contribute to motor learning by biasing motor output towards desired targets or by introducing stochastic variability required for reinforcement learning.

Acoustic Stimulation↗

Reliability of internal prediction/estimation and its application. I. Adaptive action selection reflecting reliability of value function.

This article proposes an adaptive action-selection method for a model-free reinforcement learning system, based on the concept of the 'reliability of internal prediction/estimation'. This concept is realized using an internal variable, called the Reliability Index (RI), which estimates the accuracy of the internal estimator. We define this index for a value function of a temporal difference learning system and substitute it for the temperature parameter of the Boltzmann action-selection rule. Accordingly, the weight of exploratory actions adaptively changes depending on the uncertainty of the prediction. We use this idea for tabular and weighted-sum type value functions. Moreover, we use the RI to adjust the learning coefficient in addition to the temperature parameter, meaning that the reliability becomes a general basis for meta-learning. Numerical experiments were performed to examine the behavior of the proposed method. The RI-based Q-learning system demonstrated its features when the adaptive learning coefficient and large RI-discount rate (which indicate how the RI values of future states are reflected in the RI value of the current state) were introduced. Statistical tests confirmed that the algorithm spent more time exploring in the initial phase of learning, but accelerated learning from the midpoint of learning. It is also shown that the proposed method does not work well with the actor-critic models. The limitations of the proposed method and its relationship to relevant research are discussed.

Acclimatization↗

Comparison of methyl anthranilate and denatonium benzoate as aversants for learning in chicks.

Methyl anthranilate (MeA) has been widely used as a taste aversant for domestic chicks in the one-trial passive avoidance learning (PAL) task. However, MeA has a strong smell that may be aversive to chicks. Therefore, odourless denatonium benzoate (DB) has been suggested as an alternative taste aversant in PAL. The present study was designed to compare the efficacy of MeA and DB as aversants in the one-trial PAL task. In this task, young chicks peck a visually conspicuous bead coated with a taste aversant and in a single trial learn to avoid a similar, but uncoated bead at subsequent presentation. In Experiment 1, chicks were trained using a silver-coloured bead coated with 100% MeA, 0.5% DB or distilled water. After 3 h, MeA-trained, but not DB-trained chicks, exhibited significantly higher avoidance of the test bead than water-trained chicks. In Experiment 2, three pre-training presentations of an uncoated red bead preceded training with the silver bead. MeA-trained chicks showed significantly higher avoidance of the test bead than water-trained chicks. The numbers of water- and DB-trained chicks that avoided pecking the test bead were low and not significantly different from each other. However, DB-trained chicks exhibited significantly longer latencies to peck the test bead than water-trained chicks, indicating that they had retained some memory of the task. Thus, 0.5% DB is a weaker aversant than MeA and it does not induce high levels of learning in the one-trial PAL task. However, DB may prove useful for investigating weakly reinforced learning.

Animals↗

Is the short-latency dopamine response too short to signal reward error?

Unexpected stimuli that are behaviourally significant have the capacity to elicit a short-latency, short-duration burst of firing in mesencephalic dopaminergic neurones. An influential interpretation of the experimental data that characterize this response proposes that dopaminergic neurones have a crucial role in reinforcement learning because they signal error in the prediction of future reward. In this article we propose a different functional role for this 'short-latency dopamine response' in the mechanisms that underlie associative learning. We suggest that the initial burst of dopaminergic-neurone firing could represent an essential component in the process of switching attentional and behavioural selections to unexpected, behaviourally important stimuli. This switching response could be a crucial prerequisite for associative learning and might be part of a general short-latency response that is mediated by catecholamines and prepares the organism for an appropriate reaction to biologically significant events. Any act which in a given situation produces satisfaction becomes associated with that situation so that when the situation recurs the act is more likely than before to recur also. E.L. Thorndike (1911) 1.

Action Potentials↗