PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10Linked to original sources

Modeling functions of striatal dopamine modulation in learning and planning.

The activity of midbrain dopamine neurons is strikingly similar to the reward prediction error of temporal difference reinforcement learning models. Experimental evidence and simulation studies suggest that dopamine neuron activity serves as an effective reinforcement signal for learning of sensorimotor associations in striatal matrisomes. In the current study, we simulate dopamine neuron activity with the extended temporal difference model of Pavlovian learning and examine the influences of this signal on medium spiny neurons in striatal matrisomes. The modeled influences include transient membrane effects of dopamine D(1) receptor activation, dopamine-dependent long-term adaptations of corticostriatal transmission, and effects of dopamine on rhythmic fluctuations of the membrane potential between an elevated "up-state" and a hyperpolarized "down-state". The most dominant activity in the striatal matrisomes is assumed to elicit behaviors via projections from the basal ganglia to the thalamus and the cortex. This "standard model" performs successfully when tested for sensorimotor learning and goal-directed behavior (planning). To investigate the contributions of our model assumptions to learning and planning, we test the performance of several model variants that lack one of these mechanisms. These simulations show that the adaptation of the dopamine-like signal is necessary for sensorimotor learning and planning. Sensorimotor learning requires dopamine-dependent long-term adaptation of corticostriatal transmission. Lack of dopamine-like novelty responses decreases the number of exploratory acts, which impairs planning capabilities. The model loses its planning capabilities if the dopamine-like signal is simulated with the original temporal difference model, because the original temporal difference model does not form novel associative chains. Transient membrane effects of the dopamine-like signal on striatal firing substantially shorten the reaction time in the planning task. The capability for planning is improved by influences of dopamine on the durations of membrane potential fluctuations and by manipulations that prolong the reaction time of the model. These results suggest that responses of dopamine neurons to conditioned stimuli contribute to sensorimotor reward learning, novelty responses of dopamine neurons stimulate exploration, and transient dopamine membrane effects are important for planning.

Animals↗

A cellular mechanism of reward-related learning.

Positive reinforcement helps to control the acquisition of learned behaviours. Here we report a cellular mechanism in the brain that may underlie the behavioural effects of positive reinforcement. We used intracranial self-stimulation (ICSS) as a model of reinforcement learning, in which each rat learns to press a lever that applies reinforcing electrical stimulation to its own substantia nigra. The outputs from neurons of the substantia nigra terminate on neurons in the striatum in close proximity to inputs from the cerebral cortex on the same striatal neurons. We measured the effect of substantia nigra stimulation on these inputs from the cortex to striatal neurons and also on how quickly the rats learned to press the lever. We found that stimulation of the substantia nigra (with the optimal parameters for lever-pressing behaviour) induced potentiation of synapses between the cortex and the striatum, which required activation of dopamine receptors. The degree of potentiation within ten minutes of the ICSS trains was correlated with the time taken by the rats to learn ICSS behaviour. We propose that stimulation of the substantia nigra when the lever is pressed induces a similar potentiation of cortical inputs to the striatum, positively reinforcing the learning of the behaviour by the rats.

Animals↗

Alcohol and error processing.

A recent study indicates that alcohol consumption reduces the amplitude of the error-related negativity (ERN), a negative deflection in the electroencephalogram associated with error commission. Here, we explore possible mechanisms underlying this result in the context of two recent theories about the neural system that produces the ERN - one based on principles of reinforcement learning and the other based on response conflict monitoring.

Alcohol Drinking↗

Mutations in the dopa decarboxylase gene affect learning in Drosophila.

Fruit flies synthesize several monoamine neurotransmitters. Dopa decarboxylase (Ddc) mutations affect synthesis of two of these, dopamine and serotonin. Both transmitters are implicated in vertebrate and invertebrate learning. Therefore, we bred flies of various Ddc genotypes and tested their learning ability in positively and negatively reinforced learning tasks. Mutations in the Ddc gene diminished learning acquisition approximately in proportion to their effect on enzymatic activity. Courtship and mating sequences of the mutants appeared normal, except for one aspect of male courtship that had previously been shown to be experience dependent. In contrast, the effect on behavior patterns that do not involve learning--phototaxis, geotaxis, olfactory acuity, responsiveness to sucrose--was relatively slight under these conditions. Moderate Ddc mutations affected the acquisition of learned responses while leaving memory retention unaltered. This is in contrast to the mutations dunce , rutabaga , and amnesiac , which primarily affect short-term memory.

Animals↗

Dynamic response-by-response models of matching behavior in rhesus monkeys.

We studied the choice behavior of 2 monkeys in a discrete-trial task with reinforcement contingencies similar to those Herrnstein (1961) used when he described the matching law. In each session, the monkeys experienced blocks of discrete trials at different relative-reinforcer frequencies or magnitudes with unsignalled transitions between the blocks. Steady-state data following adjustment to each transition were well characterized by the generalized matching law; response ratios undermatched reinforcer frequency ratios but matched reinforcer magnitude ratios. We modelled response-by-response behavior with linear models that used past reinforcers as well as past choices to predict the monkeys' choices on each trial. We found that more recently obtained reinforcers more strongly influenced choice behavior. Perhaps surprisingly, we also found that the monkeys' actions were influenced by the pattern of their own past choices. It was necessary to incorporate both past reinforcers and past choices in order to accurately capture steady-state behavior as well as the fluctuations during block transitions and the response-by-response patterns of behavior. Our results suggest that simple reinforcement learning models must account for the effects of past choices to accurately characterize behavior in this task, and that models with these properties provide a conceptual tool for studying how both past reinforcers and past choices are integrated by the neural systems that generate behavior.

Animals↗

Are rats with genetic absence epilepsy behaviorally impaired?

Absence seizures in humans are characterized by unresponsiveness to external stimuli and inactivity. However, in typical generalized non-convulsive epilepsy in children, intellectual capacities are considered to be normal. Wistar rats from an inbred strain with spontaneous absence-like seizures were compared with rats from the outbred control strain in various behavioral tasks in order to detect possible impairments related either to the absence epilepsy or to occurrence of spike and wave discharges (SWD). Spontaneous circadian locomotion, exploratory activity in an open field, social interactions with an unfamiliar conspecific and mouse killing behavior were similar in both strains. Avoidance learning in a shuttle box or food reinforced learning in a Skinner test were unimpaired or even improved in epileptic rats. During performance of a learned task either in the Skinner box or in a conditioned sound-bar pressing task, SWD were suppressed in epileptic rats as long as they were working for reinforcement. SWD reappeared when the motivation to perform the task had declined: unresponsiveness to a conditioned stimulus was then observed during SWD. These data are in agreement with observations commonly described in children with typical genetic absence epilepsy.

Aggression↗

Neural mechanism for stochastic behaviour during a competitive game.

Previous studies have shown that non-human primates can generate highly stochastic choice behaviour, especially when this is required during a competitive interaction with another agent. To understand the neural mechanism of such dynamic choice behaviour, we propose a biologically plausible model of decision making endowed with synaptic plasticity that follows a reward-dependent stochastic Hebbian learning rule. This model constitutes a biophysical implementation of reinforcement learning, and it reproduces salient features of behavioural data from an experiment with monkeys playing a matching pennies game. Due to interaction with an opponent and learning dynamics, the model generates quasi-random behaviour robustly in spite of intrinsic biases. Furthermore, non-random choice behaviour can also emerge when the model plays against a non-interactive opponent, as observed in the monkey experiment. Finally, when combined with a meta-learning algorithm, our model accounts for the slow drift in the animal's strategy based on a process of reward maximization.

Algorithms↗

Learning and stabilization of altruistic behaviors in multi-agent systems by reciprocity.

Optimization of performance in collective systems often requires altruism. The emergence and stabilization of altruistic behaviors are difficult to achieve because the agents incur a cost when behaving altruistically. In this paper, we propose a biologically inspired strategy to learn stable altruistic behaviors in artificial multi-agent systems, namely reciprocal altruism. This strategy in conjunction with learning capabilities make altruistic agents cooperate only between themselves, thus preventing their exploitation by selfish agents, if future benefits are greater than the current cost of altruistic acts. Our multi-agent system is made up of agents with a behavior-based architecture. Agents learn the most suitable cooperative strategy for different environments by means of a reinforcement learning algorithm. Each agent receives a reinforcement signal that only measures its individual performance. Simulation results show how the multi-agent system learns stable altruistic behaviors, so achieving optimal (or near-to-optimal) performances in unknown and changing environments.

Altruism↗

An associational model of birdsong sensorimotor learning I. Efference copy and the learning of song syllables.

Birdsong learning provides an ideal model system for studying temporally complex motor behavior. Guided by the well-characterized functional anatomy of the song system, we have constructed a computational model of the sensorimotor phase of song learning. Our model uses simple Hebbian and reinforcement learning rules and demonstrates the plausibility of a detailed set of hypotheses concerning sensory-motor interactions during song learning. The model focuses on the motor nuclei HVc and robust nucleus of the archistriatum (RA) of zebra finches and incorporates the long-standing hypothesis that a series of song nuclei, the Anterior Forebrain Pathway (AFP), plays an important role in comparing the bird's own vocalizations with a previously memorized song, or "template." This "AFP comparison hypothesis" is challenged by the significant delay that would be experienced by presumptive auditory feedback signals processed in the AFP. We propose that the AFP does not directly evaluate auditory feedback, but instead, receives an internally generated prediction of the feedback signal corresponding to each vocal gesture, or song "syllable." This prediction, or "efference copy," is learned in HVc by associating premotor activity in RA-projecting HVc neurons with the resulting auditory feedback registered within AFP-projecting HVc neurons. We also demonstrate how negative feedback "adaptation" can be used to separate sensory and motor signals within HVc. The model predicts that motor signals recorded in the AFP during singing carry sensory information and that the primary role for auditory feedback during song learning is to maintain an accurate efference copy. The simplicity of the model suggests that associational efference copy learning may be a common strategy for overcoming feedback delay during sensorimotor learning.

Algorithms↗

Differential acquisition of a "working memory" task by the Roman strains of rats.

Twenty-eight male rats of the Roman strains-fourteen RHA (Roman High Avoidance) and fourteen RLA (Roman Low Avoidance)-were submitted to a positively reinforced task, the delayed reinforced alternation test (DRA), in a T-maze. Performances of RLA rats were significantly better than those of RHA; RLA rats also had higher VTE (vicarious trial and error) and spontaneous alternation (SA) scores. These data confirm the fact that RLA may acquire positively reinforced learning as rapidly, or even more rapidly, than RHA rats, and that the differences in active avoidance behavior between these strains depend more on differential freezing behavior than on learning and memory capacities. Since the delayed reinforced alternation is considered as a working memory test, our results suggest that the Roman strains could be used as a genetic model for the neurobiological study of this form of memory.

Animals↗

Coupled replicator equations for the dynamics of learning in multiagent systems.

Starting with a group of reinforcement-learning agents we derive coupled replicator equations that describe the dynamics of collective learning in multiagent systems. We show that, although agents model their environment in a self-interested way without sharing knowledge, a game dynamics emerges naturally through environment-mediated interactions. An application to rock-scissors-paper game interactions shows that the collective learning dynamics exhibits a diversity of competitive and cooperative behaviors. These include quasiperiodicity, stable limit cycles, intermittency, and deterministic chaos-behaviors that should be expected in heterogeneous multiagent systems described by the general replicator equations we derive.

Journal Article↗

Behavioral and neural predictors of upcoming decisions.

Although it is widely known that brain regions such as the prefrontal cortex, the amygdala, and the ventral striatum play large roles in decision making, their precise contributions remain unclear. Here, we used functional magnetic resonance imaging and principles of reinforcement learning theory to investigate the relationship between current reinforcements and future decisions. In the experiment, subjects chose between high-risk (i.e., low probability of a large monetary reward) and low-risk (high probability of a small reward) decisions. For each subject, we estimated value functions that represented the degree to which reinforcements affected the value of decision options on the subsequent trial. Individual differences in value functions predicted not only trial-to-trial behavioral strategies, such as choosing high-risk decisions following high-risk rewards, but also the relationship between activity in prefrontal and subcortical regions during one trial and the decision made in the subsequent trial. These findings provide a novel link between behavior and neural activity by demonstrating that value functions are manifested both in adjustments in behavioral strategies and in the neural activity that accompanies those adjustments.

Adult↗

[Role of the hypothalamus in originating disorders of exploratory behavior and learning in rats with excised nuclei of the locus coeruleus].

Open-field behaviour and emotionally differently reinforced learning were studied in male Wistar rats with bilaterally ablated Locus coeruleus. Histochemical analysis of the hypothalamic structures was carried out. Decrease of investigating activity and attention was found as well as disturbances of learning with emotionally-negative (painful) reinforcement. By means of histochemical methods, fluorescence characteristic for catecholamines was found to decrease sharply in paraventricular and supraoptic hypothalamic nuclei, eminentia medialis and the posterior lobe of the hypophysis.

Animals↗

Anatomy and function of the orbital frontal cortex, II: Function and relevance to obsessive-compulsive disorder.

The authors review neurophysiological, neurobehavioral, and neuropsychological investigations of the orbital frontal cortex (OFC) in human and non-human primates. The article critically examines the role of the OFC in 1) recognition of reinforcers; 2) stimulus-reinforcer learning; 3) modulation of responses based on changes in reinforcement contingencies; 4) emotions, social behavior, and autonomic regulation; 5) mnemonic functions; and 6) rule learning. Examining these functional areas with reference to the OFC's anatomical and neurophysiological properties, the authors suggest ways in which the OFC might contribute to obsessive-compulsive disorder.

Humans↗

Addiction as a computational process gone awry.

Addictive drugs have been hypothesized to access the same neurophysiological mechanisms as natural learning systems. These natural learning systems can be modeled through temporal-difference reinforcement learning (TDRL), which requires a reward-error signal that has been hypothesized to be carried by dopamine. TDRL learns to predict reward by driving that reward-error signal to zero. By adding a noncompensable drug-induced dopamine increase to a TDRL model, a computational model of addiction is constructed that over-selects actions leading to drug receipt. The model provides an explanation for important aspects of the addiction literature and provides a theoretic view-point with which to address other aspects.

Animals↗

Synthesis of nonlinear control surfaces by a layered associative search network.

An approach to solving nonlinear control problems is illustrated by means of a layered associative network composed of adaptive elements capable of reinforcement learning. The first layer adaptively develops a representation in terms of which the second layer can solve the problem linearly. The adaptive elements comprising the network employ a novel type of learning rule whose properties, we argue, are essential to the adaptive behavior of the layered network. The behavior of the network is illustrated by means of a spatial learning problem that requires the formation of nonlinear associations. We argue that this approach to nonlinearity can be extended to a large class of nonlinear control problems.

Animals↗