PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7Linked to original sources

Transient calcium and dopamine increase PKA activity and DARPP-32 phosphorylation.

Reinforcement learning theorizes that strengthening of synaptic connections in medium spiny neurons of the striatum occurs when glutamatergic input (from cortex) and dopaminergic input (from substantia nigra) are received simultaneously. Subsequent to learning, medium spiny neurons with strengthened synapses are more likely to fire in response to cortical input alone. This synaptic plasticity is produced by phosphorylation of AMPA receptors, caused by phosphorylation of various signalling molecules. A key signalling molecule is the phosphoprotein DARPP-32, highly expressed in striatal medium spiny neurons. DARPP-32 is regulated by several neurotransmitters through a complex network of intracellular signalling pathways involving cAMP (increased through dopamine stimulation) and calcium (increased through glutamate stimulation). Since DARPP-32 controls several kinases and phosphatases involved in striatal synaptic plasticity, understanding the interactions between cAMP and calcium, in particular the effect of transient stimuli on DARPP-32 phosphorylation, has major implications for understanding reinforcement learning. We developed a computer model of the biochemical reaction pathways involved in the phosphorylation of DARPP-32 on Thr34 and Thr75. Ordinary differential equations describing the biochemical reactions were implemented in a single compartment model using the software XPPAUT. Reaction rate constants were obtained from the biochemical literature. The first set of simulations using sustained elevations of dopamine and calcium produced phosphorylation levels of DARPP-32 similar to that measured experimentally, thereby validating the model. The second set of simulations, using the validated model, showed that transient dopamine elevations increased the phosphorylation of Thr34 as expected, but transient calcium elevations also increased the phosphorylation of Thr34, contrary to what is believed. When transient calcium and dopamine stimuli were paired, PKA activation and Thr34 phosphorylation increased compared with dopamine alone. This result, which is robust to variation in model parameters, supports reinforcement learning theories in which activity-dependent long-term synaptic plasticity requires paired glutamate and dopamine inputs.

Calcium↗

Incentive learning following reinforcer devaluation is not conditional upon the motivational state during re-exposure.

Three experiments analysed the effect of re-exposure to the reinforcer following aversion conditioning on instrumental performance. In the first experiment, groups of hungry and thirsty rats were trained to press a lever for sucrose, which was then followed by a single injection of lithium chloride (LiCl). On the following day, half the animals in each motivational condition received re-exposure to the sucrose solution; the remaining animals were not re-exposed. In a subsequent extinction test animals that had received re-exposure to the sucrose pressed less than animals that were not re-exposed. Moreover, the effect of re-exposure to the sucrose solution was similar following training under hunger and thirst. In the remaining studies, animals were trained to lever-press for sucrose while either hungry or thirsty. They were then injected with LiCl and re-exposed to the sucrose while either hungry or thirsty, i.e. in the same or different motivational state employed during training, or they were not re-exposed. Lever pressing was then tested in extinction in the training motivational state. As in the first experiment, re-exposure to the reinforcer after aversion conditioning enhanced the magnitude of the reinforcer devaluation effect. More importantly, re-exposure to the sucrose produced a comparable effect on instrumental performance, whether re-exposure was given under the same or different motivational state to that employed during training. These results suggest that the instrumental reinforcer devaluation effect depends upon a process of incentive learning, but that this process is not conditional upon the current motivational state of the animal.

Animals↗

Personality, reinforcement and learning.

Conflict in predictions resulting from Eysenck's (1957) and Gray's (1970) theoretical formulations on personality and conditioning were tested at the behavioral level. Given conditions which do not produce over-arousal, it would be predicted from Eysenck's position that Introverts would condition better than Extraverts. From Gray's formulation it would follow that Introverts condition better if negative reinforcement is used and Extraverts condition better if positive reinforcement is used. The two opposing predictions were tested in pursuit rotor learning by either positively or negatively reinforcing the hit/miss dimension of performance by 166 males aged 14 to 15 yr. The results gave support to Gray's position but if over-arousal is assumed Eysenck's position is tenable.

Adolescent↗

A more biologically plausible learning rule for neural networks.

Many recent studies have used artificial neural network algorithms to model how the brain might process information. However, back-propagation learning, the method that is generally used to train these networks, is distinctly "unbiological." We describe here a more biologically plausible learning rule, using reinforcement learning, which we have applied to the problem of how area 7a in the posterior parietal cortex of monkeys might represent visual space in head-centered coordinates. The network behaves similarly to networks trained by using back-propagation and to neurons recorded in area 7a. These results show that a neural network does not require back propagation to acquire biologically interesting properties.

Animals↗

The primate amygdala represents the positive and negative value of visual stimuli during learning.

Visual stimuli can acquire positive or negative value through their association with rewards and punishments, a process called reinforcement learning. Although we now know a great deal about how the brain analyses visual information, we know little about how visual representations become linked with values. To study this process, we turned to the amygdala, a brain structure implicated in reinforcement learning. We recorded the activity of individual amygdala neurons in monkeys while abstract images acquired either positive or negative value through conditioning. After monkeys had learned the initial associations, we reversed image value assignments. We examined neural responses in relation to these reversals in order to estimate the relative contribution to neural activity of the sensory properties of images and their conditioned values. Here we show that changes in the values of images modulate neural activity, and that this modulation occurs rapidly enough to account for, and correlates with, monkeys' learning. Furthermore, distinct populations of neurons encode the positive and negative values of visual stimuli. Behavioural and physiological responses to visual stimuli may therefore be based in part on the plastic representation of value provided by the amygdala.

Amygdala↗

The primate amygdala: Neuronal representations of the viscosity, fat texture, temperature, grittiness and taste of foods.

The primate amygdala is implicated in the control of behavioral responses to foods and in stimulus-reinforcement learning, but only its taste representation of oral stimuli has been investigated previously. Of 1416 macaque amygdala neurons recorded, 44 (3.1%) responded to oral stimuli. Of the 44 orally responsive neurons, 17 (39%) represent the viscosity of oral stimuli, tested using carboxymethyl-cellulose in the range 1-10,000 cP. Two neurons (5%) responded to fat in the mouth by encoding its texture (shown by the responses of these neurons to a range of fats, and also to non-fat oils such as silicone oil ((Si(CH(3))(2)O)(n)) and mineral oil (pure hydrocarbon), but no or small responses to the cellulose viscosity series or to the fatty acids linoleic acid and lauric acid). Of the 44 neurons, three (7%) responded to gritty texture (produced by microspheres suspended in cellulose). Eighteen neurons (41%) responded to the temperature of liquid in the mouth. Some amygdala neurons responded to capsaicin, and some to fatty acids (but not to fats in the mouth). Some amygdala neurons respond to taste, texture and temperature unimodally, but others combine these inputs. These results provide fundamental evidence about the information channels used to represent the texture and flavor of food in a part of the brain important in appetitive responses to food and in learning associations to reinforcing oral stimuli, and are relevant to understanding the physiological and pathophysiological processes related to food intake, food selection, and the effects of variety of food texture in combination with taste and other inputs on food intake.

Action Potentials↗