PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5Linked to original sources

Phosphorylation changes following weakly reinforced learning and ACTH-induced memory consolidation for a weak learning experience.

The formation of a protein synthesis-dependent long-term memory stage in day-old chicks trained on a passive discriminated avoidance task has been shown to occur only with an adequate level of reinforcement, and is preceded by a significant change in the phosphorylation state of the forebrain synaptosomal membrane protein GAP43 protein. In the present study, it is shown that weakly reinforced training did not lead to formation of a long-term memory stage or to any change in phosphate incorporation into forebrain P2M protein bands. However, administration of ACTH immediately posttraining led to both the formation of the long-term memory stage and a preceding significant increase in the phosphorylation of GAP43. These findings are consistent with the view that a reinforcement-dependent neurohormone-mediated change to the phosphorylation of this synaptosomal membrane protein may be implicated in the triggering of long-term memory consolidation.

Animals↗

Is acetylcholine involved in memory consolidation of over-reinforced learning?

Systemic administration of anticholinergic drugs produces amnesia. To determine whether this effect can be prevented by increasing the magnitude of the learning experience, independent groups of rats were trained in passive avoidance, using a 3.0-mA footshock, and then injected with scopolamine (2, 4, 6, 8 or 12 mg/kg). When retention of the task was evaluated, a dose-dependent amnesic effect was found. When footshock intensity was increased to 6.0 and 9.0 mA, injections of 8 and 12 mg/kg of scopolamine did not produce memory impairments. These findings indicate that acetylcholine plays an important role in consolidation of passive avoidance, but it does not seem to be involved in memory processes when the magnitude of the negative reinforcer is increased.

Acetylcholine↗

The role of basal ganglia in reinforcement learning and imprinting in domestic chicks.

Effects of bilateral kainate lesions of telencephalic basal ganglia (lobus parolfactorius, LPO) were examined in domestic chicks. In the imprinting paradigm, where chicks learned to selectively approach a moving object without any explicitly associated reward, both the pre- and post-training lesions were without effects. On the other hand, in the water-reinforced pecking task, pre-training lesions of LPO severely impaired immediate reinforcement as well as formation of the association memory. However, post-training LPO lesions did not cause amnesia, and chicks selectively pecked at the reinforced color. The LPO could thus be involved specifically in the evaluation of present rewards and the instantaneous reinforcement of pecking, but not in the execution of selective behavior based on a memorized color cue.

Animals↗

Information processing, dimensionality reduction and reinforcement learning in the basal ganglia.

Modeling of the basal ganglia has played a major role in our understanding of this elusive group of nuclei. Models of the basal ganglia have undergone evolutionary and revolutionary changes over the last 20 years, as new research in the fields of anatomy, physiology and biochemistry of these nuclei has yielded new information. Early models dealt with a single pathway through the nuclei and focused on the nature of the processing performed within it, convergence of information versus parallel processing of information. Later, the Albin-DeLong "box-and-arrow" model characterized the inter-nuclei interaction as multiple pathways while maintaining a simplistic scalar representation of the nuclei themselves. This model made a breakthrough by providing key insights into the behavior of these nuclei in hypo- and hyper-kinetic movement disorders. The next generation of models elaborated the intra-nuclei interactions and focused on the role of the basal ganglia in action selection and sequence generation which form the most current consensus regarding basal ganglia function in both normal and pathological conditions. However, new findings challenge these models and point to a different neural network approach to information processing in the basal ganglia. Here, we take an in-depth look at the reinforcement driven dimensionality reduction (RDDR) model which postulates that the basal ganglia compress cortical information according to a reinforcement signal using optimal extraction methods. The model provides new insights and experimental predictions on the computational capacity of the basal ganglia and their role in health and disease.

Animals↗

Functional interaction of NMDA and group I metabotropic glutamate receptors in negatively reinforced learning in rats.

RATIONALE: The role of glutamatergic system in learning and memory has been extensively studied, and especially N-methyl-D: -aspartate (NMDA) receptors have been implicated in different learning and memory processes. Less is known, however, about group I metabotropic glutamate (mGlu) receptors in this field. Recent studies indicated that the coactivation of both NMDA and group I mGlu receptors is required for the induction of long-term potentiation (LTP) and learning. OBJECTIVE: The purpose of the study is to evaluate if there is a functional interaction between NMDA and group I mGlu receptors in two different models of aversive learning. METHODS: Effects of NMDA, mGlu1, and mGlu5 receptor antagonists on acquisition were tested after systemic coadministration of selected ineffective doses in passive avoidance (PA) and fear-potentiated startle (FPS). RESULTS: Interaction in aversive learning was investigated using selective antagonists: (3-ethyl-2-methyl-quinolin-6-yl)-(4-methoxy-cyclohexyl)-methanone methanesulfonate (EMQMCM) for mGlu1, [(2-methyl-1,3-thiazol-4-yl)ethynyl]pyridine (MTEP) for mGlu5, and (+)-5-methyl-10,11-dihydro-5H-dibenzocyclohepten-5,10-imine maleate [(+)MK-801] for NMDA receptors. In PA, the coapplication of MTEP at a dose of 5 mg/kg and (+)MK-801 at a dose of 0.1 mg/kg 30 min before training impaired the acquisition tested 24 h later. Similarly, EMQMCM (2.5 mg/kg) plus (+)MK-801 (0.1 mg/kg), given during the acquisition phase, blocked the acquisition of the PA response. In contrast, neither the combination of MTEP (1.25 mg/kg) nor EMQMCM (5 mg/kg) plus (+)MK-801 (0.05 mg/kg) was effective on the acquisition assessed in the FPS paradigm. CONCLUSION: The findings suggest differences in the interaction of the NMDA and mGlu group I receptor types in aversive instrumental conditioning vs conditioning to a discrete light cue.

Animals↗

A comparison of the effects of diazepam and scopolamine in two positively reinforced learning tasks.

In a helical maze scopolamine (0.5 and 1 mg/kg) significantly impaired the ability of rats to acquire a spatial learning task using reference memory. In contrast, diazepam (0.5-2 mg/kg) did not impair acquisition of this task and the only effect of diazepam (4 mg/kg) was likely to be secondary to sedative effects. Diazepam (0.5-4 mg/kg) did not impair 8-day retention of the helical maze. In a test of working and reference memory in which spatial processing was minimised, scopolamine (0.5 and 1 mg/kg) significantly impaired acquisition and increased the number of reference memory errors. Diazepam (1 and 4 mg/kg) did not impair acquisition of this task, but when a delay was interposed in the middle of a trial the diazepam-treated rats were slower to complete the task than the controls and made more errors of both working and reference memory. In contrast, when the rats were tested with a change of context, the diazepam-treated rats completed the task more quickly than the controls and made fewer errors of both working and reference memory.

Animals↗

Modular fuzzy-reinforcement learning approach with internal model capabilities for multiagent systems.

To date, many researchers have proposed various methods to improve the learning ability in multiagent systems. However, most of these studies are not appropriate to more complex multiagent learning problems because the state space of each learning agent grows exponentially in terms of the number of partners present in the environment. Modeling other learning agents present in the domain as part of the state of the environment is not a realistic approach. In this paper, we combine advantages of the modular approach, fuzzy logic and the internal model in a single novel multiagent system architecture. The architecture is based on a fuzzy modular approach whose rule base is partitioned into several different modules. Each module deals with a particular agent in the environment and maps the input fuzzy sets to the action Q-values; these represent the state space of each learning module and the action space, respectively. Each module also uses an internal model table to estimate actions of the other agents. Finally, we investigate the integration of a parallel update method with the proposed architecture. Experimental results obtained on two different environments of a well-known pursuit domain show the effectiveness and robustness of the proposed multiagent architecture and learning approach.

Journal Article↗

The ascending neuromodulatory systems in learning by reinforcement: comparing computational conjectures with experimental findings.

A central problem in cognitive neuroscience is how animals can manage to rapidly master complex sensorimotor tasks when the only sensory feedback they use to improve their performance is a simple reinforcing stimulus. Neural network theorists have constructed algorithms for reinforcement learning that can be used to solve a variety of biological problems and do not violate basic neurophysiological principles, in contrast to the back-propagation algorithm. A key assumption in these models is the existence of a reinforcement signal, which would be diffusively broadcast throughout one or several brain areas engaged in learning. This signal is further assumed to mediate up- and downward changes in synaptic efficacy by acting as a multiplicative factor in learning rules. The biological plausibility of these algorithms has been defended by the conjecture that the neuromodulators noradrenaline, acetylcholine or dopamine may form the neurochemical substrate of reinforcement signals. In this commentary, the predictions raised by this hypothesis are compared to anatomical, electrophysiological and behavioural findings. The experimental evidence does not support, and often argues against, a general reinforcement-encoding function of these neuromodulatory systems. Nevertheless, the broader concept of evaluative signalling between brain structures implied in learning appears to be reasonable and the available algorithms may open new avenues for constructing more realistic network architectures.

Adaptation, Psychological↗

Reinforced variability and operant learning.

Reinforcement of variability may help to explain operant learning. Three groups of rats were reinforced, in different phases, whenever the following target sequences of left (L) and right (R) lever presses occurred: LR, RLL, LLR, RRLR, RLLRL, and in Experiment 2, LLRRL. One group (variability [VAR]) was concurrently reinforced once per minute for sequence variations, a second group also once per minute but independently of variations, that is, for any sequences (ANY), and a control group (CON) received no additional reinforcers. The 3 groups learned the easiest targets equally. For the most difficult targets, CON animals' responding extinguished whereas both VAR and ANY responded at high rates. Only the VAR animals learned, however. Thus, concurrent reinforcers--contingent on variability or not--helped to maintain responding when difficult sequences were reinforced, but learning those sequences depended on reinforcement of variations.

Animals↗

Complementary roles of basal ganglia and cerebellum in learning and motor control.

The classical notion that the basal ganglia and the cerebellum are dedicated to motor control has been challenged by the accumulation of evidence revealing their involvement in non-motor, cognitive functions. From a computational viewpoint, it has been suggested that the cerebellum, the basal ganglia, and the cerebral cortex are specialized for different types of learning: namely, supervised learning, reinforcement learning and unsupervised learning, respectively. This idea of learning-oriented specialization is helpful in understanding the complementary roles of the basal ganglia and the cerebellum in motor control and cognitive functions.

Animals↗

[Learning and resistance to extinction in external-self dual reinforcement as compared with those in external reinforcement].

Learning and resistance to extinction in external-self dual reinforcement (ESR) were compared with those in external reinforcement (ER). delta ESR (delta: extinction process) was also compared with the corresponding process of other two conditions; 1) discontinuation of self reinforcement after external-self dual reinforcement (delta SR), 2) discontinuation of external reinforcement following the dual reinforcement (delta ER). Undergraduates (n = 58 in Experiment I, n = 74 in Experiment II) were randomly assigned to one of the four conditions. Subjects were given association learning in Exp. I and memory task in Exp. II, respectively. Those who responded successfully were then shifted to extinction session for testing. Effects of ESR on learning were same as those of ER in the two experiments. Resistance to extinction in Exp. I was significantly higher for delta SR and delta ESR as compared with delta ER and *ER (extinction switched from external reinforcement). Resistance for delta SR in Exp. II was significantly higher than for the remaining conditions. Based on these findings, high level of the resistance for delta ESR was discussed by attributing it to the internalization of ESR.

Adult↗

Nucleus accumbens core lesions retard instrumental learning and performance with delayed reinforcement in the rat.

BACKGROUND: Delays between actions and their outcomes severely hinder reinforcement learning systems, but little is known of the neural mechanism by which animals overcome this problem and bridge such delays. The nucleus accumbens core (AcbC), part of the ventral striatum, is required for normal preference for a large, delayed reward over a small, immediate reward (self-controlled choice) in rats, but the reason for this is unclear. We investigated the role of the AcbC in learning a free-operant instrumental response using delayed reinforcement, performance of a previously-learned response for delayed reinforcement, and assessment of the relative magnitudes of two different rewards. RESULTS: Groups of rats with excitotoxic or sham lesions of the AcbC acquired an instrumental response with different delays (0, 10, or 20 s) between the lever-press response and reinforcer delivery. A second (inactive) lever was also present, but responding on it was never reinforced. As expected, the delays retarded learning in normal rats. AcbC lesions did not hinder learning in the absence of delays, but AcbC-lesioned rats were impaired in learning when there was a delay, relative to sham-operated controls. All groups eventually acquired the response and discriminated the active lever from the inactive lever to some degree. Rats were subsequently trained to discriminate reinforcers of different magnitudes. AcbC-lesioned rats were more sensitive to differences in reinforcer magnitude than sham-operated controls, suggesting that the deficit in self-controlled choice previously observed in such rats was a consequence of reduced preference for delayed rewards relative to immediate rewards, not of reduced preference for large rewards relative to small rewards. AcbC lesions also impaired the performance of a previously-learned instrumental response in a delay-dependent fashion. CONCLUSIONS: These results demonstrate that the AcbC contributes to instrumental learning and performance by bridging delays between subjects' actions and the ensuing outcomes that reinforce behaviour.

Animals↗

[Interval effects of added sequences on reinforcement pattern learning in rats].

Three experiments examined how intervals of added sequences affected rat's learning of the reinforcement pattern. Animals were trained for runway performance corresponding to the series of one reinforcement trial (R) and two nonreinforcement trials (N) run with 30-s ITIs. The series was NNR, RNN, NRN for Experiments 1, 2, 3, respectively. Following this acquisition training, a second series of three nonreinforced trials with 30-s ITIs was added to the first series. Animals were assigned to two groups matched for performance level and they were given an added series with either long or short inter-session intervals. Subjects in Group S-ITI were given totally six trials with 30-s ITIs, while subjects in Group L-INT were given 30-min between the first and second series. Running speed for the first series differed with structure of the series (reinforcement pattern). The pattern of running speed for added series (Trials 4-6) was similar to that for original series (Trials 1-3) in Group L-INT, while running speed was kept at a low level for added series of nonreinforced trials in Group S-ITI. As is suggested by these findings, the events regularly occurred with equal ITIs can be remembered as one series, even when a new series of events is added to already experienced one with the same ITIs. However, when the second series is temporally separated from the first one by long intervals, the memory system may be reset so that events can be segregated as two sequences.

Animals↗

Testing computational models of dopamine and noradrenaline dysfunction in attention deficit/hyperactivity disorder.

We test our neurocomputational model of fronto-striatal dopamine (DA) and noradrenaline (NA) function for understanding cognitive and motivational deficits in attention deficit/hyperactivity disorder (ADHD). Our model predicts that low striatal DA levels in ADHD should lead to deficits in 'Go' learning from positive reinforcement, which should be alleviated by stimulant medications, as observed with DA manipulations in other populations. Indeed, while nonmedicated adult ADHD participants were impaired at both positive (Go) and negative (NoGo) reinforcement learning, only the former deficits were ameliorated by medication. We also found evidence for our model's extension of the same striatal DA mechanisms to working memory, via interactions with prefrontal cortex. In a modified AX-continuous performance task, ADHD participants showed reduced sensitivity to working memory contextual information, despite no global performance deficits, and were more susceptible to the influence of distractor stimuli presented during the delay. These effects were reversed with stimulant medications. Moreover, the tendency for medications to improve Go relative to NoGo reinforcement learning was predictive of their improvement in working memory in distracting conditions, suggestive of common DA mechanisms and supporting a unified account of DA function in ADHD. However, other ADHD effects such as erratic trial-to-trial switching and reaction time variability are not accounted for by model DA mechanisms, and are instead consistent with cortical noradrenergic dysfunction and associated computational models. Accordingly, putative NA deficits were correlated with each other and independent of putative DA-related deficits. Taken together, our results demonstrate the usefulness of computational approaches for understanding cognitive deficits in ADHD.

Adolescent↗