PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14Linked to original sources

Synaptic reentry reinforcement based network model for long-term memory consolidation.

The conversion of newly formed declarative memories into long-term memories is known to be dependent on the hippocampus. Recent experiments suggest that memory consolidation requires reactivation of the NMDA receptor in CA1 during the initial week(s) after training. This led to the hypothesis that the repeated post-learning reinforcement of synaptic modifications, termed synaptic reentry reinforcement (SRR), is essential for long-term memory consolidation and storage. Based on experimental observations, we have built a computational model to further illustrate and explore the effect of the SRR process on the formation of long-term memory. We show that SRR is capable of strengthening and maintaining memory traces despite inherent variability in the system due to such processes as the turnover of synaptic receptors and their associated signaling and structural proteins. Furthermore, we demonstrate that new rounds of synaptic modification triggered by memory reactivation, either during conscious recall or sleep, could lead to the selective consolidation of a subset of memory traces. Finally, we show why the SRR process in the hippocampus is required during the initial post-training weeks for synaptic reinforcement based memory consolidation in the cortex.

Animals↗

PD trivia: Making learning fun.

Nurses are educators. It is the aim of every educator that his or her teaching should translate into learning. Effective teaching is especially of importance in assuring that patients learn to perform their own peritoneal dialysis (PD). In facilitating an environment where learning can occur, making learning fun is the objective. It is with this mandate that PD Trivia was created. PD Trivia is an interactive game created to facilitate and reinforce learning. PD Trivia consists of 100 essential questions to making PD a success at home. Evaluations at the peritoneal dialysis clinic have revealed excellent quantitative and qualitative results of this simple but comprehensive teaching tool for effective learning of PD.

Adult↗

Modeling of autonomous problem solving process by dynamic construction of task models in multiple tasks environment.

Traditional reinforcement learning (RL) supposes a complex but single task to be solved. When a RL agent faces a task similar to a learned one, the agent must re-learn the task from the beginning because it doesn't reuse the past learned results. This is the problem of quick action learning, which is the foundation of decision making in the real world. In this paper, we suppose agents that can solve a set of tasks similar to each other in a multiple tasks environment, where we encounter various problems one after another, and propose a technique of action learning that can quickly solve similar tasks by reusing previously learned knowledge. In our method, a model-based RL uses a task model constructed by combining primitive local predictors for predicting task and environmental dynamics. To evaluate the proposed method, we performed a computer simulation using a simple ping-pong game with variations.

Animals↗

Patients with Parkinson's disease learn to control complex systems-an indication for intact implicit cognitive skill learning.

Implicit memory and learning mechanisms are composed of multiple processes and systems. Previous studies demonstrated a basal ganglia involvement in purely cognitive tasks that form stimulus response habits by reinforcement learning such as implicit classification learning. We will test the basal ganglia influence on two cognitive implicit tasks previously described by Berry and Broadbent, the sugar production task and the personal interaction task. Furthermore, we will investigate the relationship between certain aspects of an executive dysfunction and implicit learning. To this end, we have tested 22 Parkinsonian patients and 22 age-matched controls on two implicit cognitive tasks, in which participants learned to control a complex system. They interacted with the system by choosing an input value and obtaining an output that was related in a complex manner to the input. The objective was to reach and maintain a specific target value across trials (dynamic system learning). The two tasks followed the same underlying complex rule but had different surface appearances. Subsequently, participants performed an executive test battery including the Stroop test, verbal fluency and the Wisconsin card sorting test (WCST). The results demonstrate intact implicit learning in patients, despite an executive dysfunction in the Parkinsonian group. They lead to the conclusion that the basal ganglia system affected in Parkinson's disease does not contribute to the implicit acquisition of a new cognitive skill. Furthermore, the Parkinsonian patients were able to reach a specific goal in an implicit learning context despite impaired goal directed behaviour in the WCST, a classic test of executive functions. These results demonstrate a functional independence of implicit cognitive skill learning and certain aspects of executive functions.

Aged↗

A global bioheat model with self-tuning optimal regulation of body temperature using Hebbian feedback covariance learning.

In the lower brain, body temperature is continually being regulated almost flawlessly despite huge fluctuations in ambient and physiological conditions that constantly threaten the well-being of the body. The underlying control problem defining thermal homeostasis is one of great enormity: Many systems and sub-systems are involved in temperature regulation and physiological processes are intrinsically complex and intertwined. Thus the defining control system has to take into account the complications of nonlinearities, system uncertainties, delayed feedback loops as well as internal and external disturbances. In this paper, we propose a self-tuning adaptive thermal controller based upon Hebbian feedback covariance learning where the system is to be regulated continually to best suit its environment. This hypothesis is supported in part by postulations of the presence of adaptive optimization behavior in biological systems of certain organisms which face limited resources vital for survival. We demonstrate the use of Hebbian feedback covariance learning as a possible self-adaptive controller in body temperature regulation. The model postulates an important role of Hebbian covariance adaptation as a means of reinforcement learning in the thermal controller. The passive system is based on a simplified 2-node core and shell representation of the body, where global responses are captured. Model predictions are consistent with observed thermoregulatory responses to conditions of exercise and rest, and heat and cold stress. An important implication of the model is that optimal physiological behaviors arising from self-tuning adaptive regulation in the thermal controller may be responsible for the departure from homeostasis in abnormal states, e.g., fever. This was previously unexplained using the conventional "set-point" control theory.

Adaptation, Physiological↗

Early odor preference learning in the rat: bidirectional effects of cAMP response element-binding protein (CREB) and mutant CREB support a causal role for phosphorylated CREB.

Early odor preference learning in rats is associated with increases of phosphorylated CREB (pCREB) in mitral cells of the olfactory bulb. In the present study, herpes simplex virus expressing CREB (HSV-CREB) and dominant-negative mutant CREB (HSV-mCREB) have been injected into the bulb to assess a causal role for CREB and pCREB in this model. Odor paired with stroking or with the beta-adrenoceptor agonist isoproterenol produces odor approach 24 hr later. Isoproterenol-induced learning exhibits an inverted U curve dose-dependent learning relationship with both low and high doses failing to produce learning. pCREB increases have only been seen at the learning effective dose. In the present study, injection of an HSV vector expressing mutant CREB into the olfactory bulb prevented learning induced by stroking. Control HSV expressing LacZ was without effect. Expression of mutant CREB shifted the dose-learning curve for isoproterenol to the right such that a higher dose was required to induce learning. Expression of CREB shifted the dose-learning curve for isoproterenol to the left, with a lower dose now producing learning. As expected from this shift, CREB overexpression interfered with learning induced by stroking. When learning occurred, with either CREB or mutant CREB, pCREB was observed to be elevated relative to the nonlearning LacZ control groups. Unexpectedly, with odor plus stroking in the nonlearning CREB group, the level of pCREB was also higher than with odor plus stroking in LacZ controls that did learn. The data demonstrate a causal role for CREB and pCREB in early mammalian odor preference learning, reinforcing CREB as a "universal" memory molecule. They support evidence that CREB overexpression can be deleterious and suggest the hypothesis of an optimal pCREB window for learning.

Adrenergic beta-Agonists↗

Neuronal responses in the ventral striatum of the behaving macaque.

To analyse the functioning of the ventral striatum, the responses of more than 1,000 single neurons were recorded in a region which included the nucleus accumbens and olfactory tubercle in 5 macaque monkeys. While the monkeys performed visual discrimination and related feeding tasks, the different populations of neurons found included neurons which responded to novel visual stimuli; to reinforcement-related visual stimuli such as (for different neurons) food-related stimuli, aversive stimuli, or faces; to other visual stimuli; in relation to somatosensory stimulation and movement; or to cues which signalled the start of a task. The neurons with responses to reinforcing or novel visual stimuli may reflect the inputs to the ventral striatum from the amygdala and hippocampus, and are consistent with the hypothesis that the ventral striatum provides a route for learned reinforcing and novel visual stimuli to influence behaviour.

Amygdala↗

Specialization in multi-agent systems through learning.

Specialization is a common feature in animal societies that leads to an improvement in the fitness of the team members and to an increase in the resources obtained by the team. In this paper we propose a simple reinforcement learning approach to specialization in an artificial multi-agent system. The system is composed of homogeneous and non-communicating agents. Because there is no communication, the number of agents in the team can easily scale up. Agents have the same initial functionalities, but they learn to specialize and so cooperate to achieve a complex gathering task efficiently. Simulation experiments show how the multi-agent system specializes appropriately so as to reach optimal (or near-to-optimal) performance in unknown and changing environments.

Animal Communication↗

Space, time and dopamine.

In recent years, dopamine has emerged as a key neurotransmitter that is crucially involved in incentive motivation and reinforcement learning. Dopamine release is evoked by rewards. The extensive divergence of outputs from a small number of dopaminergic neurons suggests a spatially nonselective action of dopamine, but it reinforces the specific actions that led to reward. How is this achieved? We propose that the selectivity of dopamine effects is achieved by the timing of dopamine release in relation to the activity of glutamatergic synapses, rather than by spatial localization of the dopamine signal to specific synaptic contacts. The synaptic mechanisms of these actions are unknown but reduced levels of dopamine, for example in Parkinson's disease, leads to a paucity of behavioural output, whereas its excess production has been associated with psychiatric problems. Clearly, there are therapeutic imperatives that require a better understanding of how dopamine functions at a synaptic level.

Animals↗

Chaos in learning a simple two-person game.

We investigate the problem of learning to play the game of rock-paper-scissors. Each player attempts to improve her/his average score by adjusting the frequency of the three possible responses, using reinforcement learning. For the zero sum game the learning process displays Hamiltonian chaos. Thus, the learning trajectory can be simple or complex, depending on initial conditions. We also investigate the non-zero sum case and show that it can give rise to chaotic transients. This is, to our knowledge, the first demonstration of Hamiltonian chaos in learning a basic two-person game, extending earlier findings of chaotic attractors in dissipative systems. As we argue here, chaos provides an important self-consistency condition for determining when players will learn to behave as though they were fully rational. That chaos can occur in learning a simple game indicates one should use caution in assuming real people will learn to play a game according to a Nash equilibrium strategy.

Game Theory↗

A cautionary note on interpreting the effects of partial reinforcement on place learning performance in the water maze.

The effects of partial reinforcement on dry land and swimming pool place learning tasks have recently been compared and it has been suggested that they differ fundamentally [8]. That is, partial reinforcement impairs performance in the water maze, but not on dry land. However, other evidence suggests that partial reinforcement may have the opposite effect in the water maze, strengthening the accuracy and persistence of spatial responses. We discuss how the discrepancy may depend on 'levels' of negative reinforcement (e.g. escaping to a submerged platform before complete removal from the pool) and how experimental procedures may set up competitive contingencies that reinforce alternative behaviors. Finally, we consider data from past lesion studies and suggest ways to improve the design of future water maze experiments.

Animals↗

Opponent interactions between serotonin and dopamine.

Anatomical and pharmacological evidence suggests that the dorsal raphe serotonin system and the ventral tegmental and substantia nigra dopamine system may act as mutual opponents. In the light of the temporal difference model of the involvement of the dopamine system in reward learning, we consider three aspects of motivational opponency involving dopamine and serotonin. We suggest that a tonic serotonergic signal reports the long-run average reward rate as part of an average-case reinforcement learning model; that a tonic dopaminergic signal reports the long-run average punishment rate in a similar context; and finally speculate that a phasic serotonin signal might report an ongoing prediction error for future punishment.

Animals↗

The development and pilot testing of a multimedia CD-ROM for diabetes education.

The multimedia CD-ROM program, Take Charge of Diabetes, was found to be accurate, easy to use, and enjoyable by the clients and health professionals who completed the pilot study. Participants perceived an increase in knowledge after completing the five modules. Two of the participants verbally stated that the program clarified information for them and they wished they had had such a program when they were first diagnosed with diabetes. Further evaluation is needed to generalize the effect of the program on knowledge of diabetes because the pilot study was not designed to fully evaluate the effectiveness of the program on knowledge level or behavior change. Behavior change resulting in better control of blood sugar levels and hemoglobin A1c within normal range is the goal for diabetes education. The person who lives with diabetes must learn self-care methods. To accomplish that, the person must be able to comprehend the material presented. CAI programs provide an individualized, interactive, and interesting way to learn about diabetes and self-care, using visual effects and audio to support the written text. CAI can provide an element of excitement that is not available with other conventional methods. Providing prompt reinforcement of correct answers in quiz sections and including positive written messages can increase patients' self-confidence and self-esteem. Computer-assisted instruction is not intended to replace personal contact with physicians and diabetes educators, but rather complement this contact, reinforce learning, and possibly increase self-motivation to take charge of one's diabetes.

CD-ROM↗

Psychostimulant-induced behavioral sensitization depends on nicotinic receptor activation.

Animal studies have shown that nicotine and psychostimulant drugs (amphetamine and cocaine) share the property of inducing long-lasting behavioral and neurochemical sensitization, which is thought to contribute to their addictive properties. Neuroplasticity subserving learning and memory mechanisms is considered to be involved in psychostimulant-induced sensitization and addiction behavior. Because nicotinic receptors in the brain play a role in the storage of drug-related information underlying reinforcement learning, we evaluated the possibility that activation of central nicotinic receptors may underlie psychostimulant-induced sensitization. Repeated exposure of rats to nicotine profoundly enhanced the psychomotor effects of nicotine and amphetamine 3 weeks after nicotine pretreatment. Moreover, the nicotinic receptor antagonist mecamylamine completely blocked the induction, but not the long-term expression, of behavioral sensitization to amphetamine in amphetamine-pretreated rats. Mecamylamine also prevented the development of cocaine-induced behavioral sensitization. Behavioral sensitization induced by nicotine, amphetamine, or cocaine was associated with an increase in the electrically evoked release of [(3)H]dopamine from nucleus accumbens slices. Coadministration of mecamylamine during pretreatment with nicotine, amphetamine, or cocaine prevented the development of this long-term hyperreactivity of nucleus accumbens dopamine neurons. Similarly, the high-affinity non-alpha7 subtype nicotinic receptor antagonist dihydro-beta-erythroidine prevented the development of amphetamine-induced behavioral and neurochemical sensitization. These data indicate that nicotinic receptor activation (by endogenously released acetylcholine) is a common denominator initiating neuroplasticity involved in the development of amphetamine, as well as cocaine-induced sensitization.

Amphetamine↗

How the basal ganglia use parallel excitatory and inhibitory learning pathways to selectively respond to unexpected rewarding cues.

After classically conditioned learning, dopaminergic cells in the substantia nigra pars compacta (SNc) respond immediately to unexpected conditioned stimuli (CS) but omit formerly seen responses to expected unconditioned stimuli, notably rewards. These cells play an important role in reinforcement learning. A neural model explains the key neurophysiological properties of these cells before, during, and after conditioning, as well as related anatomical and neurophysiological data about the pedunculopontine tegmental nucleus (PPTN), lateral hypothalamus, ventral striatum, and striosomes. The model proposes how two parallel learning pathways from limbic cortex to the SNc, one devoted to excitatory conditioning (through the ventral striatum, ventral pallidum, and PPTN) and the other to adaptively timed inhibitory conditioning (through the striosomes), control SNc responses. The excitatory pathway generates CS-induced excitatory SNc dopamine bursts. The inhibitory pathway prevents dopamine bursts in response to predictable reward-related signals. When expected rewards are not received, striosomal inhibition of SNc that is unopposed by excitation results in a phasic drop in dopamine cell activity. The adaptively timed inhibitory learning uses an intracellular spectrum of timed responses that is proposed to be similar to adaptively timed cellular mechanisms in the hippocampus and cerebellum. These mechanisms are proposed to include metabotropic glutamate receptor-mediated Ca(2+) spikes that occur with different delays in striosomal cells. A dopaminergic burst in concert with a Ca(2+) spike is proposed to potentiate inhibitory learning. The model provides a biologically predictive alternative to temporal difference conditioning models and explains substantially more data than alternative models.

Animals↗

Differences in the effects of post-trial chlorpromazine, reserpine, and amphetamine on discrimination learning in rats.

Rats were trained to perform in discrimination learning reinforced by water for 6 days, and were intraperitoneally injected with chlorpromazine, reserpine, or d-amphetamine after each training session. Although chlorpromazine at the dose levels of 0.5 mg/kg or more injected immediately after training impaired learning, the drug did not affect learning when it was injected 60 min after training. Reserpine and amphetamine also impaired learning, but delaying the time intervals between training and injection to 60 min or more had no influence on this learning impairment. Post-trial chlorpromazine and amphetamine had no effect on, but reserpine decreased, motility in the subsequent training session. Chlorpromazine had no effect on water intake in the subsequent session, but reserpine and amphetamine decreased water intake at the dose levels that impaired learning. It was concluded that all three drugs impaired learning, but differed in their effects on learning; chlorpromazine impaired learning by a specific effect on learning itself; reserpine, by a non-specific effect on behavior due to a long acting sedation; and amphetamine, by an effect to decrease the motivation to drink water. The specific effect of chlorpromazine could be related to the hypothesis of "memory trace" synthesis.

Animals↗

Adaptive critic autopilot design of bank-to-turn missiles using fuzzy basis function networks.

A new adaptive critic autopilot design for bank-to-turn missiles is presented. In this paper, the architecture of adaptive critic learning scheme contains a fuzzy-basis-function-network based associative search element (ASE), which is employed to approximate nonlinear and complex functions of bank-to-turn missiles, and an adaptive critic element (ACE) generating the reinforcement signal to tune the associative search element. In the design of the adaptive critic autopilot, the control law receives signals from a fixed gain controller, an ASE and an adaptive robust element, which can eliminate approximation errors and disturbances. Traditional adaptive critic reinforcement learning is the problem faced by an agent that must learn behavior through trial-and-error interactions with a dynamic environment, however, the proposed tuning algorithm can significantly shorten the learning time by online tuning all parameters of fuzzy basis functions and weights of ASE and ACE. Moreover, the weight updating law derived from the Lyapunov stability theory is capable of guaranteeing both tracking performance and stability. Computer simulation results confirm the effectiveness of the proposed adaptive critic autopilot.

Journal Article↗

Autonomous learning based on cost assumptions: theoretical studies and experiments in robot control.

Autonomous learning techniques are based on experience acquisition. In most realistic applications, experience is time-consuming: it implies sensor reading, actuator control and algorithmic update, constrained by the learning system dynamics. The information crudeness upon which classical learning algorithms operate make such problems too difficult and unrealistic. Nonetheless, additional information for facilitating the learning process ideally should be embedded in such a way that the structural, well-studied characteristics of these fundamental algorithms are maintained. We investigate in this article a more general formulation of the Q-learning method that allows for a spreading of information derived from single updates towards a neighbourhood of the instantly visited state and converges to optimality. We show how this new formulation can be used as a mechanism to safely embed prior knowledge about the structure of the state space, and demonstrate it in a modified implementation of a reinforcement learning algorithm in a real robot navigation task.

Models, Theoretical↗