PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16Linked to original sources

Optimal spike-timing-dependent plasticity for precise action potential firing in supervised learning.

In timing-based neural codes, neurons have to emit action potentials at precise moments in time. We use a supervised learning paradigm to derive a synaptic update rule that optimizes by gradient ascent the likelihood of postsynaptic firing at one or several desired firing times. We find that the optimal strategy of up- and downregulating synaptic efficacies depends on the relative timing between presynaptic spike arrival and desired postsynaptic firing. If the presynaptic spike arrives before the desired postsynaptic spike timing, our optimal learning rule predicts that the synapse should become potentiated. The dependence of the potentiation on spike timing directly reflects the time course of an excitatory postsynaptic potential. However, our approach gives no unique reason for synaptic depression under reversed spike timing. In fact, the presence and amplitude of depression of synaptic efficacies for reversed spike timing depend on how constraints are implemented in the optimization problem. Two different constraints, control of postsynaptic rates and control of temporal locality, are studied. The relation of our results to spike-timing-dependent plasticity and reinforcement learning is discussed.

Action Potentials↗

The soft constraints hypothesis: a rational analysis approach to resource allocation for interactive behavior.

Soft constraints hypothesis (SCH) is a rational analysis approach that holds that the mixture of perceptual-motor and cognitive resources allocated for interactive behavior is adjusted based on temporal cost-benefit tradeoffs. Alternative approaches maintain that cognitive resources are in some sense protected or conserved in that greater amounts of perceptual-motor effort will be expended to conserve lesser amounts of cognitive effort. One alternative, the minimum memory hypothesis (MMH), holds that people favor strategies that minimize the use of memory. SCH is compared with MMH across 3 experiments and with predictions of an Ideal Performer Model that uses ACT-R's memory system in a reinforcement learning approach that maximizes expected utility by minimizing time. Model and data support the SCH view of resource allocation; at the under 1000-ms level of analysis, mixtures of cognitive and perceptual-motor resources are adjusted based on their cost-benefit tradeoffs for interactive behavior.

Cognition↗

Reverse replay of behavioural sequences in hippocampal place cells during the awake state.

The hippocampus has long been known to be involved in spatial navigational learning in rodents, and in memory for events in rodents, primates and humans. A unifying property of both navigation and event memory is a requirement for dealing with temporally sequenced information. Reactivation of temporally sequenced memories for previous behavioural experiences has been reported in sleep in rats. Here we report that sequential replay occurs in the rat hippocampus during awake periods immediately after spatial experience. This replay has a unique form, in which recent episodes of spatial experience are replayed in a temporally reversed order. This replay is suggestive of a role in the evaluation of event sequences in the manner of reinforcement learning models. We propose that such replay might constitute a general mechanism of learning and memory.

Action Potentials↗

AMPA/kainate, NMDA, and dopamine D1 receptor function in the nucleus accumbens core: a context-limited role in the encoding and consolidation of instrumental memory.

Neural integration of glutamate- and dopamine-coded signals within the nucleus accumbens (NAc) is a fundamental process governing cellular plasticity underlying reward-related learning. Intra-NAc core blockade of NMDA or D1 receptors in rats impairs instrumental learning (lever-pressing for sugar pellets), but it is not known during which phase of learning (acquisition or consolidation) these receptors are recruited, nor is it known what role AMPA/kainate receptors have in these processes. Here we show that pre-trial intra-NAc core administration of the NMDA, AMPA/KA, and D1 receptor antagonists AP-5 (1 microg/0.5 microL), LY293558 (0.01 or 0.1 microg/0.5 microL), and SCH23390 (1 microg/0.5 microL), respectively, impaired acquisition of a lever-pressing response, whereas post-trial administration left memory consolidation unaffected. An analysis of the microstructure of behavior while rats were under the influence of these drugs revealed that glutamatergic and dopaminergic signals contribute differentially to critical aspects of the initial, randomly emitted behaviors that enable reinforcement learning. Thus, glutamate and dopamine receptors are activated in a time-limited fashion-only being required while the animals are actively engaged in the learning context.

Animals↗

Internal models in sensorimotor integration: perspectives from adaptive control theory.

Internal models and adaptive controls are empirical and mathematical paradigms that have evolved separately to describe learning control processes in brain systems and engineering systems, respectively. This paper presents a comprehensive appraisal of the correlation between these paradigms with a view to forging a unified theoretical framework that may benefit both disciplines. It is suggested that the classic equilibrium-point theory of impedance control of arm movement is analogous to continuous gain-scheduling or high-gain adaptive control within or across movement trials, respectively, and that the recently proposed inverse internal model is akin to adaptive sliding control originally for robotic manipulator applications. Modular internal models' architecture for multiple motor tasks is a form of multi-model adaptive control. Stochastic methods, such as generalized predictive control, reinforcement learning, Bayesian learning and Hebbian feedback covariance learning, are reviewed and their possible relevance to motor control is discussed. Possible applicability of a Luenberger observer and an extended Kalman filter to state estimation problems-such as sensorimotor prediction or the resolution of vestibular sensory ambiguity-is also discussed. The important role played by vestibular system identification in postural control suggests an indirect adaptive control scheme whereby system states or parameters are explicitly estimated prior to the implementation of control. This interdisciplinary framework should facilitate the experimental elucidation of the mechanisms of internal models in sensorimotor systems and the reverse engineering of such neural mechanisms into novel brain-inspired adaptive control paradigms in future.

Adaptation, Physiological↗

Portacaval anastomosis attenuates the impairing effect of cyproheptadine on avoidance learning in rats--an involvement of the serotonergic system.

Portacaval-anastomized (PCA) rats were used to demonstrate the involvement of the serotonergic system in long-term memory formation. Significant increases in the concentration of 5-hydroxyindoleacetic acid, a metabolite of 5-hydroxytryptamine (5-HT), in all regions examined and the turnover rate of this indoleamine transmitter in the hippocampus, hypothalamus, midbrain and medulla oblongata were observed in PCA rats in comparison with sham-operated controls. Cyproheptadine, a 5-HT receptor blocking agent, impaired the retention of two-way avoidance learning reinforced by light stimuli when the drug was intraperitoneally injected immediately after the completion of training. PCA treatment attenuated the impairing effect of cyproheptadine. When cyproheptadine was injected 2 h after the completion of training, the correct response in the retention test period was not decreased. The present results suggest that memory formation is a time-requiring process and is mediated by the central serotonergic mechanism.

Animals↗

Superior water maze performance and increase in fear-related behavior in the endothelial nitric oxide synthase-deficient mouse together with monoamine changes in cerebellum and ventral striatum.

Nitric oxide (NO) has been implicated in the control of emotion, learning, and memory. We have examined endothelial NO synthase-deficient mice (eNOS-/-) in terms of habituation to an open field, elevated plus-maze behavior, Morris water maze performance, and changes in cerebral monoamines. In the open field, eNOS-/- animals were less active than wild-type controls but showed unimpaired habituation. In the plus-maze, an anxiogenic effect was observed. Proceeding from previous findings of deficits in hippocampal and neocortical long-term potentiation (LTP) in our eNOS-/- mice, we investigated whether these animals also express deficits in learning tasks that have been linked to hippocampal function and LTP. Unexpectedly, eNOS gene disruption led to accelerated place learning in the water maze. Furthermore, during long-term retention and reversal learning, eNOS-/- mice showed improved performance. In a cued version of the water maze task, eNOS-/- and control mice did not differ, implying that the superior performance of eNOS-/- animals on the former tasks cannot be attributed solely to differences in sensorimotor capacities. The neurochemical evaluation of the eNOS-/- mice revealed increases in the concentrations of the serotonin metabolite 5-HIAA in the cerebellum, together with an accelerated serotonin turnover in the frontal cortex. Furthermore, eNOS-/- mice had a higher dopamine turnover in the ventral striatum. These findings are discussed in terms of possible concomitant effects on physiological parameters, such as a decreased reactivity of GABAergic neurotransmission or changes in vascular functions, and effects on behavioral processes related to reinforcement, learning, and emotion.

3,4-Dihydroxyphenylacetic Acid↗

Development and implementation of the MTutor on-line tutorial system for diploma level research students.

This study describes the development, implementation and evaluation of MTutor, a web-based tutorial system, which was developed for students undertaking a diploma level research module. The aim of the study was to evaluate the feasibility of using a problem solving approach through an on-line system, to provide additional academic support for health care students undertaking a research module. The extent to which students found it a useful approach and resource for reinforcing learning and for developing their computer and information technology skills was also evaluated. The tutorial was developed around a single researchable problem related to communication skills and nurse-patient interactions. It was based on an experimental design and consisted of four sub-problems supported by a variety of text-based resources relevant to the sub-problems. A series of multiple-choice questions accompanied each sub-problem and an evaluation questionnaire at the end of the tutorial was used to obtain formal feedback on the system. The evaluation questionnaire revealed that students found the system to be easy to use and that the on-line resources provided sufficient information with which to answer the tutorial problems. Students felt that the tutorial reinforced the taught sessions in the module and that working through the problems enhanced their understanding of research design. Further piloting of the system is planned to obtain feedback from an additional cohort of students as well as lecturers teaching the diploma level research module.

Computer-Assisted Instruction↗

An integrate-and-fire model of prefrontal cortex neuronal activity during performance of goal-directed decision making.

The orbital frontal cortex appears to be involved in learning the rules of goal-directed behavior necessary to perform the correct actions based on perception to accomplish different tasks. The activity of orbitofrontal neurons changes dependent upon the specific task or goal involved, but the functional role of this activity in performance of specific tasks has not been fully determined. Here we present a model of prefrontal cortex function using networks of integrate-and-fire neurons arranged in minicolumns. This network model forms associations between representations of sensory input and motor actions, and uses these associations to guide goal-directed behavior. The selection of goal-directed actions involves convergence of the spread of activity from the goal representation with the spread of activity from the current state. This spiking network model provides a biological implementation of the action selection process used in reinforcement learning theory. The spiking activity shows properties similar to recordings of orbitofrontal neurons during task performance.

Action Potentials↗

The community-reinforcement approach.

The community-reinforcement approach (CRA) is an alcoholism treatment approach that aims to achieve abstinence by eliminating positive reinforcement for drinking and enhancing positive reinforcement for sobriety. CRA integrates several treatment components, including building the client's motivation to quit drinking, helping the client initiate sobriety, analyzing the client's drinking pattern, increasing positive reinforcement, learning new coping behaviors, and involving significant others in the recovery process. These components can be adjusted to the individual client's needs to achieve optimal treatment outcome. In addition, treatment outcome can be influenced by factors such as therapist style and initial treatment intensity. Several studies have provided evidence for CRA's effectiveness in achieving abstinence. Furthermore, CRA has been successfully integrated with a variety of other treatment approaches, such as family therapy and motivational interviewing, and has been tested in the treatment of other drug abuse.

Adaptation, Psychological↗

How visual stimuli activate dopaminergic neurons at short latency.

Unexpected, biologically salient stimuli elicit a short-latency, phasic response in midbrain dopaminergic (DA) neurons. Although this signal is important for reinforcement learning, the information it conveys to forebrain target structures remains uncertain. One way to decode the phasic DA signal would be to determine the perceptual properties of sensory inputs to DA neurons. After local disinhibition of the superior colliculus in anesthetized rats, DA neurons became visually responsive, whereas disinhibition of the visual cortex was ineffective. As the primary source of visual afferents, the limited processing capacities of the colliculus may constrain the visual information content of phasic DA responses.

Animals↗

Disruption of CB(1) receptor signaling impairs extinction of spatial memory in mice.

RATIONALE: A growing body of in vitro and in vivo evidence indicates that a central endocannabinoid system, consisting of CB(1) receptors and endogenous cannabinoids, modulates specific aspects of mnemonic processes. Previous research has demonstrated that either permanent or drug-induced disruption of CB(1) receptor signaling interferes with the extinction of a conditioned fear response. OBJECTIVES: In the present study, we evaluated whether the endocannabinoid system also plays a role in extinguishing learned escape behavior in a Morris water maze task. METHODS: CB(1) (-/-) mice and mice repeatedly treated with 3 mg/kg of the CB(1) receptor antagonist SR 141716 (Rimonabant) were trained to locate a hidden platform in the Morris water maze. Following acquisition, the platform was removed and subjects were assigned to either a massed (i.e., five consecutive sessions consisting of four 2-min trials/session) or a spaced (a single, 1-min trial every 2-4 weeks) extinction protocol. RESULTS: Strikingly, both 3 mg/kg SR 141716-treated mice and CB(1) (-/-) mice continued to return to the target location across all five trials in the spaced extinction procedure, while the control mice underwent extinction by the third or fourth trial. In contrast, both the 3-mg/kg SR 141716-treated and CB(1) (-/-) mice exhibited extinction in the massed extinction trial procedure. CONCLUSIONS: These findings indicate that disruption of CB(1) receptor signaling impairs extinction processes in the Morris water maze, thus lending further support to the hypothesis that the endocannabinoid system plays an integral role in the suppression of non-reinforced learned behaviors.

Animals↗

Dissociable systems for empathy.

Empathy is a lay term that is becoming increasingly used in the field of cognitive neuroscience. In this paper, it is argued that empathy is a loose collection of partially dissociable neurocognitive systems. Two forms of 'emotional' empathy were considered: First, responding to emotional expressions, particularly angry expressions, leading to response reversal. Secondly, responding to emotional expressions, particularly fearful and sad expressions, leading to stimulus-reinforcement learning. The implications of these forms of empathy for understanding specific psychiatric conditions are briefly considered.

Affect↗

Hybrid model building methodology using unsupervised fuzzy clustering and supervised neural networks.

This paper suggests a model building methodology for dealing with new processes. The methodology, called Hybrid Fuzzy Neural Networks (HFNN), combines unsupervised fuzzy clustering and supervised neural networks in order to create simple and flexible models. Fuzzy clustering was used to define relevant domains on the input space. Then, sets of multilayer perceptrons (MLP) were trained (one for each domain) to map input-output relations, creating, in the process, a set of specified sub-models. The estimated output of the model was obtained by fusing the different sub-model outputs weighted by their predicted possibilities. On-line reinforcement learning enabled improvement of the model. The determination of the optimal number of clusters is fundamental to the success of the HFNN approach. The effectiveness of several validity measures was compared to the generalization capability of the model and information criteria. The validity measures were tested with fermentation simulations and real fermentations of a yeast-like fungus, Aureobasidium pullulans. The results outline the criteria limitations. The learning capability of the HFNN was tested with the fermentation data. The results underline the advantages of HFNN over a single neural network.

Biotechnology↗

Adaptive fuzzy control of electrically stimulated muscles for arm movements.

A modified adaptive Takagi-Sugeno (TS) fuzzy logic controller (FLC) is proposed that allows a simulated elbow-like biomechanical system to accurately track sigmoidal and sinusoidal trajectories in the sagittal plane. The work is a first effort towards the implementation of a system to restore elbow movements in quadriplegics using functional neuromuscular stimulation. The single-joint musculo-skeletal system is composed of a co-contractable pair of electrically stimulated muscles; the muscle model accounts for the increase in fatigue during the tracking exercise. In the proposed controller structure, a reinforcement learning scheme is used to accomplish the parameter tuning, and the parameter projection algorithm guarantees the system stability during the adaptation process. The controller performance is evaluated using computer simulation experiments and compared with the performance achievable when a standard proportional-integrative-derivative (PID) controller is employed for the same application. The modified adaptive TSFLC outperforms the PID controller in all tested situations, with a clear-cut advantage in the case of high-frequency sinusoidal trajectories (angular frequencies spanning the interval 8-12 rad s-1). The standard controller suffers from a dramatic increase in root mean square (RMS) tracking error above the value at 8 rad s-1, e.g. ERMS > or = 0.013, whereas the correlation coefficient between the actual and desired trajectory falls almost to zero, starting from the value rho approximately equal to 0.97 at 8 rad s-1. On the other hand, the adaptive TSFLC yields ERMS < or = 0.015, with rho > or = 0.78, over the whole range of tested angular frequencies.

Arm↗

A model of the cerebellar pathways applied to the control of a single-joint robot arm actuated by McKibben artificial muscles.

This article describes an expanded version of a previously proposed motor control scheme, based on rules for combining sensory and motor signals within the central nervous system. Classical control elements of the previous cybernetic circuit were replaced by artificial neural network modules having an architecture based on the connectivity of the cerebellar cortex, and whose functioning is regulated by reinforcement learning. The resulting model was then applied to the motion control of a mechanical, single-joint robot arm actuated by two McKibben artificial muscles. Various biologically plausible learning schemes were studied using both simulations and experiments. After learning, the model was able to accurately pilot the movements of the robot arm, both in velocity and position.

Arm↗

Motor-maps, navigation and implicit space representation in the hippocampus.

Multiple sensory-motor maps located in the brainstem and the cortex are involved in spatial orientation. Guiding movements of eyes, head, neck and arms they provide an approximately linear relation between target distance and motor response. This involves especially the superior colliculus in the brainstem and the parietal cortex. There, the natural frame of reference follows from the retinal representation of the environment. A model of navigation is presented that is based on the modulation of activity in those sensory-motor maps. The actual mechanism chosen was gain-field modulation, a process of multimodal integration that has been demonstrated in the parietal cortex and superior colliculus, and was implemented as attraction to visual cues (colour). Dependent on the metric of the sensory-motor map, the relative attraction to these cues implemented as gain field modulation and their position define a fixed point attractor on the plane for locomotive behaviour. The actual implementation used Kohonen-networks in a variant of reinforcement learning that are well suited to generate such topographically organized sensory-motor maps with roughly linear visuo-motor response characteristics. In the following, it was investigated how such an implicit coding of target positions by gain-field parameters might be represented in the hippocampus formation and under what conditions a direction-invariant space representation can arise from such retinotopic representations of multiple cues. Information about the orientation in the plane--as could be provided by head direction cells--appeared to be necessary for unambiguous space representation in our model in agreement with physiological experiments. With this information, Gauss-shaped "place-cells" could be generated, however, the representation of the spatial environment was repetitive and clustered and single cells were always tuned to the gain-field parameters as well.

Algorithms↗

Contribution of the Roman strains of rats to the elaboration of animal models of memory.

Performances of male rats of the Roman High (RHA)- and Roman Low (RLA)-Avoidance strains were compared along four essential dimensions: working memory, reference memory, spontaneous locomotion and avoidance conditioning. Performances of RLA and RHA rats were significantly different for each dimension. As constantly reported, RHA rats were by far superior to RLA in avoidance conditioning. They had also higher levels of locomotor activity. On the opposite, RLA performed better than RHA in an appetitive working memory task, the delayed reinforced alternation, and were also superior in an appetitive reference memory task, the 5-unit linear maze. These results confirm the fact that RLA rats may acquire positively reinforced learning more rapidly than RHA rats and that the differences in active avoidance behavior between the two strains depend more on differential freezing behavior than on learning or memory capacities. Beyond the problem of the characterization of the Roman strains, these data might give indications on the relationships between behavioral tests widely used in rats, and on their use as memory models.

Animals↗