PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13Linked to original sources

Neural networks for continuous online learning and control.

This paper proposes a new hybrid neural network (NN) model that employs a multistage online learning process to solve the distributed control problem with an infinite horizon. Various techniques such as reinforcement learning and evolutionary algorithm are used to design the multistage online learning process. For this paper, the infinite horizon distributed control problem is implemented in the form of real-time distributed traffic signal control for intersections in a large-scale traffic network. The hybrid neural network model is used to design each of the local traffic signal controllers at the respective intersections. As the state of the traffic network changes due to random fluctuation of traffic volumes, the NN-based local controllers will need to adapt to the changing dynamics in order to provide effective traffic signal control and to prevent the traffic network from becoming overcongested. Such a problem is especially challenging if the local controllers are used for an infinite horizon problem where online learning has to take place continuously once the controllers are implemented into the traffic network. A comprehensive simulation model of a section of the Central Business District (CBD) of Singapore has been developed using PARAMICS microscopic simulation program. As the complexity of the simulation increases, results show that the hybrid NN model provides significant improvement in traffic conditions when evaluated against an existing traffic signal control algorithm as well as a new, continuously updated simultaneous perturbation stochastic approximation-based neural network (SPSA-NN). Using the hybrid NN model, the total mean delay of each vehicle has been reduced by 78% and the total mean stoppage time of each vehicle has been reduced by 84% compared to the existing traffic signal control algorithm. This shows the efficacy of the hybrid NN model in solving large-scale traffic signal control problem in a distributed manner. Also, it indicates the possibility of using the hybrid NN model for other applications that are similar in nature as the infinite horizon distributed control problem.

Algorithms↗

Making patient safety the focus: crisis resource management in the undergraduate curriculum.

BACKGROUND: This paper examines the role of high fidelity simulation and crisis resource management in bridging the gap between theory and practice. Patient safety is fundamental to healthcare professional practice and is a common goal for healthcare providers. It provides a focus to motivate practitioners. Patient safety issues are not a priority in undergraduate curricula. Raising the profile at this level is crucial to improving the safety and quality of healthcare delivery. This paper explores the role of simulation in providing a realistic, safe environment for participants with different levels of experience to manage evolving crises in the context of their work environment. METHODS: The Southern Health Simulation and Skills Centre uses a patient safety focus in delivering a specialised educational programme adapted from aviation to healthcare. The programme, crisis resource management, enables participants to consolidate knowledge, attitudes and skills to achieve a deeper understanding of how their performance impacts on patient safety and the quality of healthcare provided. Self-reported written evaluation data was collected from participants of three different courses at Southern Health. RESULTS: Participants consistently report that these courses offer unique learning experiences that address aspects of workplace learning in ways that have not previously been possible. A video-assisted reflective process powerfully reinforces learning. CONCLUSION: Crisis resource management courses demonstrate the value of simulation in bridging the gap between 'knowing' and 'doing' and keeping the focus on patient safety. Recommendations are made for ways in which the core elements of crisis resource management philosophy can influence the conceptualization of a new medical curriculum.

Clinical Competence↗

Memory consolidation of weak training experiences by hormonal treatments.

Day-old chicks trained on a single-trial passive discrimination avoidance task using a concentrated chemical aversant, methyl anthranilate (MeA), have been shown to exhibit three stages of memory processing; short-, intermediate- and long-term. A similar learning task with the aversant diluted to 20% in ethanol leads to short-and intermediate-term memory, but no long-term memory. Subcutaneous administration of selected doses of the stress-related hormones, noradrenaline, ACTH and vasopressin in close temporal proximity to the training trial, produced long-term memory in chicks trained on the weakly reinforced task, mimicking the outcome of strongly reinforced learning and of retraining with the weakly reinforced task reported previously. These effects are shown to be associated with the production of a nonenergy-dependent phase of the intermediate memory stage, postulated to be necessary for long-term memory consolidation.

2,4-Dinitrophenol↗

Passive and active avoidance behavior in the light-dark box test.

The temporal evolution of passive and active avoidance behaviors has been followed in rats, using the light-dark box test, by measuring step-through and exit latencies. The employed schedule consisted of three 7 day periods (free exploration, reinforced learning, forced extinction-retention). The data show clearly that the two learned behaviors are both rapidly established and exhibit significant differences only during extinction, active avoidance apparently depending on close temporal reinforcement. The diverse role of several behavioral and neurological mechanisms is hypothesized.

Animals↗

Idebenone improves learning and memory impairment induced by cholinergic or serotonergic dysfunction in rats.

The effects of idebenone, a cerebral metabolic enhancer, on learning and memory impairment in two rat models with central cholinergic or serotonergic dysfunction were investigated using positively reinforced learning tasks. A delayed alternation task using a T maze was employed to test the effect of idebenone on short-term memory impairment induced by a cholinergic antagonist, scopolamine. A correct response, defined as a turn toward the arm opposite to that in the forced run, was rewarded with food pellets. Scopolamine (0.2 and 0.5 mg/kg, i.p.) significantly decreased the correct responses to the chance level in the 60-s-delayed alternation task. The scopolamine (0.2 mg/kg, i.p.)-induced impairment of short-term memory was improved by idebenone (3-30 mg/kg, i.p.) or an acetylcholinesterase inhibitor, physostigmine (0.1 and 0.2 mg/kg, i.p.), administered simultaneously. The central serotonergic dysfunction model was produced by giving rats a diet deficient in tryptophan, a precursor of serotonin. The rats fed on a tryptophan-deficient diet (TDD) showed a slower learning process in the operant brightness discrimination task (mult V115 EXT) than did rats fed on a normal diet. Idebenone (60 mg/kg/day) admixed with the TDD decreased the number of lever-pressing responses emitted during the extinction periods. The percentage of correct responses was significantly higher in the idebenone-treated group than in the control TDD group. These results suggest that idebenone may improve both the impairment of short-term memory induced by a decreased cholinergic activity and the retardation of discrimination learning induced by central serotonergic dysfunction.

Acetylcholine↗

Effects of differential reinforcement expectancies on successive matching-to-sample performance in pigeons.

A series of experiments employed a symbolic variant of Konorski's delayed successive matching-to-sample task in order to determine whether differential reinforcement expectancies affect discriminative responding. One of two sample stimuli (S1 or S2) was followed, after a delay (0, 5, or 10 sec), by one of two test stimuli (T1 or T2). Pigeons' key pecking during test periods could produce food only on S1-T1 and S2-T2 (positive) trials; nonreinforcement invariably occurred on S1-T2 and S2-T1 (negative) trials. Differential reinforcement was scheduled by following the two positive trial sequences with different probabilities of reinforcement (.2 and 1.0); nondifferential reinforcement was scheduled by following the two positive trial sequences with a single, intermediate probability of reinforcement. (.6). Subjects given differential reinforcement acquired the conditional discriminaton more rapidly and reached higher terminal levels of performance than nondifferential controls (Experiment 1). Moreover, the magnitude of these differences increased as the delay between sample and test stimuli was lengthened. Reversing the probabilities of reinforcement in the differential problem produced a substantial and durable disruption of conditional discrimination performance (Experiment 2). The same general pattern of results was obtained when differential sample key pecking was eliminated (Experiment 3). These results can be parsimoniously interpreted by postulating the existence of learned reinforcement expectancies, and they detract from the merits of trace theory as a complete account of animal memory.

Animals↗

A shared system for learning serial and temporal structure of sensori-motor sequences? Evidence from simulation and human experiments.

This research investigates the influences of temporal structure on the representation of serial order. Experiments are performed in a neural network model of sequence learning and in human subjects. In the sequence learning model, a recurrent network of leaky integrator neurons encodes a succession of internal states that become associated, by reinforcement learning, with the correct sequential responses. First, the model is shown to learn a simple temporal discrimination task. The model is then exposed to two novel serial reaction time (SRT) experiments. In the standard SRT task (M.J. Nissen, P. Bullemer, Attentional requirements of learning: evidence from performance measures, Cogn. Psychol. 19 (1987) 1-32 [16]), reaction times for stimuli presented in a repeating sequence are reduced with respect to those for random stimuli, providing a measure of sequence learning. The novelty of the current experiments is that imbedded in the serial order of the sequences, there is a temporal structure of delays. The model is sensitive to both the serial structure and the temporal structure of the sequences. This observation is then confirmed in human subjects. These results demonstrate how a novel recurrent architecture encodes the interaction of temporal and serial structure and provide insight into related aspects of human sensori-motor sequence learning.

Analysis of Variance↗

Functional specificity of ventral striatal compartments in appetitive behaviors.

The nucleus accumbens and its associated circuitry subserve behaviors linked to natural or biological rewards, such as feeding, drinking, sex, exploration, and appetitive learning. We have investigated the functional role of neurotransmitter and intracellular transduction mechanisms in behaviors subserved by the core and shell subsystems within the accumbens. Local infusion of the selective NMDA antagonist, AP-5, into the accumbens core, but not the shell, completely blocked acquisition of a bar-press response for food in hungry rats. This effect was apparent only when infused during the early stages of learning. We have also recently shown that infusion of certain protein kinase inhibitors into the core also impairs learning in the same paradigm. These results suggest that plasticity-related mechanisms within the accumbens core, involving glutamate-linked intracellular second messengers, are important for response-reinforcement learning. In contrast to the core, which primarily connects to somatic motor output systems, the shell is more intimately linked to viscero-endocrine effector systems. We have shown that both AMPA and GABA receptors within the medial shell (but not the core) are critically involved in controlling the brain's feeding pathways, via activation of the lateral hypothalamus (LH). This effect is blocked by local inhibition of the LH in double-cannulae experiments and also strongly and selectively activates Fos expression in the LH. These results provide a newly emerging picture of the differentiated functions of this forebrain region and suggest an integrated role in the elaboration of adaptive motor actions.

Animals↗

Involvement of the rat anterior cingulate cortex in control of instrumental responses guided by reward expectancy.

The anterior cingulate cortex (ACC) plays a critical role in stimulus-reinforcement learning and reward-guided selection of actions. Here we conducted a series of experiments to further elucidate the role of the ACC in instrumental behavior involving effort-based decision-making and instrumental learning guided by reward-predictive stimuli. In Experiment 1, rats were trained on a cost-benefit T-maze task in which they could either choose to climb a barrier to obtain a high reward (four pellets) in one arm or a low reward (two pellets) in the other with no barrier present. In line with previous studies, our data reveal that rats with quinolinic acid lesions of the ACC selected the response involving less work and smaller reward. Experiment 2 demonstrates that breaking points of instrumental performance under a progressive ratio schedule were similar in sham-lesioned and ACC-lesioned rats. Thus, lesions of the ACC did not interfere with the effort a rat is willing to expend to obtain a specific reward in this test. In a subsequent task, we examined effort-based decision-making in a lever-press task where rats had the choice between pressing a lever to receive preferred food pellets under a progressive ratio schedule, or free feeding on a less preferred food, i.e. lab chow. Results show that sham- and ACC-lesioned animals had similar breaking points and ingested comparable amounts of less-preferred food. Together, the results of Experiment 1 and 2 suggest that the ACC plays a role in evaluating how much effort to expend for reward; however, the ACC is not necessary in all situations requiring an assessment of costs and benefits. In Experiment 3 we investigated learning and reversal learning of instrumental responses guided by reward predictive stimuli. A reaction time (RT) task demanding conditioned lever release was used in which the upcoming reward magnitude (five vs. one food pellet) was signalled in advance by discriminative visual stimuli. Results revealed that rats with ACC lesions were able to discriminate reward magnitude-predictive stimuli and to adapt instrumental behavior to reversed stimulus-reward magnitude contingencies. Thus, in a simple discrimination task as used here, the ACC appears not to be required to discriminate reward magnitude-predictive stimuli and to use the learned significance of the stimuli to guide instrumental behavior.

Animals↗

Racial differences in self-assessed health problems, depressive cognitions, and learned resourcefulness.

The purpose of this study was to investigate differences in self-assessed health problems, depressive cognitions, and learned resourcefulness in functionally independent older adults, including 30 Black and 30 White elders age 65 and over. Data were collected during structured face-to-face interviews. The two groups were similar in age, gender, education, and income. Results revealed that White elders reported more physical health problems than Blacks, but the types of problems were similar. While White elders reported more depressive cognitions, Black elders were significantly more resourceful. Most importantly, a strong association between depressive cognitions and learned resourcefulness was found for both Black and White elders. Findings highlight the importance of developing strategies to teach or reinforce learned resourcefulness in Black and White elders in order to prevent depression and promote mental health. This study recommends further investigation of racial differences in cognitive-behavior strategies constituting learned resourcefulness, using larger and randomly selected samples.

Adaptation, Psychological↗

Diabetes education for the family, patient and paramedical staff.

Diabetes education should fulfill definable objectives and be provided in an orderly way to match the child and family's ability and readiness to learn. The main aims of education are (1) gaining an understanding of diabetes (2) developing practical skills in care (3) acquiring attitudes of optimism and self-confidence (4) acquiring detailed knowledge of management and (5) developing the ability to make management decisions. Families can respond to education as they overcome their initial shock and grief. Their ability to learn is enhanced by professional support in the anxious task of assuming responsible care for their child. It is helpful to involve all the family and both parents so they can support each other, share responsibility and enjoy the satisfaction of contributing to their child's good health. The child after infancy, should participate in care as much as he is able, and is consistent with his developmental stage. People learn in different ways and it is helpful to have different teaching methods available: individual learning, group discussion, reading material, visual aids and seminars. For the older child and teenager, camps enhance self-esteem and reinforce learning and allow the family to profit by their own experience. Paramedical staff should meet regularly to maintain their competence and ensure that teaching is consistent.

Adolescent↗

Self-tuning Optimal Regulation of Respiratory Motor Output by Hebbian Covariance Learning.

The respiratory motor system is a specialized musculoskeletal system that is controlled by a small assembly of neuronal clusters in the brainstem. Its prime function is to maintain CO(2), O(2), and pH homeostasis in arterial circulation through the motor act of breathing. A longstanding dilemma is that during muscular exercise homeostatic regulation occurs automatically without any apparent feedback or feedforward signals, whereas the homeostasis is readily abolished by exogenous chemical challenge. Recently, it has been proposed that these seemingly incongruous behaviors of the respiratory controller may be a manifestation of self-tuning adaptive control. This hypothesis is supported in part by recent discoveries of various memory systems in the brainstem controller including short- and long-term potentiation and depression of synaptic transmission as well as the dramatic abolishment of homeostatic regulation in mutant mice with targeted genetic disruption of the NMDA receptors. In this paper, we propose a model of self-tuning homeostatic regulation based upon a synaptic adaptation rule-Hebbian covariance learning rule-that has been suggested to underlie many forms of learning and memory in the higher brain. We show that such an adaptation rule may be a useful neuronal substrate for reinforcement learning in optimization tasks. The model demonstrates how spontaneous oscillations and/or random-like fluctuations of neural activity may be exploited by the controller to adaptively regulate and optimize motor output. Such a self-tuning neural control paradigm, which operates without the need for any internal model of the external environment, is generally applicable to a class of steady-state optimal regulation problems with infinite time horizon. An interesting implication of the present results is that brain intelligent control can occur at a subconscious level without the need for voluntary intervention from the higher brain. Copyright 1996 Elsevier Science Ltd.

Journal Article↗

Banishing the homunculus: making working memory work.

The prefrontal cortex has long been thought to subserve both working memory and "executive" function, but the mechanistic basis of their integrated function has remained poorly understood, often amounting to a homunculus. This paper reviews the progress in our laboratory and others pursuing a long-term research agenda to deconstruct this homunculus by elucidating the precise computational and neural mechanisms underlying these phenomena. We outline six key functional demands underlying working memory, and then describe the current state of our computational model of the prefrontal cortex and associated systems in the basal ganglia (BG). The model, called PBWM (prefrontal cortex, basal ganglia working memory model), relies on actively maintained representations in the prefrontal cortex, which are dynamically updated/gated by the basal ganglia. It is capable of developing human-like performance largely on its own by taking advantage of powerful reinforcement learning mechanisms, based on the midbrain dopaminergic system and its activation via the basal ganglia and amygdala. These learning mechanisms enable the model to learn to control both itself and other brain areas in a strategic, task-appropriate manner. The model can learn challenging working memory tasks, and has been corroborated by several important empirical studies.

Amygdala↗

Dissociable roles of ventral and dorsal striatum in instrumental conditioning.

Instrumental conditioning studies how animals and humans choose actions appropriate to the affective structure of an environment. According to recent reinforcement learning models, two distinct components are involved: a "critic," which learns to predict future reward, and an "actor," which maintains information about the rewarding outcomes of actions to enable better ones to be chosen more frequently. We scanned human participants with functional magnetic resonance imaging while they engaged in instrumental conditioning. Our results suggest partly dissociable contributions of the ventral and dorsal striatum, with the former corresponding to the critic and the latter corresponding to the actor.

Adult↗

A unified model for perceptual learning.

Perceptual learning in adult humans and animals refers to improvements in sensory abilities after training. These improvements had been thought to occur only when attention is focused on the stimuli to be learned (task-relevant learning) but recent studies demonstrate performance improvements outside the focus of attention (task-irrelevant learning). Here, we propose a unified model that explains both task-relevant and task-irrelevant learning. The model suggests that long-term sensitivity enhancements to task-relevant or irrelevant stimuli occur as a result of timely interactions between diffused signals triggered by task performance and signals produced by stimulus presentation. The proposed mechanism uses multiple attentional and reinforcement systems that rely on different underlying neuromodulators. Our model provides insights into how neural modulators, attentional and reinforcement learning systems are related.

Animals↗

Right or wrong, familiar or novel in pictorial list discrimination learning.

The interaction between nonassociative learning (presentation frequencies) and associative learning (reinforcement rates) in stimulus discrimination performance was investigated. Subjects were taught to discriminate lists of visual pattern pairs. When they chose the stimulus designated as right they were symbolically rewarded and when they chose the stimulus designated as wrong they were symbolically penalised. Subjects first learned one list and then another list. For a "right" group the pairs of the second list consisted of right stimuli from the first list and of novel wrong stimuli. For a "wrong" group it was the other way round. The right group transferred some discriminatory performance from the first to the second list while the control and wrong groups initially only performed near chance with the second list. When the first list involved wrong stimuli presented twice as frequently as right stimuli, the wrong group exhibited a better transfer than the right group. In a final experiment subjects learned lists which consisted of frequent right stimuli paired with scarce wrong stimuli and frequent wrong stimuli paired with scarce right stimuli. In later test trials these stimuli were shown in new combinations and additionally combined with novel stimuli. Subjects preferred to choose the most rewarded stimuli and to avoid the most penalised stimuli when the test pairs included at least one frequent stimulus. With scarce/scarce or scarce/novel stimulus combinations they performed less well or even chose randomly. A simple mathematical model that ascribes stimulus choices to a Cartesian combination of stimulus frequency and stimulus value succeeds in matching all these results with satisfactory precision.

Adolescent↗