Hypothalamic substrates of reward.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Many behavioral tasks require goal-directed actions to obtain delayed reward. The prefrontal cortex appears to mediate many aspects of goal-directed decision making. This article presents a model of prefrontal cortex function emphasizing the influence of goal-related activity on the choice of the next motor output. The model can be interpreted in terms of key elements of Reinforcement Learning Theory. Different neocortical minicolumns represent distinct sensory input states and distinct motor output actions. The dynamics of each minicolumn include separate phases of encoding and retrieval. During encoding, strengthening of excitatory connections forms forward and reverse associations between each state, the following action, and a subsequent state, which may include reward. During retrieval, activity spreads from reward states throughout the network. The interaction of this spreading activity with a specific input state directs selection of the next appropriate action. Simulations demonstrate how these mechanisms can guide performance in a range of goal-directed tasks, and provide a functional framework for some of the neuronal responses previously observed in the medial prefrontal cortex during performance of spatial memory tasks in rats.
We address the problem of how to reinforce learning in ultracomplex environments, with huge state-spaces, where one must learn to exploit a compact structure of the problem domain. The approach we propose is to simulate the evolution of an artificial economy of computer programs. The economy is constructed based on two simple principles so as to assign credit to the individual programs for collaborating on problem solutions. We find empirically that starting from programs that are random computer code, we can develop systems that solve hard problems. In particular, our economy learned to solve almost all random Blocks World problems with goal stacks that are 200 blocks high. Competing methods solve such problems only up to goal stacks of at most 8 blocks. Our economy has also learned to unscramble about half a randomly scrambled Rubik's cube and to solve several commercially sold puzzles.
Experimental studies of reasoning and planned behavior have provided evidence that nervous systems use internal models to perform predictive motor control, imagery, inference, and planning. Classical (model-free) reinforcement learning approaches omit such a model; standard sensorimotor models account for forward and backward functions of sensorimotor dependencies but do not provide a proper neural representation on which to realize planning. We propose a sensorimotor map to represent such an internal model. The map learns a state representation similar to self-organizing maps but is inherently coupled to sensor and motor signals. Motor activations modulate the lateral connection strengths and thereby induce anticipatory shifts of the activity peak on the sensorimotor map. This mechanism encodes a model of the change of stimuli depending on the current motor activities. The activation dynamics on the map are derived from neural field models. An additional dynamic process on the sensorimotor map (derived from dynamic programming) realizes planning and emits corresponding goal-directed motor sequences, for instance, to navigate through a maze.
An important question in neuroevolution is how to gain an advantage from evolving neural network topologies along with weights. We present a method, NeuroEvolution of Augmenting Topologies (NEAT), which outperforms the best fixed-topology method on a challenging benchmark reinforcement learning task. We claim that the increased efficiency is due to (1) employing a principled method of crossover of different topologies, (2) protecting structural innovation using speciation, and (3) incrementally growing from minimal structure. We test this claim through a series of ablation studies that demonstrate that each component is necessary to the system as a whole and to each other. What results is significantly faster learning. NEAT is also an important contribution to GAs because it shows how it is possible for evolution to both optimize and complexify solutions simultaneously, offering the possibility of evolving increasingly complex solutions over generations, and strengthening the analogy with biological evolution.
This report describes the piloting mechanisms employed by honey bees during their final approach to a goal. Conceptually applying a bottom-up approach, we systematically varied the position, number and appearance landmarks associated with a rewarded target location within a large, homogenous flight tent. The flight behavior measured under various conditions is well explained with visuo-motor control loops that link perceived landmarks with appropriate turning responses. This view is consistent with the requirement of prolonged reinforcement learning for efficient goal navigation. A simple model is able to provide a comprehensive explanation for diverse flight patterns that range from convoluted searching behavior to highly idiosyncratic approaches, depending on the experimental context. Our results challenge the prevalent notion that honey bees employ image matching for visual guidance toward a goal site. Basic visuo-motor control loops may better meet the high demands for robust and fast flight control, which could serve as a powerful bio-mimetic design principle for micro-robotic aircraft.
This article presents an overview of those animal studies which so far have been performed with dedicated small animal positron emission tomographs in the field of the neurosciences. In vivo investigations focus on energy metabolism, perfusion and receptor/transporter binding in rat models of reinforcement, learning and memory, traumatic brain injury, epilepsy, depression, cardiovascular diseases--such as ischemia and focal stroke--and neurodegenerative disorders such as Alzheimer's, Parkinson's and Huntington's disease. In the majority of studies, important novel aspects arise from the fact that the investigators made use of an option inherent to in vivo studies, namely to conduct longitudinal investigations on the same animals. Relevant findings pertain to the relationship of brain metabolism/perfusion and the cholinergic system, the regulation state of dopamine receptors upon cocaine administration and withdrawal, the regulation state of dopamine receptors and transporters in animal models of Parkinson's and Huntington's disease, and potential treatments of progressive dopaminergic depletion with adenoviral vectors, embryonic grafts, stem cells and nerve growth factors.
Endocannabinoids are important mediators of short- and long-term synaptic plasticity, but the mechanisms of endocannabinoid release have not been studied extensively outside the hippocampus and cerebellum. Here, we examined the mechanisms of endocannabinoid-mediated long-term depression (eCB-LTD) in the dorsal striatum, a brain region critical for motor control and reinforcement learning. Unlike other cell types, strong depolarization of medium spiny neurons was not sufficient to yield detectable endocannabinoid release. However, when paired with postsynaptic depolarization sufficient to activate L-type calcium channels, activation of postsynaptic metabotropic glutamate receptors (mGluRs), either by high-frequency tetanic stimulation or an agonist, induced eCB-LTD. Pairing bursts of afferent stimulation with brief subthreshold membrane depolarizations that mimicked down-state to up-state transitions also induced eCB-LTD, which not only required activation of mGluRs and L-type calcium channels but also was bidirectionally modulated by dopamine D2 receptors. Consistent with network models, these results demonstrate that dopamine regulates the induction of a Hebbian form of long-term synaptic plasticity in the striatum. However, this gating of plasticity by dopamine is accomplished via an unexpected mechanism involving the regulation of mGluR-dependent endocannabinoid release.
Swarm intelligence inspired by the social behavior of ants boasts a number of attractive features, including adaptation, robustness and distributed, decentralized nature, which are well suited for routing in modern communication networks. This paper describes an adaptive swarm-based routing algorithm that increases convergence speed, reduces routing instabilities and oscillations by using a novel variation of reinforcement learning and a technique called momentum. Experiment on the dynamic network showed that adaptive swarm-based routing learns the optimum routing in terms of convergence speed and average packet latency.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
We used genetic algorithms to evolve populations of reinforcement learning (Q-learning) agents to play a repeated two-player symmetric coordination game under different risk conditions and found that evolution steered our simulated populations to the Pareto inefficient equilibrium under high-risk conditions and to the Pareto efficient equilibrium under low-risk conditions. Greater degrees of forgiveness and temporal discounting of future returns emerged in populations playing the low-risk game. Results demonstrate the utility of simulation to evolutionary psychology.
Explore the source record for details and available documents.
Long-lasting adaptations in the mesolimbic dopamine (DA) system in response to drugs of abuse likely mediate many of the behavioral changes that underlie addiction. Recent work suggests that long-term changes in synaptic strength at excitatory synapses in the two major components of this system, the nucleus accumbens (NAc) and ventral tegmental area, may be particularly important for the development of drug-induced sensitization, a process that may contribute to addiction, as well as for normal response-reinforcement learning. Using whole-cell patch-clamp recording techniques from in vitro slice preparations, we have examined the existence and basic mechanisms of long-term depression (LTD) at excitatory synapses on both GABAergic medium spiny neurons in the NAc and dopaminergic neurons in the midbrain. We find that both sets of synapses express LTD but that their basic triggering mechanisms differ. Furthermore, DA blocks the induction of LTD in the midbrain via activation of D2-like receptors but has minimal effects on LTD in the NAc. The existence of LTD in mesolimbic structures and its modulation by DA represent mechanisms that may contribute to the modifications of neural circuitry that mediate reward-related learning as well as the development of addiction.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
OBJECTIVE: Increased attention has been directed toward assessing and improving academic quality in athletic training education. The educational process has been assessed from a global level, but little is known about how athletic training students learn. The purpose of this investigation was to assess the learning styles of undergraduate athletic training students. DESIGN AND SETTING: Undergraduate students enrolled in a Committee on Accreditation of Allied Health Education Programs (CAAHEP)-accredited athletic training education program completed a learning styles inventory during a regularly scheduled athletic training class at the start of the spring semester. SUBJECTS: Twenty-seven student athletic trainers (age range, 19-30 yrs, mean age = 20.5 yrs) served as subjects. Sixteen subjects (7 male, 9 female) were in the first year of this 3-year program. Eleven subjects (7 male, 4 female) were second-year students. MEASUREMENTS: Learning style was assessed using the Productivity Environmental Preference Survey. RESULTS: Parametric and nonparametric one-way analyses of variance for each learning subscale by sex and by year in program revealed significant differences (P < .05) in light preferences for male and female students. There were also significant differences (P < .05) between first-and second-year students in preferences for afternoon learning activities. CONCLUSIONS: These findings suggest that undergraduate athletic training students function best as leamers in a well-lit leaming environment. The significance of aftemoon as the preferred time for learning reinforces the importance of the clinical setting in the introduction and mastery of skills. Athletic training educators and clinical instructors can use these results as they examine their teaching strategies and educational environments.
Knowledge of core content of nursing research and computer technology is increasingly required of baccalaureate students. This article describes a teaching model that incorporates computer technology content within a baccalaureate nursing research course. The model uses thirteen (13) linking concepts to reinforce learning about both content areas throughout the course.