Asymmetric-reinforcing events in probability learning.
Explore the source record for details and available documents.
SEARCH · PubMed Health
Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Explore the source record for details and available documents.
Traditionally, addiction research in neuroscience has focused on mechanisms involving dopamine and endogenous opioids. More recently, it has been realized that glutamate also plays a central role in processes underlying the development and maintenance of addiction. These processes include reinforcement, sensitization, habit learning and reinforcement learning, context conditioning, craving and relapse. In the past few years, some major advances have been made in the understanding of how glutamate acts and interacts with other transmitters (in particular, dopamine) in the context of processes underlying addiction. It appears that while many actions of glutamate derive their importance from a stimulatory interaction with the dopaminergic system, there are some glutamatergic mechanisms that contribute to addiction independent of dopaminergic systems. Among those, context-specific aspects of behavioral determinants (ie control over behavior by conditioned stimuli) appear to depend heavily on glutamatergic transmission. A better understanding of the underlying mechanisms might open new avenues to the treatment of addiction, in particular regarding relapse prevention.
BACKGROUND: Medical students participate in a longitudinal (3-year) primary care preceptorship to assist them in developing skills in interviewing and examining patients in an ambulatory care setting. PURPOSE: To identify from a student's perspective important context and process issues in a longitudinal preceptorship. METHODS: The investigators used an "editing" style of analysis to identify significant themes across 24 medical student focus groups held between October 1995 and December 1997. RESULTS: Significant themes emerged from the data analysis that describe important features of what makes the preceptorship work for students. The main themes are active teaching, active learning, a trusting relationship, sufficient time, and a shared understanding of preceptorship objectives. The potential benefits to students in an enhanced learning environment are comfort, confidence, responsibility, skills, knowledge, reinforcement, learning opportunities, teaching opportunities, and models for practice. CONCLUSIONS: We offer recommendations for enhancing longitudinal preceptorships for preceptors, students, and leaders in medical education.
Explore the source record for details and available documents.
This study investigated how the simulated response of dopamine neurons to reward-related stimuli could be used as reinforcement signal for learning a spatial delayed response task. Spatial delayed response tasks assess the functions of frontal cortex and basal ganglia in short-term memory, movement preparation and expectation of environmental events. In these tasks, a stimulus appears for a short period at a particular location, and after a delay the subject moves to the location indicated. Dopamine neurons are activated by unpredicted rewards and reward-predicting stimuli, are not influenced by fully predicted rewards, and are depressed by omitted rewards. Thus, they appear to report an error in the prediction of reward, which is the crucial reinforcement term in formal learning theories. Theoretical studies on reinforcement learning have shown that signals similar to dopamine responses can be used as effective teaching signals for learning. A neural network model implementing the temporal difference algorithm was trained to perform a simulated spatial delayed response task. The reinforcement signal was modeled according to the basic characteristics of dopamine responses to novel stimuli, primary rewards and reward-predicting stimuli. A Critic component analogous to dopamine neurons computed a temporal error in the prediction of reinforcement and emitted this signal to an Actor component which mediated the behavioral output. The spatial delayed response task was learned via two subtasks introducing spatial choices and temporal delays, in the same manner as monkeys in the laboratory. In all three tasks, the reinforcement signal of the Critic developed in a similar manner to the responses of natural dopamine neurons in comparable learning situations, and the learning curves of the Actor replicated the progress of learning observed in the animals. Several manipulations demonstrated further the efficacy of the particular characteristics of the dopamine-like reinforcement signal. Omission of reward induced a phasic reduction of the reinforcement signal at the time of the reward and led to extinction of learned actions. A reinforcement signal without prediction error resulted in impaired learning because of perseverative errors. Loss of learned behavior was seen with sustained reductions of the reinforcement signal, a situation in general comparable to the loss of dopamine innervation in Parkinsonian patients and experimentally lesioned animals. The striking similarities in teaching signals and learning behavior between the computational and biological results suggest that dopamine-like reward responses may serve as effective teaching signals for learning behavioral tasks that are typical for primate cognitive behavior, such as spatial delayed responding.
Heinrich's (1931) classical study implies that most industrial accidents can be characterized as a probabilistic result of human error. The present research quantifies Heinrich's observation and compares four descriptive models of decision making in the abstracted setting. The suggested quantification utilizes signal detection theory (Green & Swets, 1966). It shows that Heinrich's observation can be described as a probabilistic signal detection task. In a controlled experiment, 90 decision makers participated in 600 trials of six safety games. Each safety game was a numerical example of the probabilistic SDT abstraction of Heinrich's proposition. Three games were designed under a frame of gain to represent perception of safe choice as costless, while the other three were designed under a frame of loss to represent perception of safe choice as costly. Probabilistic penalty for Miss was given at three different levels (1, .5, .1). The results showed that decisions tended initially to be risky and that experience led to safer behavior. As the probability of being penalized was lowered decisions became riskier and the learning process was impaired. The results support the cutoff reinforcement learning model suggested by Erev et al. (1995). The hill-climbing learning model (Busemeyer & Myung, 1992) was partially supported. Theoretical and practical implications are discussed. Copyright 1998 Academic Press.
Symbiosis is the phenomenon in which organisms of different species live together in close association, resulting in a raised level of fitness for one or more of the organisms. Symbiogenesis is the name given to the process by which symbiotic partners combine and unify, that is, become genetically linked, giving rise to new morphologies and physiologies evolutionarily more advanced than their constituents. The importance of this process in the evolution of complexity is now well established. Learning classifier systems are a machine learning technique that uses both evolutionary computing techniques and reinforcement learning to develop a population of cooperative rules to solve a given task. In this article we examine the use of symbiogenesis within the classifier system rule base to improve their performance. Results show that incorporating simple rule linkage does not give any benefits. The concept of (temporal) encapsulation is then added to the symbiotic rules and shown to improve performance in ambiguous/non-Markov environments.
It has long been known that in some relatively simple reinforcement learning tasks traditional strength-based classifier systems will adapt poorly and show poor generalisation. In contrast, the more recent accuracy-based XCS, appears both to adapt and generalise well. In this work, we attribute the difference to what we call strong over general and fit over general rules. We begin by developing a taxonomy of rule types and considering the conditions under which they may occur. In order to do so an extreme simplification of the classifier system is made, which forces us toward qualitative rather than quantitative analysis. We begin with the basics, considering definitions for correct and incorrect actions, and then correct, incorrect, and overgeneral rules for both strength and accuracy-based fitness. The concept of strong overgeneral rules, which we claim are the Achilles' heel of strength-based classifier systems, are then analysed. It is shown that strong overgenerals depend on what we call biases in the reward function (or, in sequential tasks, the value function). We distinguish between strong and fit overgeneral rules, and show that although strong overgenerals are fit in a strength-based system called SB-XCS, they are not in XCS. Next we show how to design fit overgeneral rules for XCS (but not SB-XCS), by introducing biases in the variance of the reward function, and thus that each system has its own weakness. Finally, we give some consideration to the prevalence of reward and variance function bias, and note that non-trivial sequential tasks have highly biased value functions.
The nigrostriatal dopamine system has long been regarded to play an essential role in motor mechanisms, since selective damage to this system results in the severe motor disturbances known as Parkinson's syndromes. The mechanisms how the nigrostriatal dopamine system plays a part is still not clear. Based on recording of single neuron activity in the primate striatum and selective destruction of the nigrostriatal dopamine system, we propose that this system is crucially involved in behavioral learning through providing reinforcement signals to the striatum during behavioral learning. The way of involvement in learning seems to be consistent with the reinforcement learning rule.
Explore the source record for details and available documents.
Explore the source record for details and available documents.
During the learning of instrumental tasks, rats are usually fasted to increase reinforced learning. However, fasting produces several undesirable side effects. The aim of this study was to test the hypothesis that control rats, i.e. full-fed and group-reared rats, will learn an autoshaping task to the same level as fasted or singly-reared rats. The interaction between fasting and single-rearing of rats was also tested. Results showed that control rats and fasted rats acquired the autoshaping task similarly, independently of rearing condition or gender. However, fasted or singly-reared rats produced fear-like behaviour, since male rats group-reared and fasted (85% body/wt, P <0.05), male rats singly-reared (full fed, P <0.05; 12 h fasted, P <0.05; 85% body/wt, P <0.05), female rats group-reared (12 h fasted, P <0.05; 85% body/wt, P <0.05) and female rats singly reared (full fed, P <0.05; 12 h fasted, P <0.05; 85% body/wt, P <0.05) displayed reduced amounts of time exploring the open arms of the elevated plus-maze. In conclusion, control rats learned the autoshaping task to the same level as fasted or singly-reared rats. However, fasting or single-rearing produced fear-like behaviour. Thus, the training of control rats in autoshaping tasks may be an option that improves animal welfare.
The assumption that people possess a repertoire of strategies to solve the inference problems they face has been raised repeatedly. However, a computational model specifying how people select strategies from their repertoire is still lacking. The proposed strategy selection learning (SSL) theory predicts a strategy selection process on the basis of reinforcement learning. The theory assumes that individuals develop subjective expectations for the strategies they have and select strategies proportional to their expectations, which are then updated on the basis of subsequent experience. The learning assumption was supported in 4 experimental studies. Participants substantially improved their inferences through feedback. In all 4 studies, the best-performing strategy from the participants' repertoires most accurately predicted the inferences after sufficient learning opportunities. When testing SSL against 3 models representing extensions of SSL and against an exemplar model assuming a memory-based inference process, the authors found that SSL predicted the inferences most accurately.
Perceptual decisions are often made in complex social settings in which distinct observers can affect each other. To address such situations, I. Erev, D. Gopher, R. Itkin, and Y. Greenshpan (1995) proposed a formal extension of signal-detection theory and a descriptive modification of the extended theory. The current article presents 2 experiments that were designed to test these models in the context of repeated 2-person perceptual safety games. In both experiments, pairs of participants performed a simulation of an industrial-production process under distinct payoff rules. Each participant had to try to produce as much as possible while avoiding costly accidents. In line with the descriptive model's predictions, the results showed a slow adjustment to the incentive structure that can be approximated by a reinforcement learning process among different perceptual cutoff strategies. Providing players with prior information about the game had an initial effect but did not alter the pattern of the results.
The cross-entropy method is an efficient and general optimization algorithm. However, its applicability in reinforcement learning (RL) seems to be limited because it often converges to suboptimal policies. We apply noise for preventing early convergence of the cross-entropy method, using Tetris, a computer game, for demonstration. The resulting policy outperforms previous RL algorithms by almost two orders of magnitude.
Explore the source record for details and available documents.
Conditional visuo-motor learning consists in learning by trial and error to associate visual cues with correct motor responses, that have no direct link. Converging evidence supports the role of a large brain network in this type of learning, including the prefrontal and the premotor cortex, the basal ganglia BG and the hippocampus. In this paper we focus on the role of a major structure of the BG, the striatum. We first present behavioral results and electrophysiological data recorded from this structure in monkeys engaged in learning new visuo-motor associations. Visual stimuli were presented on a video screen and the animals had to learn, by trial and error, to select the correct movement of a joystick, in order to receive a liquid reward. Behavioral results revealed that the monkeys used a sequential strategy, whereby they learned the associations one by one although they were presented randomly. Human subjects, tested on the same task, also used a sequential strategy. Neuronal recordings in monkeys revealed learning-related modulations of neural activity in the striatum. We then present a mathematical model inspired by viability theory developed to implement the use of strategies during learning. This model complements existing models of the BG based on reinforcement learning RL, which do not take into account the use of strategies to reduce the dimension of the learning space.
The learning classifier system (LCS) integrates a rule-based system with reinforcement learning and genetic algorithm-based rule discovery. This investigation reports on the design, implementation, and evaluation of EpiCS, a LCS adapted for knowledge discovery in epidemiologic surveillance. Using data from a large, national child automobile passenger protection program, EpiCS was compared with C4. 5 and logistic regression to evaluate its ability to induce rules from data that could be used to classify cases and to derive estimates of outcome risk, respectively. The rules induced by EpiCS were less parsimonious than those induced by C4.5, but were potentially more useful to investigators in hypothesis generation. Classification performance of C4.5 was superior to that of EpiCS (P<0.05). However, risk estimates derived by EpiCS were significantly more accurate than those derived by logistic regression (P<0.05).