PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9Linked to original sources

Brain nicotinic receptors: structure and regulation, role in learning and reinforcement.

The introduction, in the late sixties, of the concepts and methods of molecular biology to the study of the nervous system had a profound impact on the field, primarily through the identification of its basic molecular components. These structures include, for example, the elementary units of the synapse: neurotransmitters, neuropeptides and their receptors, but also ionic channels, intracellular second messengers and the relevant enzymes, cell surface adhesion molecules, or growth and trophic factors [21,78,81, 52,79]. Attempts to establish appropriate causal relationships between these molecular components, the actual organisation of neural networks, and a defined behavior, nevertheless, still must overcome many difficulties. A first problem is the recognition of the minimum levels of organisation, from the molecular, cellular, or multicellular (circuit) to the higher cognitive levels, that determine the given physiological and/or behavioral performance under investigation. A common difficulty (and potential source of errors of interpretation) is to relate a cognitive function to a network organization which does not possess the required structural complexity and vice-versa. Another problem is to distinguish, among the components of the system, those which are actually necessary and those which, taken together, suffice for a given behavior to take place. Identification of such a minimal set of building blocks may receive decisive insights from the elaboration of neurally plausible formal models that bring together, within a single and coherent 'artificial organism', the neuronal network, the circulating activity, and the behavior they determine (see [42,43,45,72,30]). In this communication, we shall attempt, still in a preliminary fashion, to bring together: (1) our recent knowledge on the molecular biology of brain nicotinic receptors (nAChRs) and their allosteric properties and (2) integrated behaviors, such as cognitive learning, investigated for instance with delayed-response or passive avoidance tasks that are likely to involve nAChRs in particular at the level of reinforcement (or reward) mechanisms (see [18,29,135]).

Allosteric Regulation↗

Learning and decision making in monkeys during a rock-paper-scissors game.

Game theory provides a solution to the problem of finding a set of optimal decision-making strategies in a group. However, people seldom play such optimal strategies and adjust their strategies based on their experience. Accordingly, many theories postulate a set of variables related to the probabilities of choosing various strategies and describe how such variables are dynamically updated. In reinforcement learning, these value functions are updated based on the outcome of the player's choice, whereas belief learning allows the value functions of all available choices to be updated according to the choices of other players. We investigated the nature of learning process in monkeys playing a competitive game with ternary choices, using a rock-paper-scissors game. During the baseline condition in which the computer selected its targets randomly, each animal displayed biases towards some targets. When the computer exploited the pattern of animal's choice sequence but not its reward history, the animal's choice was still systematically biased by the previous choice of the computer. This bias was reduced when the computer exploited both the choice and reward histories of the animal. Compared to simple models of reinforcement learning or belief learning, these adaptive processes were better described by a model that incorporated the features of both models. These results suggest that stochastic decision-making strategies in primates during social interactions might be adjusted according to both actual and hypothetical payoffs.

Algorithms↗

Midbrain dopamine neurons encode a quantitative reward prediction error signal.

The midbrain dopamine neurons are hypothesized to provide a physiological correlate of the reward prediction error signal required by current models of reinforcement learning. We examined the activity of single dopamine neurons during a task in which subjects learned by trial and error when to make an eye movement for a juice reward. We found that these neurons encoded the difference between the current reward and a weighted average of previous rewards, a reward prediction error, but only for outcomes that were better than expected. Thus, the firing rate of midbrain dopamine neurons is quantitatively predicted by theoretical descriptions of the reward prediction error signal used in reinforcement learning models for circumstances in which this signal has a positive value. We also found that the dopamine system continued to compute the reward prediction error even when the behavioral policy of the animal was only weakly influenced by this computation.

Algorithms↗

Tonic dopamine: opportunity costs and the control of response vigor.

RATIONALE: Dopamine neurotransmission has long been known to exert a powerful influence over the vigor, strength, or rate of responding. However, there exists no clear understanding of the computational foundation for this effect; predominant accounts of dopamine's computational function focus on a role for phasic dopamine in controlling the discrete selection between different actions and have nothing to say about response vigor or indeed the free-operant tasks in which it is typically measured. OBJECTIVES: We seek to accommodate free-operant behavioral tasks within the realm of models of optimal control and thereby capture how dopaminergic and motivational manipulations affect response vigor. METHODS: We construct an average reward reinforcement learning model in which subjects choose both which action to perform and also the latency with which to perform it. Optimal control balances the costs of acting quickly against the benefits of getting reward earlier and thereby chooses a best response latency. RESULTS: In this framework, the long-run average rate of reward plays a key role as an opportunity cost and mediates motivational influences on rates and vigor of responding. We review evidence suggesting that the average reward rate is reported by tonic levels of dopamine putatively in the nucleus accumbens. CONCLUSIONS: Our extension of reinforcement learning models to free-operant tasks unites psychologically and computationally inspired ideas about the role of tonic dopamine in striatum, explaining from a normative point of view why higher levels of dopamine might be associated with more vigorous responding.

Animals↗

Opponent appetitive-aversive neural processes underlie predictive learning of pain relief.

Termination of a painful or unpleasant event can be rewarding. However, whether the brain treats relief in a similar way as it treats natural reward is unclear, and the neural processes that underlie its representation as a motivational goal remain poorly understood. We used fMRI (functional magnetic resonance imaging) to investigate how humans learn to generate expectations of pain relief. Using a pavlovian conditioning procedure, we show that subjects experiencing prolonged experimentally induced pain can be conditioned to predict pain relief. This proceeds in a manner consistent with contemporary reward-learning theory (average reward/loss reinforcement learning), reflected by neural activity in the amygdala and midbrain. Furthermore, these reward-like learning signals are mirrored by opposite aversion-like signals in lateral orbitofrontal cortex and anterior cingulate cortex. This dual coding has parallels to 'opponent process' theories in psychology and promotes a formal account of prediction and expectation during pain.

Avoidance Learning↗

The dental curriculum at North American dental institutions in 2002-03: a survey of current structure, recent innovations, and planned changes.

This study examined the current format of curricula at North American dental schools, determined curriculum evaluation strategies, and identified recently implemented changes as well as planned future innovations. The academic affairs deans of sixty-four North American dental schools received an email survey in August 2002; a second, follow-up survey was sent to nonresponders in February 2003. Online responses were collected and analyzed using SurveyTracker software. The final response rate was 87 percent, with forty-eight U.S. schools and eight Canadian schools responding. Respondents were asked to select descriptive statements about the general organization of their curricula and the degree to which problem-based learning (PBL), case-reinforced learning (CRL), curricular integration, and community-based clinical treatment experiences were incorporated. They were also requested to identify strategies employed to evaluate the curriculum and to report recently completed and desired future curriculum modifications. In regard to desired future curriculum innovations, respondents identified why they were considering curriculum changes and identified resources needed to implement the planned changes. Sixty-six percent of those who responded defined their current curriculum organization as primarily discipline-based with a few interdisciplinary courses. Nearly 60 percent of schools reported that they used PBL and CRL in specific courses or for components of certain courses, but only 5 percent of the respondents indicated that all of their courses used PBL. Regarding integration of major sections of the curriculum, only 7 percent reported that their entire curriculum was organized around themes of interrelated topics. Sixty-four percent reported that their curriculum had required community-based clinical treatment experiences for students. The most frequent innovations in the past three years were increased use of computer and web-based learning (86 percent), creation of patient care experiences early in the curriculum (84 percent), enhancement of competency evaluation methods (84 percent), and curriculum decompression (79 percent). These items plus increased community-based care were the most frequently identified future curricular innovations. There were virtually no differences between the responses of Canadian and U.S. dental schools. The results of this study help to broadly characterize dental curricula at North American dental institutions and identify curriculum modifications anticipated by the academic dean respondents.

Canada↗

Sustained enhancement of AMPA receptor- and NMDA receptor-mediated currents induced by dopamine D1/D5 receptor activation in the hippocampus: an essential role of postsynaptic Ca2+.

The dopaminergic system in the limbic system, particularly the D1/D5 receptor (D1/D5r), is important for certain forms of learning and memory, such as reinforcement learning, as well as the acute development of behavioral sensitization to psychostimulants such as cocaine and methamphetamine. Here, whole-cell patch-clamp recordings of evoked excitatory postsynaptic currents (EPSCs) mediated by ionotropic glutamate receptors were made from pyramidal neurons in the CA1 area of rat hippocampal slices. Activation of the D1/D5r by a selective D1/D5r agonist (+/-)-6-chloro-PB hydrobromide (SKF 81297; 3-100 microM) concentration-dependently induced a delayed-onset, sustained enhancement of EPSCs. The D1/D5r-induced effect was blocked by a selective D1/D5r antagonist (+)SCH 23390 hydrochloride (5 microM). Furthermore, the D1/D5r-induced effect involved a sustained enhancement of both the pharmacologically isolated alpha-amino-3-hydroxy-5-methyl-4-isoxazoleproprionate receptor-(AMPAr-) and N-methyl-D-asparate receptor- (NMDAr-) mediated EPSCs, respectively. Such persistently enhanced AMPAr- and NMDAr-mediated EPSCs were also associated with significant increases in the normalized amplitudes of the decay times of the isolated EPSCs. Blockade of NMDAr activation failed to prevent the induction of D1/D5r-induced sustained enhancement, suggesting the independence of the D1/D5r-induced effect on NMDAr activation. A rise of postsynaptic calcium was required for the induction of D1/D5r-induced sustained enhancement, being abolished by loading cells with 1,2-bis(2-aminophenoxy)ethane-N,N,N',N '-tetracetic acid (BAPTA; 10 mM). Studies indicated that the D1/D5r activation induced a sustained enhancement of both the AMPAr-and NMDAr-mediated EPSCs in the hippocampus. Moreover, a rise in postsynaptic Ca2+ was necessary for triggering the D1/D5r-induced effect in the hippocampus, although the NMDAr-dependent mechanisms, such as the calcium entry via the NMDA channels, were not.

6-Cyano-7-nitroquinoxaline-2,3-dione↗

Autonomous mental development in high dimensional context and action spaces.

Autonomous Mental Development (AMD) of robots opened a new paradigm for developing machine intelligence, using neural network type of techniques and it fundamentally changed the way an intelligent machine is developed from manual to autonomous. The work presented here is a part of SAIL (Self-Organizing Autonomous Incremental Learner) project which deals with autonomous development of humanoid robot with vision, audition, manipulation and locomotion. The major issue addressed here is the challenge of high dimensional action space (5-10) in addition to the high dimensional context space (hundreds to thousands and beyond), typically required by an AMD machine. This is the first work that studies a high dimensional (numeric) action space in conjunction with a high dimensional perception (context state) space, under the AMD mode. Two new learning algorithms, Direct Update on Direction Cosines (DUDC) and High-Dimensional Conjugate Gradient Search (HCGS), are developed, implemented and tested. The convergence properties of both the algorithms and their targeted applications are discussed. Autonomous learning of speech production under reinforcement learning is studied as an example.

Learning↗

Structured clinical teaching strategy.

This study investigated the impact of a structured exercise on low back pain, as part of a second year ambulatory course, on students' low back pain examination skills. One-hundred and eighty-eight medical students participated in one of four types of instructional intervention: 1) structured clinical exercise and reading, 2) random clinical experience and no reading, 3) reading only, and 4) no clinical experience or reading. At the end of the year, students completed an Objective Structured Clinical Exam (OSCE) in which two stations assessed back pain history and physical exam skills. An analysis of variance of the OSCE scores showed no significant difference in students' performance in relation to the type of instructional intervention. The General Professional Education of the Physician Report argues for the importance of structured clinical education. In medical education, unpredictable patient exposure and crowded curricula frequently result in students having none or only one structured learning opportunity to acquire critical skills. Unfortunately, this study found that assuring at least one structured clinical experience for a specific common problem seen in ambulatory care did not enhance student ability to select and use specific history or physical exam skills for that problem as assessed by an OSCE, as compared with students who did not have this experience. To assure that essential clinical skills are acquired, it most likely requires that both systematic instructional strategies and repeated learning opportunities are available to reinforce learning.

Ambulatory Care↗