PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15Linked to original sources

Online tuning of fuzzy inference systems using dynamic fuzzy Q-learning.

This paper presents a dynamic fuzzy Q-learning (DFQL) method that is capable of tuning fuzzy inference systems (FIS) online. A novel online self-organizing learning algorithm is developed so that structure and parameters identification are accomplished automatically and simultaneously based only on Q-learning. Self-organizing fuzzy inference is introduced to calculate actions and Q-functions so as to enable us to deal with continuous-valued states and actions. Fuzzy rules provide a natural mean of incorporating the bias components for rapid reinforcement learning. Experimental results and comparative studies with the fuzzy Q-learning (FQL) and continuous-action Q-learning in the wall-following task of mobile robots demonstrate that the proposed DFQL method is superior.

Algorithms↗

Learning to cooperate in solving the traveling salesman problem.

A cooperative team of agents may perform many tasks better than single agents. The question is how cooperation among self-interested agents should be achieved. It is important that, while we encourage cooperation among agents in a team, we maintain autonomy of individual agents as much as possible, so as to maintain flexibility and generality. This paper presents an approach based on bidding utilizing reinforcement values acquired through reinforcement learning. We tested and analyzed this approach and demonstrated that a team indeed performed better than the best single agent as well as the average of single agents.

Algorithms↗

Glutamate-mediated plasticity in corticostriatal networks: role in adaptive motor learning.

Little is known about how memories of new voluntary motor actions, also known as procedural memory, are formed at the molecular level. Our work examining acquisition of lever-pressing for food in rats has shown that activation of glutamate NMDA receptors, within broadly distributed but interconnected regions (e.g., nucleus accumbens core, prefrontal cortex, basolateral amygdala), is critical for such learning to occur. This receptor stimulation triggers intracellular cascades that involve protein phosphorylation and new protein synthesis. In support of this idea, we have found that posttrial inhibition of protein synthesis in the ventral striatum impairs learning, whereas posttrial NMDA receptor blockade does not. More recent data show extension of this network to the central amygdala, where infusions of NMDA antagonists also impair learning. We hypothesize that activity in this distributed network (including dopaminergic activity and perhaps muscarinic cholinergic activity) computes coincident events and thus enhances the probability that temporally related actions and events (e.g., lever pressing and delivery of reward) become associated. Such basic mechanisms of plasticity within this reinforcement learning network also appear to be profoundly affected in addiction.

Adaptation, Psychological↗

A novel model of motor learning capable of developing an optimal movement control law online from scratch.

A computational model of a learning system (LS) is described that acquires knowledge and skill necessary for optimal control of a multisegmental limb dynamics (controlled object or CO), starting from "knowing" only the dimensionality of the object's state space. It is based on an optimal control problem setup different from that of reinforcement learning. The LS solves the optimal control problem online while practicing the manipulation of CO. The system's functional architecture comprises several adaptive components, each of which incorporates a number of mapping functions approximated based on artificial neural nets. Besides the internal model of the CO's dynamics and adaptive controller that computes the control law, the LS includes a new type of internal model, the minimal cost (IM(mc)) of moving the controlled object between a pair of states. That internal model appears critical for the LS's capacity to develop an optimal movement trajectory. The IM(mc) interacts with the adaptive controller in a cooperative manner. The controller provides an initial approximation of an optimal control action, which is further optimized in real time based on the IM(mc). The IM(mc) in turn provides information for updating the controller. The LS's performance was tested on the task of center-out reaching to eight randomly selected targets with a 2DOF limb model. The LS reached an optimal level of performance in a few tens of trials. It also quickly adapted to movement perturbations produced by two different types of external force field. The results suggest that the proposed design of a self-optimized control system can serve as a basis for the modeling of motor learning that includes the formation and adaptive modification of the plan of a goal-directed movement.

Adaptation, Physiological↗

Actor-critic models of the basal ganglia: new anatomical and computational perspectives.

A large number of computational models of information processing in the basal ganglia have been developed in recent years. Prominent in these are actor-critic models of basal ganglia functioning, which build on the strong resemblance between dopamine neuron activity and the temporal difference prediction error signal in the critic, and between dopamine-dependent long-term synaptic plasticity in the striatum and learning guided by a prediction error signal in the actor. We selectively review several actor-critic models of the basal ganglia with an emphasis on two important aspects: the way in which models of the critic reproduce the temporal dynamics of dopamine firing, and the extent to which models of the actor take into account known basal ganglia anatomy and physiology. To complement the efforts to relate basal ganglia mechanisms to reinforcement learning (RL), we introduce an alternative approach to modeling a critic network, which uses Evolutionary Computation techniques to 'evolve' an optimal RL mechanism, and relate the evolved mechanism to the basic model of the critic. We conclude our discussion of models of the critic by a critical discussion of the anatomical plausibility of implementations of a critic in basal ganglia circuitry, and conclude that such implementations build on assumptions that are inconsistent with the known anatomy of the basal ganglia. We return to the actor component of the actor-critic model, which is usually modeled at the striatal level with very little detail. We describe an alternative model of the basal ganglia which takes into account several important, and previously neglected, anatomical and physiological characteristics of basal ganglia-thalamocortical connectivity and suggests that the basal ganglia performs reinforcement-biased dimensionality reduction of cortical inputs. We further suggest that since such selective encoding may bias the representation at the level of the frontal cortex towards the selection of rewarded plans and actions, the reinforcement-driven dimensionality reduction framework may serve as a basis for basal ganglia actor models. We conclude with a short discussion of the dual role of the dopamine signal in RL and in behavioral switching.

Animals↗

Accuracy-based learning classifier systems: models, analysis and applications to classification tasks.

Recently, Learning Classifier Systems (LCS) and particularly XCS have arisen as promising methods for classification tasks and data mining. This paper investigates two models of accuracy-based learning classifier systems on different types of classification problems. Departing from XCS, we analyze the evolution of a complete action map as a knowledge representation. We propose an alternative, UCS, which evolves a best action map more efficiently. We also investigate how the fitness pressure guides the search towards accurate classifiers. While XCS bases fitness on a reinforcement learning scheme, UCS defines fitness from a supervised learning scheme. We find significant differences in how the fitness pressure leads towards accuracy, and suggest the use of a supervised approach specially for multi-class problems and problems with unbalanced classes. We also investigate the complexity factors which arise in each type of accuracy-based LCS. We provide a model on the learning complexity of LCS which is based on the representative examples given to the system. The results and observations are also extended to a set of real world classification problems, where accuracy-based LCS are shown to perform competitively with respect to other learning algorithms. The work presents an extended analysis of accuracy-based LCS, gives insight into the understanding of the LCS dynamics, and suggests open issues for further improvement of LCS on classification tasks.

Algorithms↗

Autonomous learning based on cost assumptions: theoretical studies and experiments in robot control.

Autonomous learning techniques are based on experience acquisition. In most realistic applications, experience is time-consuming: it implies sensor reading, actuator control and algorithmic update, constrained by the learning system dynamics. The information crudeness upon which classical learning algorithms operate make such problems too difficult and unrealistic. Nonetheless, additional information for facilitating the learning process ideally should be embedded in such a way that the structural, well-studied characteristics of these fundamental algorithms are maintained. We investigate in this article a more general formulation of the Q-learning method that allows for a spreading of information derived from single updates towards a neighbourhood of the instantly visited state and converges to optimality. We show how this new formulation can be used as a mechanism to safely embed prior knowledge about the structure of the state space, and demonstrate it in a modified implementation of a reinforcement learning algorithm in a real robot navigation task.

Algorithms↗

Contract learning for self-care activities. A protocol study among chemotherapy outpatients.

Chemotherapy presents a challenge to patients and their families because of altered abilities for self-care. The nurse investigator initiated a protocol study with five patients responding to a query, "What do you need to know?," related to chemotherapy. Evaluation was done during February and March 1989 in a private oncology/hematology clinic in Hawaii. This was an evaluative study designed to measure the effectiveness of a contract learning protocol. A descriptive case study approach was used to analyze findings. Orem's self-care deficit theory of nursing provided the theoretical framework for the protocol. Based on Orem's supportive educative system, the nurse utilized the three types of self-care requisites to develop a learning needs assessment tool. Contract learning reinforced management of self-care deficits with self-care activities. Four of the subjects were able to recognize symptoms, weigh options for actions, initiate self-care behaviors, and evaluate the effectiveness of self-care activities. Findings suggest that the protocol provides a systematic and comprehensive approach to patient's self-care deficits using adult learning principles.

Adult↗

Cooperative multiagent congestion control for high-speed networks.

An adaptive multiagent reinforcement learning method for solving congestion control problems on dynamic high-speed networks is presented. Traditional reactive congestion control selects a source rate in terms of the queue length restricted to a predefined threshold. However, the determination of congestion threshold and sending rate is difficult and inaccurate due to the propagation delay and the dynamic nature of the networks. A simple and robust cooperative multiagent congestion controller (CMCC), which consists of two subsystems: a long-term policy evaluator, expectation-return predictor and a short-term rate selector composed of action-value evaluator and stochastic action selector elements has been proposed to solve the problem. After receiving cooperative reinforcement signals generated by a cooperative fuzzy reward evaluator using game theory, CMCC takes the best action to regulate source flow with the features of high throughput and low packet loss rate. By means of learning procedures, CMCC can learn to take correct actions adaptively under time-varying environments. Simulation results showed that the proposed approach can promote the system utilization and decrease packet losses simultaneously.

Algorithms↗

A study of structural and parametric learning in XCS.

The performance of a learning classifier system is due to its two main components. First, it evolves new structures by generating new rules in a genetic process; second, it adjusts parameters of existing rules, for example rule prediction and accuracy, in an evaluation step, which is not only important for applying the rules, but also for the genetic process. The two components interleave and in the case of XCS drive the population toward a minimal, fit, non-overlapping population. In this work we attempt to gain new insights as to the relative contributions of the two components. We find that the genetic component has an additional role when using the train/test approach which is not present in online learning. We compare XCS to a system in which the rule set is restricted to the initial random population (XCS-NGA, that is, XCS No Genetic Algorithm). For small Boolean functions we can give XCS-NGA all possible rules of a particular condition length. In online learning, XCS-NGA can, given sufficiently many rules, achieve a surprisingly high classification accuracy, comparable to that of XCS. In a train/test approach, however, XCS generalises better than XCS-NGA and there seem to be limitations of XCS-NGA which cannot be overcome simply by increasing the population size. This illustrates that the requirements of a function approximator tend to differ between reinforcement learning (which is typically online) and concept learning (which is typically train/test).

Algorithms↗

Learning to control a complex multistable system.

In this paper the control of a periodically kicked mechanical rotor without gravity in the presence of noise is investigated. In recent work it was demonstrated that this system possesses many competing attracting states and thus shows the characteristics of a complex multistable system. We demonstrate that it is possible to stabilize the system at a desired attracting state even in the presence of high noise level. The control method is based on a recently developed algorithm [S. Gadaleta and G. Dangelmayr, Chaos 9, 775 (1999)] for the control of chaotic systems and applies reinforcement learning to find a global optimal control policy directing the system from any initial state towards the desired state in a minimum number of iterations. Being data-based, the method does not require any information about governing dynamical equations.

Journal Article↗

Brain mechanism of reward prediction under predictable and unpredictable environmental dynamics.

In learning goal-directed behaviors, an agent has to consider not only the reward given at each state but also the consequences of dynamic state transitions associated with action selection. To understand brain mechanisms for action learning under predictable and unpredictable environmental dynamics, we measured brain activities by functional magnetic resonance imaging (fMRI) during a Markov decision task with predictable and unpredictable state transitions. Whereas the striatum and orbitofrontal cortex (OFC) were significantly activated both under predictable and unpredictable state transition rules, the dorsolateral prefrontal cortex (DLPFC) was more strongly activated under predictable than under unpredictable state transition rules. We then modelled subjects' choice behaviours using a reinforcement learning model and a Bayesian estimation framework and found that the subjects took larger temporal discount factors under predictable state transition rules. Model-based analysis of fMRI data revealed different engagement of striatum in reward prediction under different state transition dynamics. The ventral striatum was involved in reward prediction under both unpredictable and predictable state transition rules, although the dorsal striatum was dominantly involved in reward prediction under predictable rules. These results suggest different learning systems in the cortico-striatum loops depending on the dynamics of the environment: the OFC-ventral striatum loop is involved in action learning based on the present state, while the DLPFC-dorsal striatum loop is involved in action learning based on predictable future states.

Brain↗

LEAD: a methodology for learning efficient approaches to medical diagnosis.

Determining the most efficient use of diagnostic tests is one of the complex issues facing medical practitioners. With the soaring cost of healthcare, particularly in the US, there is a critical need for cutting costs of diagnostic tests, while achieving a higher level of diagnostic accuracy. This paper develops a learning based methodology that, based on patient information, recommends test(s) that optimize a suitable measure of diagnostic performance. A comprehensive performance measure is developed that accounts for the costs of testing, morbidity, and mortality associated with the tests, and time taken to reach diagnosis. The performance measure also accounts for the diagnostic ability of the tests. The methodology combines tools from the fields of data mining (rough set theory, in particular), utility theory, Markov decision processes (MDP), and reinforcement learning (RL). The rough set theory is used in extracting diagnostic information in the form of rules from the medical databases. Utility theory is used in bringing various nonhomogenous performance measures into one cost based measure. An MDP model together with an RL algorithm facilitates obtaining efficient testing strategies. The methodology is implemented on a sample problem of diagnosing solitary pulmonary nodule (SPN). The results obtained are compared with those from four alternative testing strategies. Our methodology holds significant promise to improve the process of medical diagnosis.

Algorithms↗

Context dependence of the event-related brain potential associated with reward and punishment.

The error-related negativity (ERN) is an event-related brain potential elicited by error commission and by presentation of feedback stimuli indicating incorrect performance. In this study, the authors report two experiments in which participants tried to learn to select between response options by trial and error, using feedback stimuli indicating monetary gains and losses. The results demonstrate that the amplitude of the ERN is determined by the value of the eliciting outcome relative to the range of outcomes possible, rather than by the objective value of the outcome. This result is discussed in terms of a recent theory that holds that the ERN reflects a reward prediction error signal associated with a neural system for reinforcement learning.

Adult↗

Helping expand nurse practitioner students' clinical skills repertoire: learning minor procedures.

This course is immensely popular with students. Many express a sense of accomplishment in knowing that they have performed a specific procedure at least once before facing the requirement in the clinical setting. Students who complete the course have a firm foundation in minor procedures and are knowledgeable about indications, contraindications, and methods to perform the procedure before facing that first patient in the clinical setting. They have also accumulated a set of procedures, patient discharge instructions, and a bibliography that can be reviewed before performing a procedure for the first time in the clinical setting. Combining clinical hands-on skills with the experience of writing a step-by-step procedure reinforces learning and is a valuable skill that can be used in clinical practice after graduation. Although many of the students will only apply a portion of the skills they learn in this course, they express a significant boost in self-confidence, a decrease in anxiety level, a sense of accomplishment in their skills, and a definite edge in securing a nurse practitioner position in a competitive healthcare marketplace.

Educational Measurement↗

Cerebellar learning of bio-mechanical functions of extra-ocular muscles: modeling by artificial neural networks.

A control circuit is proposed to model the command of saccadic eye movements. Its wiring is deduced from a mathematical constraint, i.e. the necessity, for motor orders processing, to compute an approximate inverse function of the bio-mechanical function of the moving plant, here the bio-mechanics of the eye. This wiring is comparable to the anatomy of the cerebellar pathways. A predicting element, necessary for inversion and thus for movement accuracy, is modeled by an artificial neural network whose structure, deduced from physical constraints expressing the mechanics of the eye, is similar to the cell connectivity of the cerebellar cortex. Its functioning is set by supervised reinforcement learning, according to learning rules aimed at reducing the errors of pointing, and deduced from a differential calculation. After each movement, a teaching signal encoding the pointing error is distributed to various learning sites, as is, in the cerebellum, the signal issued from the inferior olive and conveyed to various cell types by the climbing fibers. Results of simulations lead to predict the existence of a learning site in the glomeruli. After learning, the model is able to accurately simulate saccadic eye movements. It accounts for the function of the cerebellar pathways and for the final integrator of the oculomotor system. The novelty of this model of movement control is that its structure is entirely deduced from mathematical and physical constraints, and is consistent with general anatomy, cell connectivity and functioning of the cerebellar pathways. Even the learning rules can be deduced from calculation, and they reproduce long term depression, the learning process which takes place in the dendritic arborization of the Purkinje cells. This approach, based on the laws of mathematics and physics, appears thus as an efficient way of understanding signal processing in the motor system.

Biomechanical Phenomena↗

Evolution of the Gynecology Teaching Associate: an education specialist.

The traditional pelvic examination instruction methods were reviewed and found to be deficient: the student learning experience was compromised by the triangular setting of patient, student, and instructor for early pelvic examination instruction. Over the past decade, a new education specialist, the Gynecology Teaching Associate (GTA), has evolved to help improve the initial gynecology teaching experience. The evolution of the GTA is described. The qualities she brings to the instructional system include sensitivity as a woman, educational skill in pelvic examination instruction, knowledge of female pelvic anatomy and physiology, and, most important, sophisticated interpersonal skills to help medical students learnin in a nonthreatening environment. Reinforcement learning theory is the foundation of this educational system. Student acceptance of this system is documented.

Curriculum↗

Factors associated with the effectiveness of continuing education in long-term care.

PURPOSE: This article examines factors within the long-term-care work environment that impact the effectiveness of continuing education. DESIGN AND METHODS: In Study 1, focus group interviews were conducted with staff and management from urban and rural long-term-care facilities in southwestern Ontario to identify their perceptions of the workplace factors that affect transfer of learning into practice. Thirty-five people were interviewed across six focus groups. In Study 2, a Delphi technique was used to refine our list of factors. Consensus was achieved in two survey rounds involving 30 and 27 participants, respectively. RESULTS: Management support was identified as the most important factor impacting the effectiveness of continuing education. Other factors included resources (staff, funding, space) and the need for ongoing expert support. IMPLICATIONS: Organizational support is necessary for continuing education programs to be effective and ongoing expert support is needed to enable and reinforce learning.

Delphi Technique↗