PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4Linked to original sources

Integrating relevance feedback techniques for image retrieval using reinforcement learning.

Relevance feedback (RF) is an interactive process which refines the retrievals to a particular query by utilizing the user's feedback on previously retrieved results. Most researchers strive to develop new RF techniques and ignore the advantages of existing ones. In this paper, we propose an image relevance reinforcement learning (IRRL) model for integrating existing RF techniques in a content-based image retrieval system. Various integration schemes are presented and a long-term shared memory is used to exploit the retrieval experience from multiple users. Also, a concept digesting method is proposed to reduce the complexity of storage demand. The experimental results manifest that the integration of multiple RF approaches gives better retrieval performance than using one RF technique alone, and that the sharing of relevance knowledge between multiple query sessions significantly improves the performance. Further, the storage demand is significantly reduced by the concept digesting technique. This shows the scalability of the proposed model with the increasing-size of database.

Algorithms↗

Signal detection by human observers: a cutoff reinforcement learning model of categorization decisions under uncertainty.

Previous experimental examinations of binary categorization decisions have documented robust behavioral regularities that cannot be predicted by signal detection theory (D.M. Green & J.A. Swets, 1966/1988). The present article reviews the known regularities and demonstrates that they can be accounted for by a minimal modification of signal detection theory: the replacement of the "ideal observer" cutoff placement rule with a cutoff reinforcement learning rule. This modification is derived from a cognitive game theoretic analysis (A.E. Roth & I. Erev, 1995). The modified model reproduces all 19 experimental regularities that have been considered. In all cases,it outperforms the original explanations. Some of these previous explanations are based on important concepts such as conservatism, probability matching, and "the gambler's fallacy" that receive new meanings given the current results. Implications for decision-making research and for applications of traditional signal detection theory are discussed.

Decision Making↗

Reward, motivation, and reinforcement learning.

There is substantial evidence that dopamine is involved in reward learning and appetitive conditioning. However, the major reinforcement learning-based theoretical models of classical conditioning (crudely, prediction learning) are actually based on rules designed to explain instrumental conditioning (action learning). Extensive anatomical, pharmacological, and psychological data, particularly concerning the impact of motivational manipulations, show that these models are unreasonable. We review the data and consider the involvement of a rich collection of different neural systems in various aspects of these forms of conditioning. Dopamine plays a pivotal, but complicated, role.

Animals↗

Reinforcement learning by Hebbian synapses with adaptive thresholds.

A central problem in learning theory is how the vertebrate brain processes reinforcing stimuli in order to master complex sensorimotor tasks. This problem belongs to the domain of supervised learning, in which errors in the response of a neural network serve as the basis for modification of synaptic connectivity in the network and thereby train it on a computational task. The model presented here shows how a reinforcing feedback can modify synapses in a neuronal network according to the principles of Hebbian learning. The reinforcing feedback steers synapses towards long-term potentiation or depression by critically influencing the rise in postsynaptic calcium, in accordance with findings on synaptic plasticity in mammalian brain. An important feature of the model is the dependence of modification thresholds on the previous history of reinforcing feedback processed by the network. The learning algorithm trained networks successfully on a task in which a population vector in the motor output was required to match a sensory stimulus vector presented shortly before. In another task, networks were trained to compute coordinate transformations by combining different visual inputs. The model continued to behave well when simplified units were replaced by single-compartment neurons equipped with several conductances and operating in continuous time. This novel form of reinforcement learning incorporates essential properties of Hebbian synaptic plasticity and thereby shows that supervised learning can be accomplished by a learning rule similar to those used in physiologically plausible models of unsupervised learning. The model can be crudely correlated to the anatomy and electrophysiology of the amygdala, prefrontal and cingulate cortex and has predictive implications for further experiments on synaptic plasticity and learning processes mediated by these areas.

Adaptation, Physiological↗

Neuromuscular control of the point to point and oscillatory movements of a sagittal arm with the actor-critic reinforcement learning method.

In this study, we have used a single link system with a pair of muscles that are excited with alpha and gamma signals to achieve both point to point and oscillatory movements with variable amplitude and frequency.The system is highly nonlinear in all its physical and physiological attributes. The major physiological characteristics of this system are simultaneous activation of a pair of nonlinear muscle-like-actuators for control purposes, existence of nonlinear spindle-like sensors and Golgi tendon organ-like sensor, actions of gravity and external loading. Transmission delays are included in the afferent and efferent neural paths to account for a more accurate representation of the reflex loops.A reinforcement learning method with an actor-critic (AC) architecture instead of middle and low level of central nervous system (CNS), is used to track a desired trajectory. The actor in this structure is a two layer feedforward neural network and the critic is a model of the cerebellum. The critic is trained by state-action-reward-state-action (SARSA) method. The critic will train the actor by supervisory learning based on the prior experiences. Simulation studies of oscillatory movements based on the proposed algorithm demonstrate excellent tracking capability and after 280 epochs the RMS error for position and velocity profiles were 0.02, 0.04 rad and rad/s, respectively.

Arm↗

Effects of mGlu1 and mGlu5 receptor antagonists on negatively reinforced learning.

Effects on aversive learning of the novel highly selective mGlu5 receptor antagonist [(2-methyl-1,3-thiazol-4-yl)ethynyl]pyridine (MTEP) and mGlu1 receptor antagonist (3-ethyl-2-methyl-quinolin-6-yl)-(4-methoxy-cyclohexyl)-methanone methanesulfonate (EMQMCM) were tested, after systemic administration, in the passive avoidance (PA) and fear potentiated startle (FPS) paradigms. Both MTEP at 10 mg/kg and EMQMCM at 5 and 10 mg/kg, given 30 min before training, impaired acquisition of the passive avoidance response (PAR). Co-administration of MTEP and EMQMCM at doses ineffective when administered alone, produced anterograde amnesia when given 30 min before the acquisition phase. Neither EMQMCM (5 mg/kg) nor MTEP (10 mg/kg) impaired retention of the PAR after direct post-training injections. EMQMCM (5 mg/kg), but not MTEP (10 mg/kg) blocked the PAR when given 30 min before testing. Pre-training administration of MTEP at doses of 2.5 and 5 mg/kg inhibited fear conditioning in the FPS when tested 24 h later. In contrast, EMQMCM was ineffective. Our findings suggest diverse involvement of mGlu1 and mGlu5 receptors in negatively reinforced learning.

Amnesia, Anterograde↗

Heterarchical reinforcement-learning model for integration of multiple cortico-striatal loops: fMRI examination in stimulus-action-reward association learning.

The brain's most difficult computation in decision-making learning is searching for essential information related to rewards among vast multimodal inputs and then integrating it into beneficial behaviors. Contextual cues consisting of limbic, cognitive, visual, auditory, somatosensory, and motor signals need to be associated with both rewards and actions by utilizing an internal representation such as reward prediction and reward prediction error. Previous studies have suggested that a suitable brain structure for such integration is the neural circuitry associated with multiple cortico-striatal loops. However, computational exploration still remains into how the information in and around these multiple closed loops can be shared and transferred. Here, we propose a "heterarchical reinforcement learning" model, where reward prediction made by more limbic and cognitive loops is propagated to motor loops by spiral projections between the striatum and substantia nigra, assisted by cortical projections to the pedunculopontine tegmental nucleus, which sends excitatory input to the substantia nigra. The model makes several fMRI-testable predictions of brain activity during stimulus-action-reward association learning. The caudate nucleus and the cognitive cortical areas are correlated with reward prediction error, while the putamen and motor-related areas are correlated with stimulus-action-dependent reward prediction. Furthermore, a heterogeneous activity pattern within the striatum is predicted depending on learning difficulty, i.e., the anterior medial caudate nucleus will be correlated more with reward prediction error when learning becomes difficult, while the posterior putamen will be correlated more with stimulus-action-dependent reward prediction in easy learning. Our fMRI results revealed that different cortico-striatal loops are operating, as suggested by the proposed model.

Association Learning↗

Anatomy of a decision: striato-orbitofrontal interactions in reinforcement learning, decision making, and reversal.

The authors explore the division of labor between the basal ganglia-dopamine (BG-DA) system and the orbitofrontal cortex (OFC) in decision making. They show that a primitive neural network model of the BG-DA system slowly learns to make decisions on the basis of the relative probability of rewards but is not as sensitive to (a) recency or (b) the value of specific rewards. An augmented model that explores BG-OFC interactions is more successful at estimating the true expected value of decisions and is faster at switching behavior when reinforcement contingencies change. In the augmented model, OFC areas exert top-down control on the BG and premotor areas by representing reinforcement magnitudes in working memory. The model successfully captures patterns of behavior resulting from OFC damage in decision making, reversal learning, and devaluation paradigms and makes additional predictions for the underlying source of these deficits.

Animals↗

GiantHunter: accurate detection of giant virus in metagenomic data using reinforcement-learning and Monte Carlo tree search.

MOTIVATION: Nucleocytoplasmic large DNA viruses (NCLDVs) are notable for their large genomes and extensive gene repertoires, which contribute to their widespread environmental presence and critical roles in processes such as host metabolic reprogramming and nutrient cycling. Metagenomic sequencing has emerged as a powerful tool for uncovering novel NCLDVs in environmental samples. However, identifying NCLDV sequences in metagenomic data remains challenging due to their high genomic diversity, limited reference genomes, and shared regions with other microbes. Existing alignment-based and machine learning methods struggle with achieving optimal trade-offs between sensitivity and precision. RESULTS: In this work, we present GiantHunter, a reinforcement learning-based tool for identifying NCLDVs from metagenomic data. By employing a Monte Carlo tree search strategy, GiantHunter dynamically selects representative non-NCLDV sequences as the negative training data, enabling the model to establish a robust decision boundary. Benchmarking on rigorously designed experiments shows that GiantHunter achieves high precision while maintaining competitive sensitivity, improving the F1-score by 10% and reducing computational cost by 90% compared to the second-best method. To demonstrate its real-world utility, we applied GiantHunter to 60 metagenomic datasets collected from six cities along the Yangtze River, located both upstream and downstream of the Three Gorges Dam. The results reveal significant differences in NCLDV diversity correlated with proximity to the dam, likely influenced by reduced flow velocity caused by the dam. These findings highlight GiantHunter's potential to advance our understanding of NCLDVs and their ecological roles in diverse environments. AVAILABILITY AND IMPLEMENTATION: The source code of GiantHunter is available via: https://github.com/FuchuanQu/GiantHunter.

Metagenomics↗

Individualization of pharmacological anemia management using reinforcement learning.

Effective management of anemia due to renal failure poses many challenges to physicians. Individual response to treatment varies across patient populations and, due to the prolonged character of the therapy, changes over time. In this work, a Reinforcement Learning-based approach is proposed as an alternative method for individualization of drug administration in the treatment of renal anemia. Q-learning, an off-policy approximate dynamic programming method, is applied to determine the proper dosing strategy in real time. Simulations compare the proposed methodology with the currently used dosing protocol. Presented results illustrate the ability of the proposed method to achieve the therapeutic goal for individuals with different response characteristics and its potential to become an alternative to currently used techniques.

Algorithms↗

Chicks' maze learning reinforced by visual pitfall extending downward.

The present study examined whether visually evoked fear of depth could reinforce a particular response of animals, i.e., to special maze learning. The maze was composed of four units of Y-shaped alley. In this maze, the visual pitfalls were set behind corners of the alley in place of a physical barrier. The experiments showed that eight of 13 male chicks could achieve the initial learning and that three successful ones could also achieve reversal learning. The results suggest that the visually evoked fear of depth provided by motion parallax can act as a reinforcer.

Animals↗

Errorless learning: reinforcement contingencies and stimulus control transfer in delayed prompting.

Delayed prompting can produce errorless discrimination learning. There is inherent in the procedure a disparity in reinforcement density which favors unprompted over prompted responses. We used three schedules of reinforcement to investigate the impact of reinforcement probability on transfer of stimulus control. One schedule of reinforcement was equal prior to and following a prompt (CRF/CRF), the second favored unprompted responses (CRF/FR3), and the third favored responses following the prompt (FR3/CRF). Experimental questions concerned the probability of errors, the probability of transfer, and the rate of transfer in the context of delayed prompting. Transfer was accelerated when reinforcement probability favored anticipatory responding. The schedule that favored prompted responses did not prevent a shift to unprompted responding. Errors were infrequent across procedures. Reinforcement probability contributes to but does not entirely determine transfer of stimulus control from a delayed prompt.

Adolescent↗

Memory formation processes in weakly reinforced learning.

Day-old chicks trained on a single-trail passive avoidance learning task, with varying concentrations of the aversive stimulus (methyl anthranilate), truncated retention functions for low concentrations. The retention function for a 20% v/v dilution of methyl anthranilate in absolute ethanol yielded high retention levels until approximately 40 to 45 minutes following learning. This retention function appears to consist of only the short-term and intermediate (phase A) memory stages of Gibbs and Ng's three-stage model of memory formation, with the short-term stage susceptible to inhibition by monosodium glutamate, and the intermediate stage by ouabain and dinitrophenol. The results suggest that processing of memory into the relatively permanent long-term stage may depend on the strength of the reinforcer in aversive learning.

2,4-Dinitrophenol↗

Similar effects of a beta-carboline and of flumazenil in negatively and positively reinforced learning tasks in mice.

Methyl beta-carboline-3-carboxylate (beta-CCM) and flumazenil (Ro15-1788) are known to be respectively an inverse agonist and an antagonist of the central benzodiazepine-receptor. Surprisingly, these two drugs have shown a similar enhancing effect in a negatively reinforced multiple-trial brightness discrimination task in mice. Thus, to evaluate the role of anxiety in this task, the action of these two drugs were compared in the same learning task with a positive or a negative reinforcement. Mice were trained for sessions of ten trials per day for six consecutive days. The sessions during the first three days took place after administration of beta-CCM (0.3 mg/kg), flumazenil (15 mg/kg) or vehicles of these drugs. A negative reinforcement (electric foot-shock) was used in a first experiment, and a positive one (food reward) in a second experiment. Results showed that, whatever the reinforcement, the two drugs enhance learning in a brightness discrimination task. The hypothesis is that flumazenil could have an inverse agonist profile in learning tasks. The question remains as to whether the flumazenil enhancing learning process results from increased arousal and/or anxiogenic factors, or from a negative modulatory influence of endogenous diazepam-like ligands for benzodiazepine receptors.

Animals↗

Error-related negativity predicts reinforcement learning and conflict biases.

The error-related negativity (ERN) is an electrophysiological marker thought to reflect changes in dopamine when participants make errors in cognitive tasks. Our computational model further predicts that larger ERNs should be associated with better learning to avoid maladaptive responses. Here we show that participants who avoided negative events had larger ERNs than those who were biased to learn more from positive outcomes. We also tested for effects of response conflict on ERN magnitude. While there was no overall effect of conflict, positive learners had larger ERNs when having to choose among two good options (win/win decisions) compared with two bad options (lose/lose decisions), whereas negative learners exhibited the opposite pattern. These results demonstrate that the ERN predicts the degree to which participants are biased to learn more from their mistakes than their correct choices and clarify the extent to which it indexes decision conflict.

Adaptation, Physiological↗

The effect of two forms of learning reinforcement upon parental retention of CPR skills.

PURPOSE: To evaluate the effect of two forms of reinforcement upon the retention of CPR psychomotor skills in parents of high-risk infants. METHOD: A pretest/posttest, 3-group design was done with a sample of 69 parent volunteers. Reinforcement with hands on practice was given to one group; a second group had reinforcement by observing a videotape, and the third group had no reinforcement strategy. CPR skills were measured by a checklist. FINDINGS: Paired t-test showed significant differences in pretest and posttest scores for all three groups. Several CPR skills were missed on the final test more frequently than others. One-way ANOVA also showed a significant difference when comparing the control group with the hands-on reinforcement group. CONCLUSIONS: Reinforcement of CPR skills should be an ongoing process. The groups who had reinforcement with hands-on practice retained the most skills.

Adolescent↗