PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3Linked to original sources

Forward propagating reinforcement learning--biologically plausible learning method for multi-layer networks.

We introduce a biologically plausible method of implementing reinforcement learning to multi-layer neural networks. The key idea is to spatially localize the synaptic modulation induced by reinforcement signals, proceeding downstream from the initial layer to the final layer. Since reinforcement signals are known to be broadcast signals in the actual brain, we need two key assumptions, inhibitory backward connections and bypass to output units, to spatially localize the effect of delayed reinforcement without breaking the basic laws of neurophysiology.

Animals↗

Complex sensory-motor sequence learning based on recurrent state representation and reinforcement learning.

A novel neural network model is presented that learns by trial-and-error to reproduce complex sensory-motor sequences. One subnetwork, corresponding to the prefrontal cortex (PFC), is responsible for generating unique patterns of activity that represent the continuous state of sequence execution. A second subnetwork, corresponding to the striatum, associates these state-encoding patterns with the correct response at each point in the sequence execution. From a neuroscience perspective, the model is based on the known cortical and subcortical anatomy of the primate oculomotor system. From a theoretical perspective, the architecture is similar to that of a finite automaton in which outputs and state transitions are generated as a function of inputs and the current state. Simulation results for complex sequence reproduction and sequence discrimination are presented.

Animals↗

Reinforcement learning algorithms for robotic navigation in dynamic environments.

The purpose of this study was to examine improvements to reinforcement learning (RL) algorithms in order to successfully interact within dynamic environments. The scope of the research was that of RL algorithms as applied to robotic navigation. Proposed improvements include: addition of a forgetting mechanism, use of feature based state inputs, and hierarchical structuring of an RL agent. Simulations were performed to evaluate the individual merits and flaws of each proposal, to compare proposed methods to prior established methods, and to compare proposed methods to theoretically optimal solutions. Incorporation of a forgetting mechanism did considerably improve the learning times of RL agents in a dynamic environment. However, direct implementation of a feature-based RL agent did not result in any performance enhancements, as pure feature-based navigation results in a lack of positional awareness, and the inability of the agent to determine the location of the goal state. Inclusion of a hierarchical structure in an RL agent resulted in significantly improved performance, specifically when one layer of the hierarchy included a feature-based agent for obstacle avoidance, and a standard RL agent for global navigation. In summary, the inclusion of a forgetting mechanism, and the use of a hierarchically structured RL agent offer substantially increased performance when compared to traditional RL agents navigating in a dynamic environment.

Algorithms↗

The asymptotic equipartition property in reinforcement learning and its relation to return maximization.

We discuss an important property called the asymptotic equipartition property on empirical sequences in reinforcement learning. This states that the typical set of empirical sequences has probability nearly one, that all elements in the typical set are nearly equi-probable, and that the number of elements in the typical set is an exponential function of the sum of conditional entropies if the number of time steps is sufficiently large. The sum is referred to as stochastic complexity. Using the property we elucidate the fact that the return maximization depends on two factors, the stochastic complexity and a quantity depending on the parameters of environment. Here, the return maximization means that the best sequences in terms of expected return have probability one. We also examine the sensitivity of stochastic complexity, which is a qualitative guide in tuning the parameters of action-selection strategy, and show a sufficient condition for return maximization in probability.

Algorithms↗

Adaptive internal state space construction method for reinforcement learning of a real-world agent.

One of the difficulties encountered in the application of the reinforcement learning to real-world problems is the construction of a discrete state space from a continuous sensory input signal. In the absence of a priori knowledge about the task, a straightforward approach to this problem is to discretize the input space into a grid, and to use a lookup table. However, this method suffers from the curse of dimensionality. Some studies use continuous function approximators such as neural networks instead of lookup tables. However, when global basis functions such as sigmoid functions are used, convergence cannot be guaranteed. To overcome this problem, we propose a method in which local basis functions are incrementally assigned depending on the task requirement. Initially, only one basis function is allocated over the entire space. The basis function is divided according to the statistical property of locally weighted temporal difference error (TD error) of the value function. We applied this method to an autonomous robot collision avoidance problem, and evaluated the validity of the algorithm in simulation. The proposed algorithm, which we call adaptive basis division (ABD) algorithm, achieved the task using a smaller number of basis functions than the conventional methods. Moreover, we applied the method to a goal-directed navigation problem of a real mobile robot. The action strategy was learned using a database of sensor data, and it was then used for navigation of a real machine. The robot reached the goal using a smaller number of internal states than with the conventional methods.

Journal Article↗

Hippocampal stimulation limits performance but does not retard acquisition of food-reinforced learning.

Facilitation of the usually very slow acquisition of hippocampal (HPC) self-stimulation (SS) by prior, non-contingent HPC stimulation may be due to the progressive attenuation of a disruptive effect of the stimulation on learning. We attempted to answer two questions: (1) Does HPC stimulation disrupt food-reinforced learning when it follows each lever-press as in SS experiments? (2) Does prior exposure to the stimulation attenuate any such disruptive effect? We observed no significant differences in the rate of acquisition when each food-reinforced response was paired contingently with a 0.5 sec train of dorsal HPC stimulation (CS group), or when 0.5 sec trains were administered randomly throughout the session (RS), compared with an implanted control group (IC). Although acquisition was rapid in all groups, performance on the second day was significantly lower in the CS group than in the IC animals. These same electrode placements later supported SS. The same results were obtained with 3 similarly-treated groups that had previously received a program of daily HPC stimulation (kindling). The results imply that kindling does not produce its facilitating effect on acquisition of HPC SS by removing a disruptive effect of the stimulation.

Animals↗

Fuzzy OLAP association rules mining-based modular reinforcement learning approach for multiagent systems.

Multiagent systems and data mining have recently attracted considerable attention in the field of computing. Reinforcement learning is the most commonly used learning process for multiagent systems. However, it still has some drawbacks, including modeling other learning agents present in the domain as part of the state of the environment, and some states are experienced much less than others, or some state-action pairs are never visited during the learning phase. Further, before completing the learning process, an agent cannot exhibit a certain behavior in some states that may be experienced sufficiently. In this study, we propose a novel multiagent learning approach to handle these problems. Our approach is based on utilizing the mining process for modular cooperative learning systems. It incorporates fuzziness and online analytical processing (OLAP) based mining to effectively process the information reported by agents. First, we describe a fuzzy data cube OLAP architecture which facilitates effective storage and processing of the state information reported by agents. This way, the action of the other agent, not even in the visual environment. of the agent under consideration, can simply be predicted by extracting online association rules, a well-known data mining technique, from the constructed data cube. Second, we present a new action selection model, which is also based on association rules mining. Finally, we generalize not sufficiently experienced states, by mining multilevel association rules from the proposed fuzzy data cube. Experimental results obtained on two different versions of a well-known pursuit domain show the robustness and effectiveness of the proposed fuzzy OLAP mining based modular learning approach. Finally, we tested the scalability of the approach presented in this paper and compared it with our previous work on modular-fuzzy Q-learning and ordinary Q-learning.

Algorithms↗

Effect of retraining trials on memory consolidation in weakly reinforced learning.

Day-old chicks given a single weakly reinforced (20% v/v methyl anthranilate in absolute ethanol) passive avoidance learning trial showed no evidence of long-term memory. A second learning trial given at 15 minutes after initial resulted in consolidation of the learning experience into long-term memory. The retention function resulting from two learning trials is similar to that observed with a single strongly reinforced learning trial, and consists of the stages postulated by Gibbs and Ng. With a dilution of 10% methyl anthranilate in ethanol, four training trials were needed to yield unequivocal evidence of long-term memory consolidation.

2,4-Dinitrophenol↗

Perspectives of probabilistic inferences: Reinforcement learning and an adaptive network compared.

The assumption that people possess a strategy repertoire for inferences has been raised repeatedly. The strategy selection learning theory specifies how people select strategies from this repertoire. The theory assumes that individuals select strategies proportional to their subjective expectations of how well the strategies solve particular problems; such expectations are assumed to be updated by reinforcement learning. The theory is compared with an adaptive network model that assumes people make inferences by integrating information according to a connectionist network. The network's weights are modified by error correction learning. The theories were tested against each other in 2 experimental studies. Study 1 showed that people substantially improved their inferences through feedback, which was appropriately predicted by the strategy selection learning theory. Study 2 examined a dynamic environment in which the strategies' performances changed. In this situation a quick adaptation to the new situation was not observed; rather, individuals got stuck on the strategy they had successfully applied previously. This "inertia effect" was most strongly predicted by the strategy selection learning theory.

Adult↗

CRISPGen: A deep generative framework for multi-objective CRISPR/Cas9 guide RNA design via Conditional Latent Diffusion and Dual-Critic Reinforcement Learning.

MOTIVATION: The CRISPR-Cas9 system offers transformative potential for precision genome editing, yet its clinical translation remains constrained by the risk of unintended off-target double-strand breaks. While current discriminative models excel at evaluating pre-specified candidate guides, resolving the fundamental antagonism between on-target cleavage efficiency and off-target specificity within a fixed sequence search space remains a major challenge. RESULTS: We present CRISPGen, a unified deep generative framework that reframes sgRNA design as a multi-objective constrained sequence synthesis problem. It integrates (i) DNABERT-2 genomic-language embeddings, (ii) a conditional latent diffusion generator conditioned on a user-specified on-target efficiency target, and (iii) a dual-critic reinforcement-learning (RL) stage that couples a frozen on-target efficiency critic with a cross-attention off-target discriminator (validation Pearson R=0.8157) trained on a unified corpus of experimental off-target events from six detection platforms. Across 1000 generated sgRNAs, CRISPGen reduces the mean off-target discriminator score by 99.7% relative to the pre-RL baseline and, under an exhaustive whole-genome screen of all 302,631,056 NGG PAM sites in GRCh38, yields zero perfect-match and only 55 one-mismatch genomic hits. We further show, transparently, that the internal on-target critic saturates under RL optimization - an instance of Goodhart's Law - and therefore assess on-target viability using an independent external CRISPRon screen (mean 47.10/100). Repeating the RL fine-tuning stage under three random seeds (with the diffusion generator, DNABERT-2 embeddings, and off-target discriminator held fixed) yields a stable operating point across seeds. Full diversity, per-mismatch, and reproducibility statistics are reported in the Results. AVAILABILITY: Source code is available at https://github.com/malekpouri/CRISPGen; the pre-trained checkpoints and the 3,000,000-sequence library are hosted on Hugging Face (https://huggingface.co/malekpouri/CRISPGen-Checkpoints) and archived on Zenodo under DOI 10.5281/zenodo.21428641.

CRISPR-Cas9↗

By carrot or by stick: cognitive reinforcement learning in parkinsonism.

To what extent do we learn from the positive versus negative outcomes of our decisions? The neuromodulator dopamine plays a key role in these reinforcement learning processes. Patients with Parkinson's disease, who have depleted dopamine in the basal ganglia, are impaired in tasks that require learning from trial and error. Here, we show, using two cognitive procedural learning tasks, that Parkinson's patients off medication are better at learning to avoid choices that lead to negative outcomes than they are at learning from positive outcomes. Dopamine medication reverses this bias, making patients more sensitive to positive than negative outcomes. This pattern was predicted by our biologically based computational model of basal ganglia-dopamine interactions in cognition, which has separate pathways for "Go" and "NoGo" responses that are differentially modulated by positive and negative reinforcement.

Aged↗

Reinforcement learning in random neural networks for cascaded decisions.

The Random Neural Network (RNN) model, in which signals travel as voltage spikes rather than as fixed signal levels, represents more closely the manner in which signals are transmitted in biophysical neural networks. In this paper a reinforcement learning strategy is proposed to make a sequence of cascaded decisions to achieve a goal while aiming to optimize the total cost of the cascaded decisions. For this purpose, RANs are used to model the system and a weight update rule together with a reinforcement function is provided. The performance of the learning strategy is analysed by applying it to the maze learning problem. The simulation results show that the performance of the system is highly dependent on the chosen reinforcement function and quite satisfactory results are obtained when the reinforcement function takes the recency effect into consideration.

Action Potentials↗

On adaptation, maximization, and reinforcement learning among cognitive strategies.

Analysis of binary choice behavior in iterated tasks with immediate feedback reveals robust deviations from maximization that can be described as indications of 3 effects: (a) a payoff variability effect, in which high payoff variability seems to move choice behavior toward random choice; (b) underweighting of rare events, in which alternatives that yield the best payoffs most of the time are attractive even when they are associated with a lower expected return; and (c) loss aversion, in which alternatives that minimize the probability of losses can be more attractive than those that maximize expected payoffs. The results are closer to probability matching than to maximization. Best approximation is provided with a model of reinforcement learning among cognitive strategies (RELACS). This model captures the 3 deviations, the learning curves, and the effect of information on uncertainty avoidance. It outperforms other models in fitting the data and in predicting behavior in other experiments.

Adaptation, Psychological↗

Cognitive navigation based on nonuniform Gabor space sampling, unsupervised growing networks, and reinforcement learning.

We study spatial learning and navigation for autonomous agents. A state space representation is constructed by unsupervised Hebbian learning during exploration. As a result of learning, a representation of the continuous two-dimensional (2-D) manifold in the high-dimensional input space is found. The representation consists of a population of localized overlapping place fields covering the 2-D space densely and uniformly. This space coding is comparable to the representation provided by hippocampal place cells in rats. Place fields are learned by extracting spatio-temporal properties of the environment from sensory inputs. The visual scene is modeled using the responses of modified Gabor filters placed at the nodes of a sparse Log-polar graph. Visual sensory aliasing is eliminated by taking into account self-motion signals via path integration. This solves the hidden state problem and provides a suitable representation for applying reinforcement learning in continuous space for action selection. A temporal-difference prediction scheme is used to learn sensorimotor mappings to perform goal-oriented navigation. Population vector coding is employed to interpret ensemble neural activity. The model is validated on a mobile Khepera miniature robot.

Cognition↗

Computer simulation of FES standing up in paraplegia: a self-adaptive fuzzy controller with reinforcement learning.

Using computer simulation, the theoretical feasibility of functional electrical stimulation (FES) assisted standing up is demonstrated using a closed-loop self-adaptive fuzzy logic controller based on reinforcement machine learning (FLC-RL). The control goal was to minimize upper limb forces and the terminal velocity of the knee joint. The reinforcement learning (RL) technique was extended to multicontroller problems in continuous state and action spaces. The validated algorithms were used to synthesize FES controllers for the knee and hip joints in simulated paraplegic standing up. The FLC-RL controller was able to achieve the maneuver with only 22% of the upper limb force required to stand-up without FES and to simultaneously reduce the terminal velocity of the knee joint close to zero. The FLC-RL controller demonstrated, as expected, the closed loop fuzzy logic control and on-line self-adaptation capability of the RL was able to accommodate for simulated disturbances due to voluntary arm forces, FES induced muscle fatigue and anthropometric differences between individuals. A method of incorporating a priori heuristic rule based knowledge is described that could reduce the number of the learning trials required to establish a usable control strategy. We also discuss how such heuristics may also be incorporated into the initial FLC-RL controller to ensure safe operation from the onset.

Arm↗

Using reinforcement learning to understand the emergence of "intelligent" eye-movement behavior during reading.

The eye movements of skilled readers are typically very regular (K. Rayner, 1998). This regularity may arise as a result of the perceptual, cognitive, and motor limitations of the reader (e.g., limited visual acuity) and the inherent constraints of the task (e.g., identifying the words in their correct order). To examine this hypothesis, reinforcement learning was used to allow an artificial "agent" to learn to move its eyes to read as efficiently as possible. The resulting patterns of simulated eye movements resembled those of skilled readers and suggest that important aspects of eye-movement behavior might emerge as a consequence of satisfying the constraints that are imposed on readers. These results also suggest novel interpretations of some contentious empirical results, such as the fixation duration costs associated with word skipping (R. Kliegl & R. Engbert, 2005), and theoretical assumptions, for example the familiarity check in the E-Z Reader model of eye-movement control (E. D. Reichle, A. Pollatsek, D. L. Fisher, & K. Rayner, 1998).

Algorithms↗

Kalman filter control embedded into the reinforcement learning framework.

There is a growing interest in using Kalman filter models in brain modeling. The question arises whether Kalman filter models can be used on-line not only for estimation but for control. The usual method of optimal control of Kalman filter makes use of off-line backward recursion, which is not satisfactory for this purpose. Here, it is shown that a slight modification of the linear-quadratic-gaussian Kalman filter model allows the on-line estimation of optimal control by using reinforcement learning and overcomes this difficulty. Moreover, the emerging learning rule for value estimation exhibits a Hebbian form, which is weighted by the error of the value estimation.

Algorithms↗

A fuzzy reinforcement learning approach to power control in wireless transmitters.

We address the issue of power-controlled shared channel access in wireless networks supporting packetized data traffic. We formulate this problem using the dynamic programming framework and present a new distributed fuzzy reinforcement learning algorithm (ACFRL-2) capable of adequately solving a class of problems to which the power control problem belongs. Our experimental results show that the algorithm converges almost deterministically to a neighborhood of optimal parameter values, as opposed to a very noisy stochastic convergence of earlier algorithms. The main tradeoff facing a transmitter is to balance its current power level with future backlog in the presence of stochastically changing interference. Simulation experiments demonstrate that the ACFRL-2 algorithm achieves significant performance gains over the standard power control approach used in CDMA2000. Such a large improvement is explained by the fact that ACFRL-2 allows transmitters to learn implicit coordination policies, which back off under stressful channel conditions as opposed to engaging in escalating "power wars."

Algorithms↗