PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6Linked to original sources

Reinforcement learning-based dynamic ensemble for missense variant effect prediction and tiered prioritization of VUS.

BACKGROUND: Accurate classification of missense variants remains a challenging task despite major advances in genomics. Numerous computational models have been developed to assist in variant classification, but often require repeated integration and benchmarking efforts. Ensemble methods have been proposed to overcome the limitations of single predictors, but mostly rely on fixed, predefined weights that constrain their ability to capture interactions among predictive signals. METHODS: We present GenixRL, a dynamic ensemble framework that reformulates model fusion as a reinforcement learning optimization problem. GenixRL uses a Q-learning agent to learn a policy that dynamically weights the probabilistic outputs of complementary predictors, including BayesDel (addAF and noAF), ClinPred, and MetaRNN. Replacing static weighting with policy learning allows GenixRL to adaptively identify optimal weightings and substantially improve classification accuracy. RESULTS: In benchmark evaluation against 25 state-of-the-art predictors, GenixRL achieved an AUROC of 0.9644 on an independent ClinVar dataset. On saturation genome editing assays for BRCA1 and BRCA2, GenixRL achieved the best performance and ranked highest on 14 of 17 clinically significant genes in a zero-shot evaluation. Applied to uncertain and conflicting ClinVar variants, GenixRL enabled tiered, evidence-based prioritization of hundreds of thousands of variants as likely pathogenic or pathogenic with high confidence, supported by orthogonal population evidence from gnomAD. CONCLUSION: GenixRL advances pathogenicity prediction for missense variants and provides an adaptive ensemble that sorts variants of uncertain significance into tiered candidates for expert curation and functional validation.

Mutation, Missense↗

[A study of relation between phrasing and intertrial interval in reinforcement pattern learning in rats].

Effects of the temporal interval on reinforcement pattern learning were investigated in four experiments using a runway. The reinforced trial was always the first trial for R5N sequence (Experiments 1a and 1b), and the fourth trial for 3NR2N sequence (Experiments 2a and 2b). Group S-ITI received the given sequence at 30-s ITI, Group L-INT received the same procedure as Group S-ITI except for a 30-min ITI inserted between Trials 3 and 4 (Experiments 1a and 2a). Group L-ITI received at 30-min ITI, Group S-INT did as Group L-ITI except for the ITI between Trials 3 and 4 was 30 seconds (Experiments 1b and 2b). It was found that the running speed on each of Trials 1 and 4 was faster than any other trials under R5N and 3NR2N sequences for Group L-INT. That is, the running pattern for Trials 1-3 was similar to that for Trials 4-6. On the other hand, the running speed on Trial 4 under R5N sequence and that on Trial 1 under 3NR2N sequence did not increase for Group S-INT. These results suggest that a longer or shorter ITI plays the roles of both phrasing cue and discriminative stimulus. For Group S-ITI and Group L-ITI, there was little evidence that sequences were phrased. Therefore, when trials are separated by equal ITIs, neither 30-s ITI nor 30-min ITI becomes a phrasing cue.

Animals↗

Anxiolytic-like action of neurokinin substance P administered systemically or into the nucleus basalis magnocellularis region.

There is evidence that the neurokinin substance P plays a role in neural mechanisms governing learning and reinforcement. Reinforcing and memory-promoting effects of substance P were found after it was injected into several parts of the brain and intraperitoneally. With regard to the close link between anxiety and memory processes for negative reinforcement learning, the aim of the present study was to gauge the effect of substance P on anxiety-related behaviors in the rat elevated plus-maze and social interaction test. Substance P was tested at injection sites where the neurokinin has been shown to promote learning and to serve as a reinforcer, namely in the periphery (after i.p. administration) and after injection into the nucleus basalis magnocellularis region. When administered i.p., substance P had a biphasic dose-response effect on behavior in the plus-maze with an anxiolytic-like action at 50 microg/kg and an anxiogenic-like one at 500 microg/kg. After unilateral microinjection into the nucleus basalis magnocellularis region, substance P (1 ng) was found to exert anxiolytic-like effects, because substance P-treated rats spent more time on the open arms of the plus-maze and showed an increase in time spent in social interaction. Furthermore, the anxiolytic effects of intrabasalis substance P were sequence-specific since injection of a compound with the inverse amino acid sequence of substance P (0.1 to 100 ng) did not influence anxiety parameters. These results show that substance P has anxiolytic-like properties in addition to its known promnestic and reinforcing effects, supporting the hypothesis of a close relationship between anxiety, memory and reinforcement processes.

Animals↗

A reinforcement learning-enhanced fuzzy multi-objective equilibrium optimization framework for multiple sequence alignment.

Multiple sequence alignment (MSA) is a fundamental task in bioinformatics, underpinning comparative genomics, structural analysis, and evolutionary inference. However, MSA remains a challenging multi-objective optimization problem due to the need to simultaneously maximize alignment accuracy, preserve conserved regions, and control gap proliferation, particularly in large and heterogeneous sequence collections. In this work, we propose MOFSACEO-MSA, a novel hybrid optimization framework for multiple sequence alignment that integrates a fuzzy multi-objective evaluation scheme with the Equilibrium Optimizer (EO) and a Soft Actor-Critic (SAC)-based adaptive control mechanism. The proposed framework formulates MSA as a dynamic multi-objective optimization problem, in which alignment quality is assessed using complementary residue-level and column-level criteria, including Sum-of-Pairs score, column conservation, entropy, and gap statistics. Fuzzy membership functions are employed to harmonize competing objectives into a unified optimization landscape, while EO provides robust global exploration. To further enhance adaptability, SAC dynamically regulates key EO parameters during the search process, enabling an effective balance between exploration and exploitation across datasets of varying size and heterogeneity. Extensive experiments werew conducted on diverse biological sequence datasets, with a primary focus on RNA benchmarks, including structured families from Rfam, large-scale repositories from RNAcentral and GenBank, and organism-specific tRNA datasets from GtRNAdb. Comparative evaluations against classical alignment tools (ClustalW, MAFFT, MUSCLE, PRANK, KAlign, and T-Coffee), metaheuristic methods (SAGA, Sequoya and EAFSA), and a reinforcement learning-based approach (RLALIGN) demonstrate that MOFSACEO-MSA consistently achieves competitive or superior Sum-of-Pairs scores while significantly reducing gap proportions and maintaining compact alignment lengths. Notably, the proposed framework exhibits improved robustness on large and highly heterogeneous datasets, where existing methods often suffer from excessive gap insertion or unstable convergence. Overall, MOFSACEO-MSA provides a flexible and extensible optimization paradigm that effectively bridges evolutionary search and reinforcement learning for high-quality multiple sequence alignment, with demonstrated effectiveness on challenging RNA alignment tasks.

Sequence Alignment↗

Evaluating theories of bird song learning: implications for future directions.

Studies of birdsong learning have stimulated extensive hypotheses at all levels of behavioral and physiological organization. This hypothesis building is valuable for the field and is consistent with the remarkable range of issues that can be rigorously addressed in this system. The traditional instructional (template) theory of song learning has been challenged on multiple fronts, especially at a behavioral level by evidence consistent with selectional hypotheses. In this review I highlight the caveats associated with these theories to better define the limits of our knowledge and identify important experiments for the future. The sites and representational forms of the various conceptual entities posited by the template theory are unknown. The distinction between instruction and selection in vocal learning is not well established at a mechanistic level. There is as yet insufficient neurophysiological data to choose between competing mechanisms of error-driven learning and reinforcement learning. Both may obtain for vocal learning. The possible role of sleep in acoustic or procedural memory consolidation, while supported by some physiological observations, does not yet have support in the behavioral literature. The remarkable expansion of knowledge in the past 20 years and the recent development of new technologies for physiological and behavioral experiments should permit direct tests of these theories in the coming decade.

Animal Communication↗

[Short intertrial interval as a phrasing cue in the reinforcement-pattern learning in rats].

The effect of short intertrial interval as a phrasing cue on the reinforcement-pattern learning was investigated in a rat experiment using a runway. In the acquisition session, 20 rats were trained for NNR sequence in which two nonreinforced trials were followed by a reinforced one with equal ITI of 30 min. In the sequence-addition session, a new sequence with consisting of three nonreinforced trials with 30-min ITI was added 30 min and 30 s after the original sequence of trials for Group L-ITI and for Group S-INT, respectively. Rats of Group S-INT responded on trials 4-6 with the similar pattern of running speed as they showed on Trials 1-3 from the second day. In Group L-ITI, the running speed on Trial 4 did not decrease. This suggests that Group S-INT rats learn to use the ITI of 30 s as a phrasing cue not on the first day but from the second day on. The phrasing effect of short ITI (30 s) was discussed in comparison with that of longer ITI (30 min).

Animals↗

Instrumental learning and relearning in individuals with psychopathy and in patients with lesions involving the amygdala or orbitofrontal cortex.

Previous work has shown that individuals with psychopathy are impaired on some forms of associative learning, particularly stimulus-reinforcement learning (Blair et al., 2004; Newman & Kosson, 1986). Animal work suggests that the acquisition of stimulus-reinforcement associations requires the amygdala (Baxter & Murray, 2002). Individuals with psychopathy also show impoverished reversal learning (Mitchell, Colledge, Leonard, & Blair, 2002). Reversal learning is supported by the ventrolateral and orbitofrontal cortex (Rolls, 2004). In this paper we present experiments investigating stimulus-reinforcement learning and relearning in patients with lesions of the orbitofrontal cortex or amygdala, and individuals with developmental psychopathy without known trauma. The results are interpreted with reference to current neurocognitive models of stimulus-reinforcement learning, relearning, and developmental psychopathy.

Adult↗

Astrocytic energy metabolism consolidates memory in young chicks.

In a single trial discrimination avoidance learning task, chicks learn to distinguish between beads of two colors, which are dipped in either a strong or weak tasting aversant (methyl anthranilate) to induce strongly-reinforced and weakly-reinforced learning, respectively. Consolidation of strongly-reinforced learning can be prevented by inhibitors of glycolysis, such as 2-deoxyglucose and iodoacetate and by inhibitors of oxidative metabolism and the consolidation of weakly-reinforced learning can be promoted by administration of glucose. In the present study we show that bilateral, intracerebral injection of 30 nmol acetate can act like glucose to consolidate labile memory and to restore memory impaired by 2-deoxyglucose administration. Acetate is a metabolic substrate that feeds into the tricarboxylic acid cycle, it is oxidized in astrocytes, but not in neurones. Our data suggest that effects of glucose administered 15-25 min post-training on memory consolidation are mediated via astrocytes not neurons.

Acetates↗

What are the computations of the cerebellum, the basal ganglia and the cerebral cortex?

The classical notion that the cerebellum and the basal ganglia are dedicated to motor control is under dispute given increasing evidence of their involvement in non-motor functions. Is it then impossible to characterize the functions of the cerebellum, the basal ganglia and the cerebral cortex in a simplistic manner? This paper presents a novel view that their computational roles can be characterized not by asking what are the "goals" of their computation, such as motor or sensory, but by asking what are the "methods" of their computation, specifically, their learning algorithms. There is currently enough anatomical, physiological, and theoretical evidence to support the hypotheses that the cerebellum is a specialized organism for supervised learning, the basal ganglia are for reinforcement learning, and the cerebral cortex is for unsupervised learning.This paper investigates how the learning modules specialized for these three kinds of learning can be assembled into goal-oriented behaving systems. In general, supervised learning modules in the cerebellum can be utilized as "internal models" of the environment. Reinforcement learning modules in the basal ganglia enable action selection by an "evaluation" of environmental states. Unsupervised learning modules in the cerebral cortex can provide statistically efficient representation of the states of the environment and the behaving system. Two basic action selection architectures are shown, namely, reactive action selection and predictive action selection. They can be implemented within the anatomical constraint of the network linking these structures. Furthermore, the use of the cerebellar supervised learning modules for state estimation, behavioral simulation, and encapsulation of learned skill is considered. Finally, the usefulness of such theoretical frameworks in interpreting brain imaging data is demonstrated in the paradigm of procedural learning.

Journal Article↗

Olfactory learning induces differential long-lasting changes in rat central olfactory pathways.

In the present work, we investigated lasting changes induced by olfactory learning at different levels of the olfactory pathways. For this, evoked field potentials induced by electrical stimulation of the olfactory bulb were recorded simultaneously in the anterior piriform cortex, the posterior piriform cortex, the lateral entorhinal cortex and the dentate gyrus. The amplitude of the evoked field potential's main component was measured in each site before, immediately after, and 20 days after completion of associative learning. Evoked field potential recordings were carried out under two experimental conditions in the same animals: awake and anesthetized. In the learning task, rats were trained to associate electrical stimulation of one olfactory bulb electrode with the delivery of sucrose (positive reward), and stimulation of a second olfactory bulb electrode with the delivery of quinine (negative reward). In this way, stimulation of the same olfactory bulb electrodes used for inducing field potentials served as a discriminative cue in the learning paradigm. The data showed that positively reinforced learning resulted in a lasting increase in evoked field potential amplitude restricted to posterior piriform cortex and lateral entorhinal cortex. In contrast, negatively reinforced learning was mainly accompanied by a decrease in evoked field potential amplitude in the dentate gyrus. Moreover, the expression of these learning-related changes occurred to be modulated by the animals arousal state. Indeed, the comparison between anesthetized versus awake animals showed that although globally similar, the changes were expressed earlier with respect to learning, under anesthesia than in the awake state. From these data we suggest that associative olfactory learning involves different neural circuits depending on the acquired value of the stimulus. Furthermore, they show the existence of a functional dissociation between anterior and posterior piriform cortex in mnesic processes, and stress the importance of the animal's arousal state on the expression of learning-induced plasticity.

Anesthetics↗

TD models of reward predictive responses in dopamine neurons.

This article focuses on recent modeling studies of dopamine neuron activity and their influence on behavior. Activity of midbrain dopamine neurons is phasically increased by stimuli that increase the animal's reward expectation and is decreased below baseline levels when the reward fails to occur. These characteristics resemble the reward prediction error signal of the temporal difference (TD) model, which is a model of reinforcement learning. Computational modeling studies show that such a dopamine-like reward prediction error can serve as a powerful teaching signal for learning with delayed reinforcement, in particular for learning of motor sequences. Several lines of evidence suggest that dopamine is also involved in 'cognitive' processes that are not addressed by standard TD models. I propose the hypothesis that dopamine neuron activity is crucial for planning processes, also referred to as 'goal-directed behavior', which select actions by evaluating predictions about their motivational outcomes.

Animals↗