PubMed Health⌕ Search

SEARCH · PubMed Health

Results for “Reinforcement learning”

Explore indexed PubMed citations for clinical trials, systematic reviews and public health research. Read source abstracts and follow each citation to its original PubMed record.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17Linked to original sources

Postsynaptic integration of glutamatergic and dopaminergic signals in the striatum.

The aim of this study was to achieve a better understanding of the integration in striatal medium-sized spiny neurons (MSNs) of converging signals from glutamatergic and dopaminergic afferents. The review of the literature in the first section shows that these two types of afferents not only contact the same striatal cell type, but that individual MSNs receive both a corticostriatal and a dopaminergic terminal. The most common sites of convergence are dendritic shafts and spines of MSNs with a distance between the terminals of less than 1-2 microns. The second section focuses on synaptic transmission and second messenger activation. Glutamate, the candidate transmitter of corticostriatal terminals, via different types of glutamate receptors can evoke an increase in intracellular free calcium concentrations. The net effect of dopamine in the striatum is a stimulation of adenylate cyclase activity leading to an increase in cAMP. The subsequent sections present information on calcium- and cAMP-sensitive biochemical pathways and review the regional and subcellular distribution of the components in the striatum. The specific biochemical reaction steps were formalized as simplified equilibrium equations. Parameter values of the model were chosen from published experimental data. Major results of this analysis are: at intracellular free calcium concentrations below 1 microM the stimulation of adenylate cyclase by calcium and dopamine is at least additive in the steady state. Free calcium concentrations exceeding 1 microM inhibit adenylate cyclase, which is not overcome by dopaminergic stimulation. The kinases and phosphatases studied can be divided in those that are almost exclusively calcium-sensitive (PP2B and CaMPK), and others that are modulated by both calcium and dopamine (PKA and PP1). Maximal threonine-phosphorylation of the phosphoprotein DARPP requires optimal concentrations of calcium (about 0.3 microM) and dopamine (above 5 microM). It seems favourable if the glutamate signal precedes phasic dopamine release by approximately 100 msec. The phosphorylation of MAP2 is under essentially calcium-dependent control of at least five kinases and phosphatases, which differentially affect its heterogeneous phosphorylation sites. Therefore, MAP2 could respond specifically to the spatio-temporal characteristics of different intracellular calcium fluxes. The quantitative description of the calcium- and dopamine-dependent regulation of DARPP and MAP2 provides insights into the crosstalk between glutamatergic and dopaminergic signals in striatal MSNs. Such insights constitute an important step towards a better understanding of the links between biochemical pathways, physiological processes, and behavioural consequences connected with striatal function. The relevance to long-term potentiation, reinforcement learning, and Parkinson's disease is discussed.

Afferent Pathways↗

Intrastriatal injection of choline accelerates the acquisition of positively rewarded behaviors.

The prediction was made that by increasing the synthesis of striatal acetylcholine, through local injection of its precursor choline, the acquisition of a lever-pressing response in two different autoshaping situations would be accelerated. In the first experiment, choline was injected into the striatum or parietal cortex of rats immediately after dipper training; 24 h later and during 5 consecutive days the animals were submitted to an autoshaping procedure of the operant kind. In the second experiment, choline was administered to the same regions shortly after each of three classical-operant autoshaping sessions; during the next two sessions, autoshaping contingencies of the operant kind were in effect. In both experiments choline injection into the striatum induced a marked facilitation of acquisition of the conditioned responses, although cortical injection of choline produced a milder improvement only in the first experiment. These results indicate that striatal cholinergic activity is, indeed, involved in the early phases of positively reinforced learning.

Animals↗

Synaptic plasticity and dysconnection in schizophrenia.

Current pathophysiological theories of schizophrenia highlight the role of altered brain connectivity. This dysconnectivity could manifest 1) anatomically, through structural changes of association fibers at the cellular level, and/or 2) functionally, through aberrant control of synaptic plasticity at the synaptic level. In this article, we review the evidence for these theories, focusing on the modulation of synaptic plasticity. In particular, we discuss how dysconnectivity, observed between brain regions in schizophrenic patients, could result from abnormal modulation of N-methyl-D-aspartate (NMDA)-dependent plasticity by other neurotransmitter systems. We focus on the implication of the dysconnection hypothesis for functional imaging at the systems level. In particular, we review recent advances in measuring plasticity in the human brain using functional magnetic resonance imaging (fMRI) and electroencephalography (EEG) that can be used to address dysconnectivity in schizophrenia. Promising experimental paradigms include perceptual and reinforcement learning. We describe how theoretical and causal models of brain responses might contribute to a mechanistic understanding of synaptic plasticity in schizophrenia.

Acetylcholine↗

Mechanistic insights into Claudin-14 dysfunction implicated in veins of Galen malformation.

Claudin-14 (CLDN14) is a key component of tight junctions (TJs) critical for maintaining paracellular barrier function. Variants of CLDN14 have been linked to Vein of Galen malformations (VOGMs), a rare cerebrovascular disorder; however, the molecular mechanisms underlying their pathogenicity remain unknown. Here, we investigate the mechanistic effects of two VOGM-associated mutations, A113P and V143M, using reinforcement-learning driven enhanced sampling molecular dynamics simulations combined with DiffNets-based deep learning and independent trajectory-wide structural analyses. Our analysis reveals that A113P induces broader structural disruption of CLDN14, perturbing paracellular sealing, pore symmetry, and inter-protomer communication, whereas V143M induces structural rearrangements centred around TM3 and the TM3-ECL2 region. In both cases, mutation-specific alterations are observed in structural stability and interfacial organization across oligomeric assemblies. Notably, these effects are qualitatively consistent across different modelled architectures, despite variability in local responses. In the absence of experimentally resolved structures, the structural perturbations reported here provide a mechanistic understanding of how VOGM-associated variants may influence CLDN14 structure and dynamics.

Aneurysm↗

Connexin mRNA expression in single dopaminergic neurons of substantia nigra pars compacta.

Dopaminergic neurons of the substantia nigra pars compacta play a major role in goal-directed behavior and reinforcement learning. The study of their local interactions has revealed that they are connected by electrical synapses. Connexins, the molecular substrate of electrical synapses, constitute a multigenic family of 20 proteins in rodents. The permeability and regulation properties of electrical synapses depend on their connexin composition. Therefore, the knowledge of the molecular composition of electrical synapses is fundamental to the understanding of their specific functions. We have investigated the connexin mRNA expression pattern of dopaminergic neurons by single-cell RT-PCR analysis, during two periods in which dopaminergic neurons are electrically coupled in vitro (P7-P10 and P17-P21). Our results show that dopaminergic neurons express mRNAs of various connexins (Cx26, Cx30, Cx31.1, Cx32, Cx36 and Cx43) in a developmentally regulated manner. Furthermore, we have observed that dopaminergic neurons display different connexin expression patterns, and that multiple connexins can be expressed in a single dopaminergic neuron. These observations underline the importance of electrical coupling in the development of dopaminergic neurons and raise the question of the existence of functionally distinct electrically coupled networks in the substantia nigra pars compacta.

Animals↗

Shape variability of the human striatum--Effects of age and gender.

Human striatum is involved in the regulation of movement, reinforcement, learning, reward, cognitive functioning, and addiction. Previous classical volumetric MRI studies have implicated age-, disease- and medication-related changes in striatal structures. Yet, no studies to date have addressed the effects of these factors on the shape variability and local structural alterations in the striatum. The local alterations may provide meaningful additional information in the context of functional neuroanatomy and brain connectivity. We developed image analysis methodology for the measurement of the volume and local shape variability of the human striatum. The method was applied in a group of 43 healthy controls to study the effects of age and gender on striatal shape variability. In the volume analysis, the volume of the striatum was normalized using the volume of the whole brain. In the local shape analysis, the deviations from a mean surface were studied for each surface point using high-dimensional mapping. Also, discriminant functions were constructed from a statistical shape model. The accuracy and reproducibility of the methods used were evaluated. The results confirmed that the volume of the striatum decreases as a function of age. However, the volume decrease was not uniform and age-related shape differences were observed in several subregions of the human striatum whereas no local gender differences were seen. Examination of the variability of striatal shape in the healthy population will pave the way for applying this method in clinical settings. This method will be particularly useful for investigating neuropsychiatric disorders that are associated with subtle morphological alterations of the brain, such as schizophrenia.

Adult↗

Eye movements in natural behavior.

The classic experiments of Yarbus over 50 years ago revealed that saccadic eye movements reflect cognitive processes. But it is only recently that three separate advances have greatly expanded our understanding of the intricate role of eye movements in cognitive function. The first is the demonstration of the pervasive role of the task in guiding where and when to fixate. The second has been the recognition of the role of internal reward in guiding eye and body movements, revealed especially in neurophysiological studies. The third important advance has been the theoretical developments in the fields of reinforcement learning and graphic simulation. All of these advances are proving crucial for understanding how behavioral programs control the selection of visual information.

Behavior↗

Strategic behavior in monkeys.

In a recent paper, Lee et al. examined adaptive decision-making processes by training monkeys to play a competitive game against a computer programmed to play using various strategies. They found that the monkeys' responses were sensitive to the computer's strategies and consistent with reinforcement learning. Research such as this strongly complements current research in behavioral economics. We propose some potential future directions for this work, and put forward conjectures about what might be learned about decision-making in humans.

Animals↗

A neural network approach to hippocampal function in classical conditioning.

Hippocampal participation in classical conditioning in terms of Grossberg's (1975) attentional theory is described. According to the present rendition of this theory, pairing of a conditioned stimulus (CS) with an unconditioned stimulus (US) causes both an association of the sensory representation of the CS with the US (conditioned reinforcement learning) and an association of the sensory representation of the CS with the drive representation of the US (incentive motivation learning). Sensory representations compete among themselves for a limited-capacity short-term memory (STM) that is reflected in a long-term memory storage. The STM regulation hypothesis, which proposes that the hippocampus controls incentive motivation, self-excitation, and competition among sensory representations thereby regulating the contents of a limited capacity STM, is introduced. Under the STM regulation hypothesis, nodes and connections in Grossberg's neural network are mapped onto regional hippocampal-cerebellar circuits. The resulting neural model provides (a) a framework for understanding the dynamics of information processing and storage in the hippocampus and cerebellum during classical conditioning of the rabbit's nictitating membrane, (b) principles for understanding the effect of different hippocampal manipulations on classical conditioning, and (c) numerous novel and testable predictions.

Animals↗

A nitric oxide agonist stimulates consolidation of long-term memory in the 1-day-old chick.

Chicks, age 1 to 2 days, that have been trained on a passive avoidance task with a strongly reinforced training trial yield a memory trace that is composed of 3 behaviorally and pharmacologically distinguishable stages, with the final long-term memory stage being dependent on protein synthesis. In contrast, chicks trained with a weakly reinforced learning trial typically do not demonstrate this final stage of memory. Sodium nitroprusside 150 microM intracranially administered immediately after a weak training trial promoted the formation of long-term memory, whereas saline did not. The results suggest that nitric oxide synthesis is either itself critical or stimulates other processes that are critical for the consolidation of long-term memory.

Age Factors↗

Evidence for striatal dopamine release during a video game.

Dopaminergic neurotransmission may be involved in learning, reinforcement of behaviour, attention, and sensorimotor integration. Binding of the radioligand 11C-labelled raclopride to dopamine D2 receptors is sensitive to levels of endogenous dopamine, which can be released by pharmacological challenge. Here we use 11C-labelled raclopride and positron emission tomography scans to provide evidence that endogenous dopamine is released in the human striatum during a goal-directed motor task, namely a video game. Binding of raclopride to dopamine receptors in the striatum was significantly reduced during the video game compared with baseline levels of binding, consistent with increased release and binding of dopamine to its receptors. The reduction in binding of raclopride in the striatum positively correlated with the performance level during the task and was greatest in the ventral striatum. These results show, to our knowledge for the first time, behavioural conditions under which dopamine is released in humans, and illustrate the ability of positron emission tomography to detect neurotransmitter fluxes in vivo during manipulations of behaviour.

Adult↗

Computational roles for dopamine in behavioural control.

Neuromodulators such as dopamine have a central role in cognitive disorders. In the past decade, biological findings on dopamine function have been infused with concepts taken from computational theories of reinforcement learning. These more abstract approaches have now been applied to describe the biological algorithms at play in our brains when we form value judgements and make choices. The application of such quantitative models has opened up new fields, ripe for attack by young synthesizers and theoreticians.

Animals↗

Cortical substrates for exploratory decisions in humans.

Decision making in an uncertain environment poses a conflict between the opposing demands of gathering and exploiting information. In a classic illustration of this 'exploration-exploitation' dilemma, a gambler choosing between multiple slot machines balances the desire to select what seems, on the basis of accumulated experience, the richest option, against the desire to choose a less familiar option that might turn out more advantageous (and thereby provide information for improving future decisions). Far from representing idle curiosity, such exploration is often critical for organisms to discover how best to harvest resources such as food and water. In appetitive choice, substantial experimental evidence, underpinned by computational reinforcement learning (RL) theory, indicates that a dopaminergic, striatal and medial prefrontal network mediates learning to exploit. In contrast, although exploration has been well studied from both theoretical and ethological perspectives, its neural substrates are much less clear. Here we show, in a gambling task, that human subjects' choices can be characterized by a computationally well-regarded strategy for addressing the explore/exploit dilemma. Furthermore, using this characterization to classify decisions as exploratory or exploitative, we employ functional magnetic resonance imaging to show that the frontopolar cortex and intraparietal sulcus are preferentially active during exploratory decisions. In contrast, regions of striatum and ventromedial prefrontal cortex exhibit activity characteristic of an involvement in value-based exploitative decision making. The results suggest a model of action selection under uncertainty that involves switching between exploratory and exploitative behavioural modes, and provide a computationally precise characterization of the contribution of key decision-related brain systems to each of these functions.

Choice Behavior↗

Uncertainty-based competition between prefrontal and dorsolateral striatal systems for behavioral control.

A broad range of neural and behavioral data suggests that the brain contains multiple systems for behavioral choice, including one associated with prefrontal cortex and another with dorsolateral striatum. However, such a surfeit of control raises an additional choice problem: how to arbitrate between the systems when they disagree. Here, we consider dual-action choice systems from a normative perspective, using the computational theory of reinforcement learning. We identify a key trade-off pitting computational simplicity against the flexible and statistically efficient use of experience. The trade-off is realized in a competition between the dorsolateral striatal and prefrontal systems. We suggest a Bayesian principle of arbitration between them according to uncertainty, so each controller is deployed when it should be most accurate. This provides a unifying account of a wealth of experimental evidence about the factors favoring dominance by either system.

Animals↗