PubMed Health⌕ Search

Biomedical subjects

Danil V Prokhorov

Publications and source records attributed to Danil V Prokhorov.

3 recordsLinked to original sources

Training recurrent neurocontrollers for robustness with derivative-free Kalman filter.

We are interested in training neurocontrollers for robustness on discrete-time models of physical systems. Our neurocontrollers are implemented as recurrent neural networks (RNNs). A model of the system to be controlled is known to the extent of parameters and/or signal uncertainties. Parameter values are drawn from a known distribution. For each instance of the model with specified parameters, a recurrent neurocontroller is trained by evaluating sensitivities of the model outputs to perturbations of the neurocontroller weights and incrementally updating the weights. Our training process strives to minimize a quadratic cost function averaged over many different models. In the end, the process yields a robust recurrent neurocontroller, which is ready for deployment with fixed weights. We employ a derivative-free Kalman filter algorithm proposed by Norgaard et al. and extended by Feldkamp et al. (2001) and Feldkamp et al. (2002) to neural network training. Our training algorithm combines effectiveness of a second-order training method with universal applicability to both differentiable and nondifferentiable systems. Our approach is that of model reference control, and it extends significantly the capabilities proposed by Prokhorov et al. (2001). We illustrate it with two examples.

Algorithms↗

A model of evolution and learning.

We study a model of evolving populations of self-learning agents and analyze the interaction between learning and evolution. We consider an agent-broker that predicts stock price changes and uses its predictions for selecting actions. Each agent is equipped with a neural network adaptive critic design for behavioral adaptation. We discuss three cases in which either evolution or learning, or both, are active in our model. We show that the Baldwin effect can be observed in our model, viz. originally acquired adaptive policy of best agent-brokers becomes inherited over the course of the evolution. We also compare the behavioral tactics of our agents to the searching behavior of simple animals.

Algorithms↗

Simple and conditioned adaptive behavior from Kalman filter trained recurrent networks.

We illustrate the ability of a fixed-weight neural network, trained with Kalman filter methods, to perform tasks that are usually entrusted to an explicitly adaptive system. Following a simple example, we demonstrate that such a network can be trained to exhibit input-output behavior that depends on which of two conditioning tasks was performed a substantial number of time steps in the past. This behavior can also be made to survive an intervening interference task.

Adaptation, Psychological↗