论文信息 - Adaptive Bases for Reinforcement Learning

Adaptive Bases for Reinforcement Learning

We consider the problem of reinforcement learning using function approximation, where the approximating basis can change dynamically while interacting with the environment. A motivation for such an approach is maximizing the value function fitness to the problem faced. Three errors are considered: approximation square error, Bellman residual, and projected Bellman residual. Algorithms under the actorcritic framework are presented, and shown to converge. The advantage of such an adaptive basis is demonstrated in simulations.

Shie Mannor | Dotan Di Castro

[1] John N. Tsitsiklis,et al. Neuro-Dynamic Programming , 1996, Encyclopedia of Machine Learning.

[2] Richard S. Sutton,et al. Introduction to Reinforcement Learning , 1998 .

[3] Shalabh Bhatnagar,et al. Fast gradient-descent methods for temporal-difference learning with linear function approximation , 2009, ICML '09.

[4] Philippe Preux,et al. Basis Expansion in Natural Actor Critic Methods , 2008, EWRL.

[5] V. Borkar. Stochastic approximation with two time scales , 1997 .

[6] Shie Mannor,et al. Basis Function Adaptation in Temporal Difference Reinforcement Learning , 2005, Ann. Oper. Res..

[7] Dimitri P. Bertsekas,et al. Dynamic Programming and Optimal Control, Two Volume Set , 1995 .

[8] Martin L. Puterman,et al. Markov Decision Processes: Discrete Stochastic Dynamic Programming , 1994 .

[9] E. J. Collins,et al. Convergent multiple-timescales reinforcement learning algorithms in normal form games , 2003 .

[10] Philippe Preux,et al. Recent Advances in Reinforcement Learning: 8th European Workshop, EWRL 2008, Villeneuve d'Ascq, France, June 30-July 3, 2008, Revised and Selected Papers , 2008 .

[11] S. Hyakin,et al. Neural Networks: A Comprehensive Foundation , 1994 .