论文信息 - Prioritized Sweeping: Reinforcement Learning with Less Data and Less Time

Prioritized Sweeping: Reinforcement Learning with Less Data and Less Time

We present a new algorithm, Prioritized Sweeping, for e cient prediction and control of stochastic Markov systems. Incremental learning methods such as Temporal Di erencing and Qlearning have fast real time performance. Classical methods are slower, but more accurate, because they make full use of the observations. Prioritized Sweeping aims for the best of both worlds. It uses all previous experiences both to prioritize important dynamic programming sweeps and to guide the exploration of state-space. We compare Prioritized Sweeping with other reinforcement learning schemes for a number of di erent stochastic optimal control problems. It successfully solves large state-space real time problems with which other methods have

C. Atkeson | A. Moore

[1] Arthur L. Samuel,et al. Some Studies in Machine Learning Using the Game of Checkers , 1967, IBM J. Res. Dev..

[2] G. Siouris,et al. Optimum systems control , 1979, Proceedings of the IEEE.

[3] Richard S. Sutton,et al. Temporal credit assignment in reinforcement learning , 1984 .

[4] David L. Waltz,et al. Toward memory-based reasoning , 1986, CACM.

[5] P. W. Jones,et al. Bandit Problems, Sequential Allocation of Experiments , 1987 .

[6] MITSUO SATO,et al. Learning control of finite Markov chains with an explicit trade-off between estimation and control , 1988, IEEE Trans. Syst. Man Cybern..

[7] A. Barto,et al. Learning and Sequential Decision Making , 1989 .

[8] John N. Tsitsiklis,et al. Parallel and distributed computation , 1989 .

[9] Richard E. Korf,et al. Real-Time Heuristic Search , 1990, Artif. Intell..

[10] Michael I. Jordan,et al. Advances in Neural Information Processing Systems 30 , 1995 .

[11] Richard S. Sutton,et al. Time-Derivative Models of Pavlovian Reinforcement , 1990 .