A Universal Empirical Dynamic Programming Algorithm for Continuous State MDPs

We propose universal randomized function approximation-based empirical value learning (EVL) algorithms for Markov decision processes. The "empirical" nature comes from each iteration being done empirically from samples available from simulations of the next state. This makes the Bellman op...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on automatic control Jg. 65; H. 1; S. 115 - 129
Hauptverfasser:	Haskell, William B., Jain, Rahul, Sharma, Hiteshi, Yu, Pengqian
Format:	Journal Article
Sprache:	Englisch
Veröffentlicht:	New York IEEE 01.01.2020 The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Schlagworte:	Algorithms Approximation Approximation algorithms Basis functions Complexity theory Computer simulation Continuous state-space Markov decision processes (MDPs) Convergence Dynamic programming dynamic programming (DP) Empirical analysis Function approximation Function space Heuristic algorithms Hilbert space Iterative methods Machine learning Markov processes Mathematical analysis Operators (mathematics) Probabilistic logic reinforcement learning (RL)
ISSN:	0018-9286, 1558-2523
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	We propose universal randomized function approximation-based empirical value learning (EVL) algorithms for Markov decision processes. The "empirical" nature comes from each iteration being done empirically from samples available from simulations of the next state. This makes the Bellman operator a random operator. A parametric and a nonparametric method for function approximation using a parametric function space and a reproducing kernel Hilbert space respectively are then combined with EVL. Both function spaces have the universal function approximation property. Basis functions are picked randomly. Convergence analysis is performed using a random operator framework with techniques from the theory of stochastic dominance. Finite time sample complexity bounds are derived for both universal approximate dynamic programming algorithms. Numerical experiments support the versatility and computational tractability of this approach.
Bibliographie:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 14
ISSN:	0018-9286 1558-2523
DOI:	10.1109/TAC.2019.2907414