Adaptive Dynamic Programming for Discrete-Time Zero-Sum Games

In this paper, a novel adaptive dynamic programming (ADP) algorithm, called "iterative zero-sum ADP algorithm," is developed to solve infinite-horizon discrete-time two-player zero-sum games of nonlinear systems. The present iterative zero-sum ADP algorithm permits arbitrary positive semid...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transaction on neural networks and learning systems Jg. 29; H. 4; S. 957 - 969
Hauptverfasser:	Wei, Qinglai, Liu, Derong, Lin, Qiao, Song, Ruizhuo
Format:	Journal Article
Sprache:	Englisch
Veröffentlicht:	United States IEEE 01.04.2018 The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Schlagworte:	Adaptive algorithms Adaptive critic designs adaptive dynamic programming (ADP) Algorithms approximate dynamic programming Computer simulation Convergence Dynamic programming Equilibrium Game theory Games Iterative algorithms neurodynamic programming Nonlinear systems Optimal control Performance analysis Performance indices Saddle points zero-sum game
ISSN:	2162-237X, 2162-2388, 2162-2388
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	In this paper, a novel adaptive dynamic programming (ADP) algorithm, called "iterative zero-sum ADP algorithm," is developed to solve infinite-horizon discrete-time two-player zero-sum games of nonlinear systems. The present iterative zero-sum ADP algorithm permits arbitrary positive semidefinite functions to initialize the upper and lower iterations. A novel convergence analysis is developed to guarantee the upper and lower iterative value functions to converge to the upper and lower optimums, respectively. When the saddle-point equilibrium exists, it is emphasized that both the upper and lower iterative value functions are proved to converge to the optimal solution of the zero-sum game, where the existence criteria of the saddle-point equilibrium are not required. If the saddle-point equilibrium does not exist, the upper and lower optimal performance index functions are obtained, respectively, where the upper and lower performance index functions are proved to be not equivalent. Finally, simulation results and comparisons are shown to illustrate the performance of the present method.
Bibliographie:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 14 content type line 23
ISSN:	2162-237X 2162-2388 2162-2388
DOI:	10.1109/TNNLS.2016.2638863