Error controlled actor-critic

Publisher:
Elsevier
Publication Type:
Journal Article
Citation:
Information Sciences, 2022, 612, pp. 62-74
Issue Date:
2022-10
Filename Description Size
1-s2.0-S0020025522009896-main.pdfPublished version2.09 MB
Adobe PDF
Full metadata record
The approximation inaccuracy of the value function in reinforcement learning (RL) algorithms unavoidably leads to an overestimation phenomenon, which has negative effects on the convergence of the algorithms. To limit the negative effects of the approximation error, we propose error controlled actor-critic (ECAC) which ensures the approximation error is limited within the value function. We present an investigation of how approximation inaccuracy can impair the optimization process of actor-critic approaches. In addition, we derive an upper bound for the approximation error of the Q function approximator and discover that the error can be reduced by limiting the KL- divergence between every two consecutive policies during policy training. Experiments on a variety of continuous control tasks demonstrate that the proposed actor-critic approach decreases approximation error and outperforms previous model-free RL algorithms by a significant margin.
Please use this identifier to cite or link to this item: