Error controlled actor-critic
- Publisher:
- Elsevier
- Publication Type:
- Journal Article
- Citation:
- Information Sciences, 2022, 612, pp. 62-74
- Issue Date:
- 2022-10
Closed Access
Filename | Description | Size | |||
---|---|---|---|---|---|
1-s2.0-S0020025522009896-main.pdf | Published version | 2.09 MB |
Copyright Clearance Process
- Recently Added
- In Progress
- Closed Access
This item is closed access and not available.
The approximation inaccuracy of the value function in reinforcement learning (RL) algorithms unavoidably leads to an overestimation phenomenon, which has negative effects on the convergence of the algorithms. To limit the negative effects of the approximation error, we propose error controlled actor-critic (ECAC) which ensures the approximation error is limited within the value function. We present an investigation of how approximation inaccuracy can impair the optimization process of actor-critic approaches. In addition, we derive an upper bound for the approximation error of the Q function approximator and discover that the error can be reduced by limiting the KL- divergence between every two consecutive policies during policy training. Experiments on a variety of continuous control tasks demonstrate that the proposed actor-critic approach decreases approximation error and outperforms previous model-free RL algorithms by a significant margin.
Please use this identifier to cite or link to this item: