Outline

Ingegneria Sismica

Ingegneria Sismica

Task-Consistent Bayesian Domain Inference via Performance Distributions for Deep Reinforcement Learning Policy Deployment

Author(s): Xiang Fu1, Kewei Chen1
1Faculty of Mechanical Engineering & Mechanics, Ningbo University, Ningbo, 315211, China
Fu, Xiang. and Chen, Kewei. “Task-Consistent Bayesian Domain Inference via Performance Distributions for Deep Reinforcement Learning Policy Deployment.” Ingegneria Sismica Volume 43 Issue 3: 1-23, doi:10.65102/is20261261.

Abstract

Although Deep reinforcement learning (DRL) has achieved remarkable success in robotics, the policies learnt in simulation often experience severe performance drops in the real world because of the reality gap. To reduce this performance gap, we propose a Task-Consistent Bayesian Inference (TCBI) framework for sim-to-real transfer. Rather than relying on intractable dynamics likelihoods or matching on high-dimensional trajectories, TCBI builds the task-level pseudo-likelihood based on the divergence of simulated and real performance distribution. In our formulation, the reward statistics, the body posture distributions and the contact time ratios are all compact task-oriented performance statistics of the distribution that characterizes task-specific domain discrepancy. Our design thus supports likelihood-free Bayesian inference effectively and robustly and also improves computational efficiency. We demonstrate the proposed framework on a six-legged robot in both balance task and forward locomotion task. Our experimental results show that TCBI always lowers the reward distribution disparity and also achieves better real-world performance than domain randomization, ABC, and simulation optimization (SimOpt) do. Ablation studies further show that incorporating reward, posture, and contact statistics can further improve the posterior identifiability and policy stability compared with using the reward distributions. Moreover, posterior variance analysis tells us that parameter concentrations in the inference process are progressive, and wall-clock time comparison also demonstrates that the computational cost of the method is much lower than that of trajectory-based methods. Robustness experiments under sensor noises further verify the stability and generalization capability of the method that we presented. All of these experimental results clearly point out that task-level probabilistic inference gives us an efficient, robust and scalable solution for sim-to-real deployment of reinforcement learning methods.

Keywords
Deep Reinforcement Learning; Sim-to-Real Transfer; Domain Adaptation; Bayesian Inference; Simulation Optimization

Related Articles

Zhihao Jiang1,2, Limi Chen1,2, Jing Yang1
1Hainan Vocational University of Science and Technology, Haikou 571126, China
2Institute for Mathematical Research, Universiti Putra Malaysia, Serdang 43400, Malaysia
Limi Chen1,2, Zhihao Jiang1,2, Jing Yang1
1Hainan Vocational University of Science and Technology, Haikou 571126, China
2Institute for Mathematical Research, Universiti Putra Malaysia, Serdang 43400, Malaysia
Hui Yuan1, Minjie Chai2, Siqing Xu1, Jinsong Li1, Jinwan Zheng1
1Electric Power Research Institute, State Grid Shanxi Electric Power Co., Ltd., Taiyuan, 030001, Shanxi, China
2Jincheng Power Supply Branch, State Grid Shanxi Electric Power Co., Ltd., Jincheng, 048000, Shanxi, China
Yanhan Zhu1,2
1China Academy of Cultural Heritage, Chaoyang District, 100029, Beijing, China
2Beijing University of Civil Engineering and Architecture, Xicheng District, 100044, Beijing, China
Ken Wang1, Jinhan Shu2, Kan Yuan1
1School of Digital Media, Shenzhen Polytechnic University, Shenzhen 518055, Guangdong, China
2Postdoctoral Mobile Station of Journalism and communication, Fudan University, Shanghai 200433, Shanghai, China