Human-machine collaboration is a promising training framework aimed at learning optimal strategies in high-cost exploration scenarios. However, such work is challenging. On one hand, current research on human-machine collaboration primarily focuses on imitation learning, overlooking the optimization of interactions between the collaborative entities. Hence, we propose a conceptual framework and modeling approach for collaborative learning based on imitation learning. On the other hand, the difficulty lies in explaining the contributions of humans and machines in the learning process and the lack of metrics for measuring the learned strategies and the uncertainty of decision gradients. To address these issues, we introduce an RL (Reinforcement Learning) framework for human-machine collaboration, known as Human-Machine RL. This framework employs reward shaping techniques for offline policy learning. In order to assess the policies, we design an estimation algorithm tailored for human-machine collaboration scenarios, based on reinforcement learning. Additionally, we incorporate Shapley as a mathematical interpretive tool for policy rewards. We tackle the issue of gradient variance that may arise from Shapley. The feasibility of our approach is theoretically demonstrated, and we have made the source code available for result reproducibility.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Shapley-Optimized Reinforcement Learning for Human-Machine Collaboration Policy

  • Jie Zhang,
  • Yiqun Niu,
  • Wei He,
  • Cheng Jin,
  • Chongjun Wang

摘要

Human-machine collaboration is a promising training framework aimed at learning optimal strategies in high-cost exploration scenarios. However, such work is challenging. On one hand, current research on human-machine collaboration primarily focuses on imitation learning, overlooking the optimization of interactions between the collaborative entities. Hence, we propose a conceptual framework and modeling approach for collaborative learning based on imitation learning. On the other hand, the difficulty lies in explaining the contributions of humans and machines in the learning process and the lack of metrics for measuring the learned strategies and the uncertainty of decision gradients. To address these issues, we introduce an RL (Reinforcement Learning) framework for human-machine collaboration, known as Human-Machine RL. This framework employs reward shaping techniques for offline policy learning. In order to assess the policies, we design an estimation algorithm tailored for human-machine collaboration scenarios, based on reinforcement learning. Additionally, we incorporate Shapley as a mathematical interpretive tool for policy rewards. We tackle the issue of gradient variance that may arise from Shapley. The feasibility of our approach is theoretically demonstrated, and we have made the source code available for result reproducibility.