Data Poisoning Attack Against Reinforcement Learning from Human Feedback in Robot Control Tasks
摘要
Reinforcement Learning from Human Feedback (RLHF) technology provides a method for agents to learn human preferences and perform actions that satisfy human desires. This technology was originally used to complete robot control tasks in situations where it was difficult to design a reward function. However, in the process of collecting human feedback, malicious human annotators may launch attacks on RLHF, posing a significant challenge to the technology. Previous research has mostly focused on studying the harm caused by different attacks on RLHF in fine-tuning large language models (LLMs). However, our study is focused on the field of robot control. We designed two data poisoning attacks against human feedback datasets and implemented our attacks in three different offline reinforcement learning (RL) environments. The experimental results show that RLHF is vulnerable to data poisoning attacks in robot control tasks.