The advent of ChatGPT has brought significant attention to research on reinforcement learning using text data as input. A major challenge in this field is achieving advanced decision-making through the learning of commonsense knowledge, which aids in understanding both the context and meaning of text data. This paper proposes a novel method for learning commonsense knowledge by applying Deep Q-Network to an existing learning environment called ScriptWorld and designing rewards to evaluate learning for significant decision-making in parts of the scenario. Specifically, we input a sequence concatenated of the state and all choices into the Deep Q-Network to output an appropriate choice and set a reward midway through the scenario graph as a subgoal, considering the graph structure. In experiments, we compared the performance of our method with Q-learning across multiple scenarios in ScriptWorld. The experimental results demonstrated that the proposed method outperformed Q-learning by learning a policy that includes commonsense knowledge to distinguish the semantics of events. Furthermore, in the analysis of the reward design effects, setting rewards for nodes with fewer detours improved performance.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Natural Language-Based Reinforcement Learning for Acquiring Commonsense Knowledge from Partial Scenarios

  • Fumito Uwano,
  • Ryota Kubo,
  • Manabu Ohta

摘要

The advent of ChatGPT has brought significant attention to research on reinforcement learning using text data as input. A major challenge in this field is achieving advanced decision-making through the learning of commonsense knowledge, which aids in understanding both the context and meaning of text data. This paper proposes a novel method for learning commonsense knowledge by applying Deep Q-Network to an existing learning environment called ScriptWorld and designing rewards to evaluate learning for significant decision-making in parts of the scenario. Specifically, we input a sequence concatenated of the state and all choices into the Deep Q-Network to output an appropriate choice and set a reward midway through the scenario graph as a subgoal, considering the graph structure. In experiments, we compared the performance of our method with Q-learning across multiple scenarios in ScriptWorld. The experimental results demonstrated that the proposed method outperformed Q-learning by learning a policy that includes commonsense knowledge to distinguish the semantics of events. Furthermore, in the analysis of the reward design effects, setting rewards for nodes with fewer detours improved performance.