Empowered by the development of multimodal information processing and large language model-related technologies, robots that can operate autonomously in human living spaces are being realized. When such robots communicate with human users, they need a framework to appropriately select what to say to the user among the recognition results, memories, and decisions obtained from various multimodal environments. In this study, we constructed a framework that enables robots to appropriately select events to be verbalized by calculating the mutual information between the robot’s verbalization texts and its own memory or common sense knowledge. User evaluation results using crowdsourcing suggested that the proposed framework improves the necessity and sufficiency of the robot’s speech. This ability will contribute to improving the usability of the robot.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

What Should Autonomous Robots Verbalize and What Should They Not?

  • Daichi Yoshihara,
  • Akishige Yuguchi,
  • Seiya Kawano,
  • Takamasa Iio,
  • Koichiro Yoshino

摘要

Empowered by the development of multimodal information processing and large language model-related technologies, robots that can operate autonomously in human living spaces are being realized. When such robots communicate with human users, they need a framework to appropriately select what to say to the user among the recognition results, memories, and decisions obtained from various multimodal environments. In this study, we constructed a framework that enables robots to appropriately select events to be verbalized by calculating the mutual information between the robot’s verbalization texts and its own memory or common sense knowledge. User evaluation results using crowdsourcing suggested that the proposed framework improves the necessity and sufficiency of the robot’s speech. This ability will contribute to improving the usability of the robot.