Monitoring the cooking process with autonomous robots requires recognition of food states, such as mixing and doneness levels. A previous approach applied CLIP [1] to cooking videos but needed to exclude frames where food regions were occluded by cooking utensils to maintain recognition accuracy. However, in cooking processes that involve frequent mixing, the number of usable frames for state recognition decreases (Challenge 1). Additionally, in the training process, images were divided into two categories—before and after completion—based on a manually specified completion timing. However, the pre-completion category contained images ranging from the start of cooking to just before completion, leading to high intra-category variance and training inconsistencies, which degraded recognition accuracy (Challenge 2). To address these challenges, our proposed method reduces the learning contribution of occluded regions, allowing all frames to be used for processing (Challenge 1). Furthermore, instead of using only two categories (completed and incomplete), we also incorporate 3 to 6 intermediate categories into the training process to capture the gradual progression of cooking, mitigating training inconsistencies and enhancing recognition accuracy (Challenge 2). Experimental results showed that the model with reduced occlusion contribution achieved 30.4% recognition accuracy, a 7.4% improvement over the model trained by removing occluded frames. Additionally, the model trained with intermediate categories achieved 31.5% accuracy, a 10.8% improvement over the conventional model trained with only two categories.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Food State Recognition from Recipes Using Multimodal Model for Task Monitoring in Autonomous Cooking Robots

  • Rina Tagami,
  • Hiroki Kobayashi,
  • Shuichi Akizuki,
  • Manabu Hashimoto

摘要

Monitoring the cooking process with autonomous robots requires recognition of food states, such as mixing and doneness levels. A previous approach applied CLIP [1] to cooking videos but needed to exclude frames where food regions were occluded by cooking utensils to maintain recognition accuracy. However, in cooking processes that involve frequent mixing, the number of usable frames for state recognition decreases (Challenge 1). Additionally, in the training process, images were divided into two categories—before and after completion—based on a manually specified completion timing. However, the pre-completion category contained images ranging from the start of cooking to just before completion, leading to high intra-category variance and training inconsistencies, which degraded recognition accuracy (Challenge 2). To address these challenges, our proposed method reduces the learning contribution of occluded regions, allowing all frames to be used for processing (Challenge 1). Furthermore, instead of using only two categories (completed and incomplete), we also incorporate 3 to 6 intermediate categories into the training process to capture the gradual progression of cooking, mitigating training inconsistencies and enhancing recognition accuracy (Challenge 2). Experimental results showed that the model with reduced occlusion contribution achieved 30.4% recognition accuracy, a 7.4% improvement over the model trained by removing occluded frames. Additionally, the model trained with intermediate categories achieved 31.5% accuracy, a 10.8% improvement over the conventional model trained with only two categories.