Human activities are inherently task-oriented, and integrating explicit task learning into action segmentation models is hypothesized to enhance performance. However, empirical evaluations using the Ego4D Goal Step dataset reveal a paradox: the inclusion of learning tasks deteriorated model performance. This issue partially arises from limited task samples, i.e., over 50% of tasks have less than two training samples, leading to bias and overfitting in the training process. To address this, we propose a novel grammar induction method to accurately capture the hierarchical decomposition of a task with limited task samples, and use the induced grammar to guide neural predictions. Experiments demonstrate that our induction method achieves comparable results with SOTA on the Breakfast dataset with as few as two training samples for each task. Additionally, incorporating our grammar significantly boosts temporal action segmentation results for Ego4D Goal Step dataset by 8%. This approach not only mitigates data scarcity but also enhances the robustness and accuracy of action segmentation and action detection models.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Temporal Action Segmentation and Detection with Hierarchical Task Grammar

  • Qiu Yihui,
  • Deepu Rajan

摘要

Human activities are inherently task-oriented, and integrating explicit task learning into action segmentation models is hypothesized to enhance performance. However, empirical evaluations using the Ego4D Goal Step dataset reveal a paradox: the inclusion of learning tasks deteriorated model performance. This issue partially arises from limited task samples, i.e., over 50% of tasks have less than two training samples, leading to bias and overfitting in the training process. To address this, we propose a novel grammar induction method to accurately capture the hierarchical decomposition of a task with limited task samples, and use the induced grammar to guide neural predictions. Experiments demonstrate that our induction method achieves comparable results with SOTA on the Breakfast dataset with as few as two training samples for each task. Additionally, incorporating our grammar significantly boosts temporal action segmentation results for Ego4D Goal Step dataset by 8%. This approach not only mitigates data scarcity but also enhances the robustness and accuracy of action segmentation and action detection models.