<p>Temporal action detection (TAD) aims to locate action positions and recognize action categories in long-term untrimmed videos. Despite the fact that numerous methods have achieved encouraging outcomes, their robustness has not been comprehensively investigated. In practice, it is observed that temporal information in videos can sometimes be compromised, such as by missing or blurred frames. Notably, existing methods are highly vulnerable to these scenarios, often experiencing a significant decline in performance even when only a single frame is corrupted. In this paper, we take the first step towards benchmarking the temporal robustness of TAD models and aim to identify the factors influencing temporal robustness. This, in turn, provides insights for designing more robust TAD models. To formally assess robustness, we establish three temporal corruption robustness benchmarks, namely THUMOS14-C, ActivityNet-1.3-C and MultiTHUMOS-C, which consider eight common types of corruption encountered during video recording and transmission. Each type of corruption is applied at three levels of severity, resulting in a total of 24 distinct corruptions, which comprehensively cover different durations of corruption, serving as a controllable representative of practical real-world scenarios. On these benchmarks, we conduct an extensive analysis of the robustness of 12 leading TAD methods and uncover several noteworthy findings: 1) Existing methods are particularly vulnerable to temporal corruptions, with end-to-end methods likely being more susceptible than those employing a pre-trained feature extractor on THUMOS14-C; 2) The primary source of vulnerability is localization error rather than classification error; and 3) TAD models tend to exhibit the most significant performance degradation when corruptions occur in the middle of an action instance. Furthermore, we investigate the impact of diverse TAD model designs on temporal robustness by evaluating eight key factors across three crucial stages: feature representation, architecture design, and training strategy. Based on the insights gained from these explorations, we propose a recipe with six steps to build a strong TAD baseline. Experiments conducted on three benchmark datasets demonstrate that our recipe not only improves robustness against corruption but also results in enhancements on clean data. Specifically, on the THUMOS14-C dataset, we achieve an 11.56% improvement in relative robustness and a 1.32% increase in clean mAP. We believe that this study will play a crucial role in shaping future research on robust video analysis. The benchmark dataset is available at <a href="https://github.com/Alvin-Zeng/temporal-robustness-benchmark">https://github.com/Alvin-Zeng/temporal-robustness-benchmark</a>.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards Robust Temporal Action Detection: Benchmark and A Strong Baseline

  • Runhao Zeng,
  • Jiaming Liang,
  • Jiaqi Mao,
  • Xiaoyong Chen,
  • Wei Wang,
  • Yong Guo,
  • Limin Wang,
  • Victor C. M. Leung,
  • Xiping Hu

摘要

Temporal action detection (TAD) aims to locate action positions and recognize action categories in long-term untrimmed videos. Despite the fact that numerous methods have achieved encouraging outcomes, their robustness has not been comprehensively investigated. In practice, it is observed that temporal information in videos can sometimes be compromised, such as by missing or blurred frames. Notably, existing methods are highly vulnerable to these scenarios, often experiencing a significant decline in performance even when only a single frame is corrupted. In this paper, we take the first step towards benchmarking the temporal robustness of TAD models and aim to identify the factors influencing temporal robustness. This, in turn, provides insights for designing more robust TAD models. To formally assess robustness, we establish three temporal corruption robustness benchmarks, namely THUMOS14-C, ActivityNet-1.3-C and MultiTHUMOS-C, which consider eight common types of corruption encountered during video recording and transmission. Each type of corruption is applied at three levels of severity, resulting in a total of 24 distinct corruptions, which comprehensively cover different durations of corruption, serving as a controllable representative of practical real-world scenarios. On these benchmarks, we conduct an extensive analysis of the robustness of 12 leading TAD methods and uncover several noteworthy findings: 1) Existing methods are particularly vulnerable to temporal corruptions, with end-to-end methods likely being more susceptible than those employing a pre-trained feature extractor on THUMOS14-C; 2) The primary source of vulnerability is localization error rather than classification error; and 3) TAD models tend to exhibit the most significant performance degradation when corruptions occur in the middle of an action instance. Furthermore, we investigate the impact of diverse TAD model designs on temporal robustness by evaluating eight key factors across three crucial stages: feature representation, architecture design, and training strategy. Based on the insights gained from these explorations, we propose a recipe with six steps to build a strong TAD baseline. Experiments conducted on three benchmark datasets demonstrate that our recipe not only improves robustness against corruption but also results in enhancements on clean data. Specifically, on the THUMOS14-C dataset, we achieve an 11.56% improvement in relative robustness and a 1.32% increase in clean mAP. We believe that this study will play a crucial role in shaping future research on robust video analysis. The benchmark dataset is available at https://github.com/Alvin-Zeng/temporal-robustness-benchmark.