<p>Recently, diffusion models have shown satisfactory performance in various fields, especially in image and video generations. Inspired by the diffusion model, we try to explore the application of the diffusion model in temporal action detection which aims to real-time understand human action in videos. To fill this gap, we propose a high-quality generative denoising (HQGD). Specifically, the action proposals for ground truth are first confused by adding multiple steps of random Gaussian noise (e.g., 300 steps). Then, we design a transformer-based denoiser to denoise the chaotic action proposal step by step, thereby restoring the initial state of the action proposals. In the process of denoising, video information is used as guidance for the denoiser. In order to achieve high-quality noise denoising, we propose a hybrid gated attention mechanism (HGAM), which can effectively fuse RGB and optical flow features to obtain higher-quality fusion features. HGAM includes a gated suppression module (GSM) and an attention refinement module (ARM). GSM uses dual-branch gated recurrent units and naive gated mechanisms to capture remote context and suppress background information. ARM uses fusion features as query vectors to refine RGB and optical flow features by attention mechanisms, respectively. We conduct extensive experiments to validate our HQGD on four challenging datasets, i.e., THUMOS14, ActivityNet<InlineEquation ID="IEq1"> <EquationSource Format="TEX">\(-\)</EquationSource> <EquationSource Format="MATHML"><math> <mo>-</mo> </math></EquationSource> </InlineEquation>1.3, MultiTHUMOS, and EPIC-Kitchens 100.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

HQGD: high-quality generative denoising for human action understanding in videos

  • Tianwei Qu,
  • Qixian Zhang,
  • Ya Li

摘要

Recently, diffusion models have shown satisfactory performance in various fields, especially in image and video generations. Inspired by the diffusion model, we try to explore the application of the diffusion model in temporal action detection which aims to real-time understand human action in videos. To fill this gap, we propose a high-quality generative denoising (HQGD). Specifically, the action proposals for ground truth are first confused by adding multiple steps of random Gaussian noise (e.g., 300 steps). Then, we design a transformer-based denoiser to denoise the chaotic action proposal step by step, thereby restoring the initial state of the action proposals. In the process of denoising, video information is used as guidance for the denoiser. In order to achieve high-quality noise denoising, we propose a hybrid gated attention mechanism (HGAM), which can effectively fuse RGB and optical flow features to obtain higher-quality fusion features. HGAM includes a gated suppression module (GSM) and an attention refinement module (ARM). GSM uses dual-branch gated recurrent units and naive gated mechanisms to capture remote context and suppress background information. ARM uses fusion features as query vectors to refine RGB and optical flow features by attention mechanisms, respectively. We conduct extensive experiments to validate our HQGD on four challenging datasets, i.e., THUMOS14, ActivityNet \(-\) - 1.3, MultiTHUMOS, and EPIC-Kitchens 100.