<p>In driving scenarios, videos recorded in rainy weather conditions are often distorted by rain streaks and raindrops, posing a significant challenge in recovering the obscured background details. The inherent temporal redundancy in videos offers stability advantages for rain removal. Traditional video deraining techniques primarily depend on optical flow estimation and kernel-based methods, which are constrained by a limited receptive field. Although transformer architectures can capture long-term dependencies, they introduce substantial computational complexity. Recently, the Receptance Weighted Key Value Model (RWKV), characterized by its linear computational complexity, has emerged as an effective tool for efficient long-term temporal modeling, which is essential for the removal of rain streaks and raindrops in video sequences. To optimize RWKV for video deraining, we introduce a wavelet transform shift mechanism that enhances low-frequency features by targeting distinct frequency bands. Additionally, we present a tubelet embedding mechanism for RWKVs, augmenting the model’s capacity to capture high-frequency details by integrating the spatiotemporal context of input frames. Extensive experiments demonstrate that our approach achieves superior performance over state-of-the-art methods.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

RainRWKV: a deep RWKV model for video deraining

  • Xijun Wang,
  • Xin Zhou,
  • Yi Wang,
  • Songto Zeng,
  • Xinyu Liu,
  • Haobo Shen,
  • Xianying Wang,
  • Ping Li,
  • Lei Zhu

摘要

In driving scenarios, videos recorded in rainy weather conditions are often distorted by rain streaks and raindrops, posing a significant challenge in recovering the obscured background details. The inherent temporal redundancy in videos offers stability advantages for rain removal. Traditional video deraining techniques primarily depend on optical flow estimation and kernel-based methods, which are constrained by a limited receptive field. Although transformer architectures can capture long-term dependencies, they introduce substantial computational complexity. Recently, the Receptance Weighted Key Value Model (RWKV), characterized by its linear computational complexity, has emerged as an effective tool for efficient long-term temporal modeling, which is essential for the removal of rain streaks and raindrops in video sequences. To optimize RWKV for video deraining, we introduce a wavelet transform shift mechanism that enhances low-frequency features by targeting distinct frequency bands. Additionally, we present a tubelet embedding mechanism for RWKVs, augmenting the model’s capacity to capture high-frequency details by integrating the spatiotemporal context of input frames. Extensive experiments demonstrate that our approach achieves superior performance over state-of-the-art methods.