Event cameras, as advanced bionic vision sensors, have garnered widespread attention due to their exceptional pixel-level temporal resolution and high dynamic range. They generate a novel data format that encapsulates spatial and temporal information, typically represented through frame-based and point-based views. Current research often focuses on optimizing feature extraction within a single view. Although frame-based views provide a dense and easily processable format, they significantly lose temporal resolution during projection. In contrast, event-based views maintain geometric accuracy and information integrity but their inherent disorder hampers effective local feature extraction. To utilize different views’ advantages and further explore the potential of the inherently diverse representations in event data, we propose the Dual-Stream Spatio-Temporal Fusion Network (DSTF-Net) for the first time, innovatively integrating multiple representations of event data, including 2D event feature maps and 3D spatio-temporal event clouds. DSTF-Net employs a phase-based motion information enhancement strategy to minimize information loss in frame-based representations, initially identifying motion trends through temporal residual maps, then enhancing these trends using event clouds. Additionally, we develop a lightweight plug-and-play method that converts event clouds into pseudo-images for integration with feature maps, thereby facilitating efficient local feature extraction. Comprehensive experiments on various datasets demonstrate that our model achieves state-of-the-art (SOTA) performance in event camera object recognition tasks. The source code can be found at: DSTF .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

DSTF: Dual-Stream Spatio-Temporal Fusion Network for Event-Based Data

  • Xusheng Gu,
  • Changjie Qiu,
  • Xiuhong Lin,
  • Xinjie Yang,
  • Yu Zang,
  • Cheng Wang

摘要

Event cameras, as advanced bionic vision sensors, have garnered widespread attention due to their exceptional pixel-level temporal resolution and high dynamic range. They generate a novel data format that encapsulates spatial and temporal information, typically represented through frame-based and point-based views. Current research often focuses on optimizing feature extraction within a single view. Although frame-based views provide a dense and easily processable format, they significantly lose temporal resolution during projection. In contrast, event-based views maintain geometric accuracy and information integrity but their inherent disorder hampers effective local feature extraction. To utilize different views’ advantages and further explore the potential of the inherently diverse representations in event data, we propose the Dual-Stream Spatio-Temporal Fusion Network (DSTF-Net) for the first time, innovatively integrating multiple representations of event data, including 2D event feature maps and 3D spatio-temporal event clouds. DSTF-Net employs a phase-based motion information enhancement strategy to minimize information loss in frame-based representations, initially identifying motion trends through temporal residual maps, then enhancing these trends using event clouds. Additionally, we develop a lightweight plug-and-play method that converts event clouds into pseudo-images for integration with feature maps, thereby facilitating efficient local feature extraction. Comprehensive experiments on various datasets demonstrate that our model achieves state-of-the-art (SOTA) performance in event camera object recognition tasks. The source code can be found at: DSTF .