VeMamba: Voxel-Based Multi-Scale State Space Model Network for Event Stream Recognition
摘要
Event cameras are innovative neuromorphic sensors that capture dynamic changes in scenes, recording millions of events per second. Recent sparse computational models on event stream recognition have achieved notable successes by utilizing graph convolution or attention mechanisms to model local dependencies. However, these methods face significant challenges in constructing effective event representations and modeling global dependencies when confronted with millions of events in spacetime. To surmount the above challenges, we present a novel voxel-based State Space Model (SSM) network, termed VeMamba, which can establish multi-scale spatiotemporal dependencies in event voxels with linear complexity. Specifically, we design a time-aware enhanced voxelization method that enriches the spatiotemporal expression within event voxels while preserving sparse computation. Then, we propose a multi-scale modeling module with linear complexity that integrates local attention into the global dual-scale SSMs to establish spatiotemporal dependencies from local to global within serialized voxels. Furthermore, leveraging a hierarchical structure grounded in voxel merging, we can extract deep semantic and motion cues from the voxels. Extensive experiments demonstrate that VeMamba achieves state-of-the-art (SOTA) performance with low model complexity and computational cost on event stream recognition tasks.