GlintNet: A Lightweight Global-Local Integration Network with Spatial-Channel Mixed Attention for ReID
摘要
Existing object re-identification models struggle with parameter inefficiency and computational overhead under complex scenarios involving viewpoint variations or occlusions. To address this, we propose a lightweight Global-Local Feature Integration Network (GlintNet) which integrating CNN-extracted local features and attention-based global representations through a simplified fusion framework. Its bottleneck architecture ensures parameter efficiency, while the novel Spatial-Channel Mixed Attention (SCMA) mechanism—comprising Spatial Embedding Attention (SEA) and Channel Mixed Attention (CMA)—enhances discriminative feature learning. SEA captures spatial patterns using depthwise separable convolutions, while CMA refines channel-wise dependencies, both maintaining linear complexity via grouped computation. Evaluations across multiple benchmarks demonstrate GlintNet’s superiority achieving state-of-the-art accuracy, particularly excelling in occlusion and cross-view scenarios.