S-TsNet: a continuous sign language recognition network via Spatial-Temporal stage Network
摘要
In continuous sign language recognition, traditional spatial feature extraction often relies on frame-by-frame static processing, which struggles to effectively capture dynamic changes between frames and limits the mining of key semantic information. Temporal modeling commonly employs fixed-scale convolutions, making it difficult to adapt to the variability of action boundaries, leading to blurred semantic segmentation and weakened temporal feature representation. To address these issues, this paper proposes a continuous sign language recognition network, Spatial-Temporal stage Network (S-TsNet). S-TsNet consists of two main stages: the spatial feature detection stage and the temporal feature detection stage. In the spatial feature extraction stage, a Mixed Spatiotemporal Block is designed, combining residual modules with an Element-level Multi-Scale Temporal Module (EMT) to model pixel-level inter-frame correlations and effectively capture dynamic information in key regions such as the hands. In the temporal modeling stage, a Multi-Scale Temporal Perception Module (MTP) is introduced, which employs hierarchical modeling and cross-scale fusion to adaptively model action boundaries. Experiments on the PHOENIX14 and PHOENIX14-T datasets demonstrate that S-TsNet significantly outperforms mainstream methods in recognition accuracy. Ablation studies and visualization further validate its effectiveness.