In our prior work, we proposed a Global-Temporal Enhancement (GTE) based continuous sign language recognition algorithm. However, we observed that GTE, serving as the intermediary network layer, lacks adequate feature representational capability, and its arbitrarily teaching of shallower module can hinder shallower feature learning. Consequently, this paper proposes an additional distillation loss between the two teacher modules, i.e., GTE and BiLSTM, which will mitigate the interference of erroneous information in GTE module, form a progressive self distillation structure, and select appropriate distillation temperatures for each distillation module. Furthermore, considering the network’s three distinct self-distillation losses, we apply separate and different temperatures to each loss based on their respective network-level characteristics, aligning with the potential feature learning capability across network levels. Additionally, in the network’s output classifier, we adopt a weight normalization linear layer in lieu of a fully connected layer for prediction classification, enhancing data distribution and training effectiveness. This Multi-Temperature Progressive Self-Distillation (MTPSD) based continuous sign language recognition algorithm improves upon our prior GTE algorithm, reducing the WER to 19.4%/20.2% and 18.2%/20.1% on the development and test sets of PHOENIX14 and PHOENIX14-T datasets respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multi-temperature Progressive Self-distillation Algorithm for Continuous Sign Language Recognition

  • Xiaofei Qin,
  • Anluo Yi,
  • Hui Wang,
  • Chengcheng Xu,
  • Huixia Zhao

摘要

In our prior work, we proposed a Global-Temporal Enhancement (GTE) based continuous sign language recognition algorithm. However, we observed that GTE, serving as the intermediary network layer, lacks adequate feature representational capability, and its arbitrarily teaching of shallower module can hinder shallower feature learning. Consequently, this paper proposes an additional distillation loss between the two teacher modules, i.e., GTE and BiLSTM, which will mitigate the interference of erroneous information in GTE module, form a progressive self distillation structure, and select appropriate distillation temperatures for each distillation module. Furthermore, considering the network’s three distinct self-distillation losses, we apply separate and different temperatures to each loss based on their respective network-level characteristics, aligning with the potential feature learning capability across network levels. Additionally, in the network’s output classifier, we adopt a weight normalization linear layer in lieu of a fully connected layer for prediction classification, enhancing data distribution and training effectiveness. This Multi-Temperature Progressive Self-Distillation (MTPSD) based continuous sign language recognition algorithm improves upon our prior GTE algorithm, reducing the WER to 19.4%/20.2% and 18.2%/20.1% on the development and test sets of PHOENIX14 and PHOENIX14-T datasets respectively.