错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Utilizing Video Word Boundaries and Feature-Based Knowledge Distillation Improving Sentence-Level Lip Reading

  • Hongzhong Zhen,
  • Chenglong Jiang,
  • Jiyong Zhou,
  • Liming Liang,
  • Ying Gao

摘要

Lip reading is to recognize the spoken content from silent video of lip movement. There is a general problem in sentence-level lip reading that the length of predicted text is inconsistent with actual text. To alleviate this problem, we introduce video word boundary information into sentence-level lip reading and propose TLiM-VWB model. Besides, to deal with the situation that video word boundaries can not be obtained in wild environment, we propose LiM-VWB-KD method with two knowledge distillation strategies utilizing video word boundary information implicitly. We evaluate our model and method on CMLR and LRS2 datasets with metrics of CER/WER and our proposed length difference rate (LDR). We verify the effectiveness of video word boundary information to improve sentence-level lip reading accuracy through the results of TLiM-VWB model. We also show the effectiveness of LiM-VWB-KD method especially with feature-based strategy. Our LiM-VWB-KD method achieves the best result on Chinese sentence-level lip reading among methods using Transformer architecture and achieves the new state-of-the-art performance in speaker-independent setting on CMLR.