End-to-End Streaming Customizable Keyword Spotting Based on Text-Adaptive Neural Search
摘要
Streaming keyword spotting (KWS) is an important technique for voice assistant wake-up. While KWS with a preset fixed keyword has been well studied, test-time customizable keyword spotting in streaming mode remains a great challenge due to the lack of pre-collected keyword-specific training data and the requirement of streaming detection output. In this paper, we propose a novel end-to-end text-adaptive neural search architecture with a multi-label trigger mechanism to allow any pre-trained ASR acoustic model to be effectively used for fast streaming customizable keyword spotting. Evaluation results on various datasets show that our approach significantly outperforms both traditional post-processing baseline and the neural search baseline, meanwhile achieving a 44x search speedup compared to the traditional post-processing method.