Context-Sensitive Adapter: Contextual Biasing for Personalized End-to-End Speech Recognition with Attention Fusion and Bias Filtering
摘要
Despite improvements in the generalization performance of Automatic Speech Recognition models, accurately recognizing infrequent words remains a challenging task, for example, language assistants in smart homes. A straightforward and viable approach to enhance the recognition accuracy of such rare vocabularies is to incorporate contextual information into the model. Consequently, the area of contextual biasing has increasingly garnered the attention of researchers. In this work, we introduce the Context-Sensitive Adapter, which leverages an attention mechanism to extract pertinent information from the hidden vectors of acoustic and contextual data. For the first time in the field of context bias, we introduce Hyperconformer, exploring its potential for novel applications. We propose a dual-thread architecture to train our model that ensures the accuracy of general speech recognition while also bolstering the recognition of context-specific words. Experimental results demonstrate that our method, employing the Hyperconformer-based Context-Sensitive Adapter, outperforms both non-contextual models and shallow fusion models. Compared to the baseline, our method achieved a maximum relative error rate reduction of 5.9% and 2.98%. Notably, against the current state-of-the-art (SOTA) models, our method achieved a performance increase of up to 41.72%.