Small footprint automatic speech recognition (ASR)-free keyword spotting using multivariate time series analysis
摘要
In the domain of speech analysis, keyword spotting (KWS) is referred to as one of the most important challenges, where the task is to determine the presence of a particular word in the audio data. With its increasing applications in a variety of fields, various methods have been proposed for KWS in recent times, predominantly using deep learning-based techniques. However, despite their appreciable performance, the recent state-of-the-art deep learning-based KWS methods are often computationally demanding, involve a large number of parameters, and hence are not appropriate for deployment on devices with limited resources like edge devices. To address this challenge, this paper proposed a novel small-footprint KWS approach that uniquely formulated KWS as a time series classification problem wherein the audio data was modeled as a univariate or multivariate time series and the target keywords formed the classes. The time series classification problem was then solved using efficient unsupervised time series feature extractors coupled with simple machine learning classifiers, thus eliminating the need for huge deep learning architectures. By exploiting the unique unsupervised temporal features in audio data, along with the acoustic features, the proposed approach provided a lightweight, state-of-the-art solution for KWS with the accuracy reaching to