<p>The growth in Internet audio data has highlighted the need for accurate and efficient search methodologies. In this context, query-by-example spoken term detection (QbE-STD) plays a pivotal role, mainly when dealing with multilingual datasets lacking proper transcription information. Recent advancements in the field include integrating convolutional neural networks (CNN) as classifiers, emphasizing the computation of matching matrices for accurate term detection. This research conducts a comparative analysis of similarity measures for QbE-STD, aiming to highlight the performance differences across different metrics. The results reveal distinctive patterns in accuracy, with kernel-based measures leading the pack. Distance-based measures demonstrate comparable accuracies, albeit slightly lower than other metrics, while DTW yields the lowest accuracy. The implications of these findings extend to optimizing QbE-STD systems, guiding practitioners in selecting suitable similarity measures based on their specific requirements. The study contributes valuable insights to spoken term detection, paving the way for enhanced methodologies in audio search applications.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring the Impact of Different Similarity Measures on Query-by-Example Spoken Term Detection

  • Manisha Naik Gaonkar,
  • Veena Thenkanidiyoor,
  • A. D. Dileep

摘要

The growth in Internet audio data has highlighted the need for accurate and efficient search methodologies. In this context, query-by-example spoken term detection (QbE-STD) plays a pivotal role, mainly when dealing with multilingual datasets lacking proper transcription information. Recent advancements in the field include integrating convolutional neural networks (CNN) as classifiers, emphasizing the computation of matching matrices for accurate term detection. This research conducts a comparative analysis of similarity measures for QbE-STD, aiming to highlight the performance differences across different metrics. The results reveal distinctive patterns in accuracy, with kernel-based measures leading the pack. Distance-based measures demonstrate comparable accuracies, albeit slightly lower than other metrics, while DTW yields the lowest accuracy. The implications of these findings extend to optimizing QbE-STD systems, guiding practitioners in selecting suitable similarity measures based on their specific requirements. The study contributes valuable insights to spoken term detection, paving the way for enhanced methodologies in audio search applications.