Practical Considerations for Selecting Metrics to Evaluate Automatic Speech Recognition Systems
摘要
Automatic speech recognition (ASR) systems have become integral to various applications, from virtual assistants and real-time transcription to accessibility technologies and customer support automation. Evaluating ASR systems effectively is essential for ensuring their performance across diverse use cases, each with unique requirements such as transcription accuracy, semantic fidelity, operational efficiency, and reliability. This paper provides a review and analysis of evaluation metrics for ASR systems, categorizing them into edit-distance-based metrics, classification task metrics, semantic metrics, and performance-based metrics. Each category is examined for its strengths, limitations, and applicability to specific tasks, with an emphasis on practical trade-offs such as accuracy versus computational complexity and latency versus semantic understanding. Recommendations are provided to guide the selection of metrics tailored to specific ASR applications, from verbatim transcription in high-stakes domains to conversational AI and real-time systems. By highlighting the nuances of ASR evaluation, this work aims to bridge the gap between theoretical advancements in metric design and the practical demands of real-world ASR applications.