The integration of Automatic Speech Recognition (ASR) technologies into judicial systems offers a promising solution to the longstanding challenges of manual court transcription—namely, high costs, delays, and human error. This paper presents a comparative analysis of five cutting-edge ASR models, Whisper Large V3, Whisper Large V3 Turbo, NVIDIA Parakeet, NVIDIA Canary 1B Flash, and Deepgram, evaluated on the “india-supreme-court-audio” dataset. Using a comprehensive set of metrics including Word Error Rate (WER), Character Error Rate (CER), Match Error Rate (MER), Word Information Lost (WIL), and Word Information Preserved (WIP), we assess each model’s suitability for the legal domain. Our results show that NVIDIA Parakeet outperforms other models in accuracy and semantic fidelity, with Whisper Turbo close behind. Deepgram, while offering flexible deployment options, recorded the highest error rates, potentially limiting its usefulness in high-stakes legal contexts. The study also emphasizes the varying impact of substitution, deletion, and insertion errors, particularly how misrecognition can alter legal meaning. The findings support the use of advanced ASR systems for courtroom transcription, provided that concerns around data privacy, licensing, and model customization are addressed. This research serves as a foundation for future implementation and refinement of ASR technologies in judicial applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comparative Analysis of State-of-the-Art Speech-to-Text Models for Court Applications

  • Ali Alameer,
  • Hamid Kouhpeimay Jahromi,
  • Zeshan Afzal,
  • Manzar Malik,
  • Sean Morphy,
  • Taha Mansouri

摘要

The integration of Automatic Speech Recognition (ASR) technologies into judicial systems offers a promising solution to the longstanding challenges of manual court transcription—namely, high costs, delays, and human error. This paper presents a comparative analysis of five cutting-edge ASR models, Whisper Large V3, Whisper Large V3 Turbo, NVIDIA Parakeet, NVIDIA Canary 1B Flash, and Deepgram, evaluated on the “india-supreme-court-audio” dataset. Using a comprehensive set of metrics including Word Error Rate (WER), Character Error Rate (CER), Match Error Rate (MER), Word Information Lost (WIL), and Word Information Preserved (WIP), we assess each model’s suitability for the legal domain. Our results show that NVIDIA Parakeet outperforms other models in accuracy and semantic fidelity, with Whisper Turbo close behind. Deepgram, while offering flexible deployment options, recorded the highest error rates, potentially limiting its usefulness in high-stakes legal contexts. The study also emphasizes the varying impact of substitution, deletion, and insertion errors, particularly how misrecognition can alter legal meaning. The findings support the use of advanced ASR systems for courtroom transcription, provided that concerns around data privacy, licensing, and model customization are addressed. This research serves as a foundation for future implementation and refinement of ASR technologies in judicial applications.