Crosstalk in audio recordings, where multiple speakers’ voices overlap, presents a significant challenge for accurate automatic speech recognition and dialogue transcription in real-time communication systems. This paper introduces an audio processing component that achieves effective crosstalk suppression through the application of sidechain compression and noise gating techniques. Implemented within the AI-powered communication assistance platform CoSy (Communication Support System), the system enhances transcription quality while maintaining low computational resource consumption, enabling real-time operation on edge devices. Evaluation using the custom LibriDialogue dataset demonstrates that the CoSy system achieves superior computational efficiency compared to state-of-the-art models Conv-TasNet and MossFormer2, while maintaining competitive transcription accuracy and reasonable signal quality. The results indicate that the proposed approach effectively mitigates crosstalk and is particularly suitable for practical deployment in scenarios requiring real-time operation on CPU-limited edge devices.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Real-Time Transcription Quality with Sidechain Compression and Gating Techniques

  • Finn Stoldt,
  • Jakob Behnke,
  • Mathias Eulers,
  • Andreas Schrader

摘要

Crosstalk in audio recordings, where multiple speakers’ voices overlap, presents a significant challenge for accurate automatic speech recognition and dialogue transcription in real-time communication systems. This paper introduces an audio processing component that achieves effective crosstalk suppression through the application of sidechain compression and noise gating techniques. Implemented within the AI-powered communication assistance platform CoSy (Communication Support System), the system enhances transcription quality while maintaining low computational resource consumption, enabling real-time operation on edge devices. Evaluation using the custom LibriDialogue dataset demonstrates that the CoSy system achieves superior computational efficiency compared to state-of-the-art models Conv-TasNet and MossFormer2, while maintaining competitive transcription accuracy and reasonable signal quality. The results indicate that the proposed approach effectively mitigates crosstalk and is particularly suitable for practical deployment in scenarios requiring real-time operation on CPU-limited edge devices.