A Multi-layered Approach to Evaluating Speech Translation Performance of Meetings
摘要
Evaluating only the final output of cascaded speech translation systems offers a limited understanding of the individual performance of each component in the cascade. This limitation burdens the identification and improvement of problematic components. To address this issue, we present a multi-layer evaluation suite for automatic speech translation of meetings. Our data features public-domain English, Latvian, and Lithuanian recordings, augmented with multiple annotation layers ranging from raw speech transcription to translation. We further present how to use our data sets and annotations to evaluate components involved in cascaded speech translation systems: speaker diarisation, speech segmentation, automatic speech recognition, punctuation restoration and sentence splitting, speech normalisation, and machine translation. We also demonstrate an ablation study that allows us to analyse each component’s error contribution to the overall speech translation error. We publish our data and evaluation scripts, making our evaluation suite the first of its kind for Latvian and Lithuanian languages.