错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Attention to Phonetics: A Visually Informed Explanation of Speech Transformers

  • Erfan A. Shams,
  • Julie Carson-Berndsen

摘要

Self-supervised learning based on the transformer architecture has improved the performance of Automatic Speech Recognition (ASR) systems hugely in recent years while the interpretability of such transformer-based models, by design, has received less attention. Considering this, we investigate post-hoc explainability methodologies to explore the types of phonetic information that are encoded within the black box of transformer-based ASR models. We propose an exploratory visual environment based on the encoded parameters in the self-attention (SA) component of the models as a first step in explaining transformer-based ASRs via interactive exploration of the SA heads. The visualisations reveal clues about the functionality of specific SA heads that in turn support the choice of a suitable domain-informed post-hoc explainability method for a deeper analysis. We apply this method to identify the impact of certain SA heads in encoding sub-phonetic information in the model embeddings and demonstrate that specialised SA heads can be potentially identified via combined visualisation and post-hoc analysis.