错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Is Transformer-Based Attention Agnostic of the Pretraining Language and Task?

  • R. H. J. Martin,
  • R. Visser,
  • M. Dunaiski

摘要

Since the introduction of the Transformer by Vaswani et al. in 2017, the attention mechanism has been used in multiple state-of-the-art large language models (LLMs), such as BERT, ELECTRA, and various GPT versions. Due to the complexity and the large size of LLMs and deep neural networks in general, intelligible explanations for specific model outputs can be difficult to formulate. However, mechanistic interpretability research aims to make this problem more tractable. In this paper, we show that models with different training objectives—namely, masked language modelling and replaced token detection—have similar internal patterns of attention, even when trained for different languages, in our case English, Afrikaans, Xhosa, and Zulu. This result suggests that, on a high level, the learnt role of attention is language-agnostic.