Transformers are widely recognized as leading models for NLP tasks due to their attention-based architecture. However, their complexity and numerous parameters hinder the understanding of their decision-making processes, restricting their use in high-risk domains where accurate explanations are crucial. To overcome this challenge, a technique named Optimus was introduced recently. Optimus provides an adaptive selection of head, layer, and matrix operations, to provide feature importance based interpretations for transformers. This work extends Optimus, adapting to two new transformer models, as well as the new task of multi-class classification, while also optimizing the time response of the technique. Experiments showed that the performance of Optimus remains consistent through different encoder-based transformer models and classification tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On the Adaptability of Attention-Based Interpretability in Different Transformer Architectures for Multi-class Classification Tasks

  • Sofia Katsaki,
  • Christos Aivazidis,
  • Nikolaos Mylonas,
  • Ioannis Mollas,
  • Grigorios Tsoumakas

摘要

Transformers are widely recognized as leading models for NLP tasks due to their attention-based architecture. However, their complexity and numerous parameters hinder the understanding of their decision-making processes, restricting their use in high-risk domains where accurate explanations are crucial. To overcome this challenge, a technique named Optimus was introduced recently. Optimus provides an adaptive selection of head, layer, and matrix operations, to provide feature importance based interpretations for transformers. This work extends Optimus, adapting to two new transformer models, as well as the new task of multi-class classification, while also optimizing the time response of the technique. Experiments showed that the performance of Optimus remains consistent through different encoder-based transformer models and classification tasks.