Mechanistic Interpretability (MI) is an emerging field focused on demystifying black-box models by reverse engineering their learned internal algorithms. Most current MI research on transformers operates within dictionary-based settings, meaning that a fixed token vocabulary is used. Recent efforts have begun to address this, but they often neglect the concept of circuits in MI, as circuits are token-based and non-dictionary settings lack tokens. We propose modifications to the theory, developing position-based circuits to overcome this limitation. However, these adjustments require input data to meet certain properties. A common strategy in MI is to employ toy models and toy datasets for simplifying analysis, yet there is an absence of suitable data that meets required criteria. In response, we have designed two novel toy datasets, “Dots” and “Moving Dots,” which aim to facilitate MI research in non-dictionary settings. We show that transformers can be trained to high accuracy scores on these datasets even when they are forced to compute both spatial and temporal attention. Our study shows that position-based MI analysis is feasible in this context. In particular, we provide evidence of inherent compositionality in transformer models within a non-dictionary setting using composed circuits. Our analysis primarily focuses on attention-only, encoder-only transformers with bidirectional attention. Additionally, we performed comparative studies using skeleton-based action data, a real-world dataset analogous to the Moving Dots dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Mechanistic Interpretability of Transformers in Non-dictionary Setting Using Toy Datasets

  • Vishnu S Nair,
  • Matcha Naga Gayathri,
  • Akash Sharma,
  • Jayaraj Joseph,
  • Mohanasankar Sivaprakasam

摘要

Mechanistic Interpretability (MI) is an emerging field focused on demystifying black-box models by reverse engineering their learned internal algorithms. Most current MI research on transformers operates within dictionary-based settings, meaning that a fixed token vocabulary is used. Recent efforts have begun to address this, but they often neglect the concept of circuits in MI, as circuits are token-based and non-dictionary settings lack tokens. We propose modifications to the theory, developing position-based circuits to overcome this limitation. However, these adjustments require input data to meet certain properties. A common strategy in MI is to employ toy models and toy datasets for simplifying analysis, yet there is an absence of suitable data that meets required criteria. In response, we have designed two novel toy datasets, “Dots” and “Moving Dots,” which aim to facilitate MI research in non-dictionary settings. We show that transformers can be trained to high accuracy scores on these datasets even when they are forced to compute both spatial and temporal attention. Our study shows that position-based MI analysis is feasible in this context. In particular, we provide evidence of inherent compositionality in transformer models within a non-dictionary setting using composed circuits. Our analysis primarily focuses on attention-only, encoder-only transformers with bidirectional attention. Additionally, we performed comparative studies using skeleton-based action data, a real-world dataset analogous to the Moving Dots dataset.