How Do Transformers Integrate Meanings? An Investigation Using Interpretable Brain-Based Componential Semantics in Two-Word Phrases
摘要
Transformer models have proven to be an efficient architecture for large language models, enabling them to exhibit human-like behaviors. However, it remains unclear how they combine the properties of individual words. To investigate this, we focus on the simplest compositional unit—two-word phrases—and employ four simple, parameter-free composition operations: addition, multiplication, first word, and second word. This approach serves as the first step toward understanding the composition mechanisms in transformer modules. We propose straightforward interpretation methods based on brain-based componential semantics. Specifically, we map the distributed vector space in transformers to an interpretable brain-based componential space to explore the intrinsic properties of representation and their semantic compositionality. Our findings show that, like humans, phrase types and semantic features influence the combination process in transformers. However, unlike humans, most phrase types paired with semantic features require more complex combination operations than simple addition or multiplication.