Transformer-based language models have achieved outstanding results on a wide range of tasks, but often at the expense of high computational demands. In response, we introduce Seek2Skim, a novel token pruning method seamlessly integrated into the sequence-to-sequence architecture. Unlike previous efforts that dynamically adjust computations in either the encoder or decoder, Seek2Skim further improves computational efficiency by simultaneously enhancing both components. Our method incorporates trainable modules that assess token importance, pruning less crucial tokens to effectively reduce computational overhead. Additionally, we introduce dynamic cross-attention filtering to preserve model performance even with reduced token representation updates. Extensive experiments confirm the effectiveness of Seek2Skim, demonstrating substantial speed improvements across diverse tasks. Notably, Seek2Skim preserves over 99% of the backbone model’s performance with a 1.91 \(\times \) and 1.53 \(\times \) inference speedup for text classification and generation tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Dynamic Layer-Wise Token Pruning for Sequence-to-Sequence Transformer Inference

  • Ji Hun Keom,
  • Yeachan Kim,
  • Sung Ju Lee,
  • Sang Keun Lee

摘要

Transformer-based language models have achieved outstanding results on a wide range of tasks, but often at the expense of high computational demands. In response, we introduce Seek2Skim, a novel token pruning method seamlessly integrated into the sequence-to-sequence architecture. Unlike previous efforts that dynamically adjust computations in either the encoder or decoder, Seek2Skim further improves computational efficiency by simultaneously enhancing both components. Our method incorporates trainable modules that assess token importance, pruning less crucial tokens to effectively reduce computational overhead. Additionally, we introduce dynamic cross-attention filtering to preserve model performance even with reduced token representation updates. Extensive experiments confirm the effectiveness of Seek2Skim, demonstrating substantial speed improvements across diverse tasks. Notably, Seek2Skim preserves over 99% of the backbone model’s performance with a 1.91 \(\times \) and 1.53 \(\times \) inference speedup for text classification and generation tasks.