We present an informed approach to augment existing contradiction detection datasets with prototypical examples for language model training. The samples are created by combining linguistic knowledge with the generative capabilities of current large language models. Specifically, we investigate three approaches that employ rule-based augmentation, data generation using GPT models and few-shot-prompting, as well as a combination of both. We find that adding prototypical samples to the training helps to significantly reduce the training set size, while maintaining or even improving performance on the downstream task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Improving Language Model Performance by Training on Prototypical Contradictions

  • Maren Pielka,
  • Marie-Christin Freischlad,
  • Svetlana Schmidt,
  • Rafet Sifa

摘要

We present an informed approach to augment existing contradiction detection datasets with prototypical examples for language model training. The samples are created by combining linguistic knowledge with the generative capabilities of current large language models. Specifically, we investigate three approaches that employ rule-based augmentation, data generation using GPT models and few-shot-prompting, as well as a combination of both. We find that adding prototypical samples to the training helps to significantly reduce the training set size, while maintaining or even improving performance on the downstream task.