This study investigates the improvement of intent detection tasks, focusing on the Facebook Multilingual Task Oriented Dataset, which encompasses 12 classes, including specific intents like weather and alarms. To address the scarcity of datasets in the Amharic language for intent detection, we first translated the English dataset into Amharic using Google Translator. Additionally, we expanded the dataset to include Lithuanian, German, French, and Czech languages. Next, we balance the dataset through data augmentation using ChatGPT, particularly targeting the four least-covered classes by adding 50 sentences to each. For text vectorization, the language-agnostic BERT sentence embedding model was used for its multilingual capabilities and Amharic language support. The classification was based on Cosine similarity, incorporating a majority voting mechanism from the most similar training instances. We conducted multilingual experiments to evaluate the performance of intent detection across various languages, both before and after data augmentation. Our findings highlight the effectiveness of using ChatGPT for data augmentation in intent detection tasks across all investigated languages, particularly in enhancing underrepresented languages and specific intents. We achieved improvement in intent detection performance for minority classes. This shows the benefit of using ChatGPT with imbalanced datasets for multilingual tasks and applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Intent Detection Through ChatGPT-Driven Data Augmentation

  • Senait Gebremichael Tesfagergish,
  • Robertas Damasevicius,
  • Jurgita Kapociute-Dzikiene

摘要

This study investigates the improvement of intent detection tasks, focusing on the Facebook Multilingual Task Oriented Dataset, which encompasses 12 classes, including specific intents like weather and alarms. To address the scarcity of datasets in the Amharic language for intent detection, we first translated the English dataset into Amharic using Google Translator. Additionally, we expanded the dataset to include Lithuanian, German, French, and Czech languages. Next, we balance the dataset through data augmentation using ChatGPT, particularly targeting the four least-covered classes by adding 50 sentences to each. For text vectorization, the language-agnostic BERT sentence embedding model was used for its multilingual capabilities and Amharic language support. The classification was based on Cosine similarity, incorporating a majority voting mechanism from the most similar training instances. We conducted multilingual experiments to evaluate the performance of intent detection across various languages, both before and after data augmentation. Our findings highlight the effectiveness of using ChatGPT for data augmentation in intent detection tasks across all investigated languages, particularly in enhancing underrepresented languages and specific intents. We achieved improvement in intent detection performance for minority classes. This shows the benefit of using ChatGPT with imbalanced datasets for multilingual tasks and applications.