This research explores the effectiveness of data augmentation in enhancing multi-label text classification for Hindi news data. We address the inherent challenge of limited and imbalanced data in Hindi Natural Language Processing (NLP) tasks by collecting and annotating news articles across various categories like economy, education, and entertainment. A key innovation lies in applying data augmentation techniques specifically tailored for the Hindi language. This includes synonym replacement, random insertion, and swap, all operating on Hindi words. We demonstrate the efficacy of this approach by comparing a machine learning model trained on the original dataset to one trained on the augmented dataset. Our results showcase significant improvements in model performance across micro-averaged F1-score, macro-averaged F1-score, and hamming loss metrics. This study highlights the potential of data augmentation for improving multi-label text classification, particularly for resource-constrained languages like Hindi language.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Data Augmentation to Achieve Robust Multi-label Classification of Hindi News

  • Swati Mathur,
  • Pratistha Mathur

摘要

This research explores the effectiveness of data augmentation in enhancing multi-label text classification for Hindi news data. We address the inherent challenge of limited and imbalanced data in Hindi Natural Language Processing (NLP) tasks by collecting and annotating news articles across various categories like economy, education, and entertainment. A key innovation lies in applying data augmentation techniques specifically tailored for the Hindi language. This includes synonym replacement, random insertion, and swap, all operating on Hindi words. We demonstrate the efficacy of this approach by comparing a machine learning model trained on the original dataset to one trained on the augmented dataset. Our results showcase significant improvements in model performance across micro-averaged F1-score, macro-averaged F1-score, and hamming loss metrics. This study highlights the potential of data augmentation for improving multi-label text classification, particularly for resource-constrained languages like Hindi language.