<p>Classifying textual data is crucial in the expanding digital landscape, especially for underrepresented cursive languages like Urdu, which pose unique challenges due to their intricate linguistic features and vast digital content. Despite the growing demand, research in effective Urdu text classification remains limited. This study introduces CuMeta, a novel meta-stacking classifier designed specifically to address these challenges. CuMeta integrates the strengths of various deep learning algorithms, including CNN and LSTM, alongside multiple Transformer-based models, to provide a comprehensive solution to issues such as tokenization, spatial inconsistencies, lack of case distinctions, and complex dialects inherent in Urdu. Through rigorous evaluation across diverse datasets, CuMeta has demonstrated its superiority over traditional and deep learning models, achieving an impressive F1-score of 97.52%. The innovative approach of CuMeta not only enhances text classification for Urdu but also offers significant implications for other low-resource, cursive-script languages like Arabic and Persian. This research marks a significant advancement in NLP, particularly for languages that have historically been underrepresented in digital and academic platforms.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging CuMeta for enhanced document classification in cursive languages with transformer stacking

  • Muhammad Shahid,
  • Muhammad Amjad Iqbal,
  • Muhammad Umair

摘要

Classifying textual data is crucial in the expanding digital landscape, especially for underrepresented cursive languages like Urdu, which pose unique challenges due to their intricate linguistic features and vast digital content. Despite the growing demand, research in effective Urdu text classification remains limited. This study introduces CuMeta, a novel meta-stacking classifier designed specifically to address these challenges. CuMeta integrates the strengths of various deep learning algorithms, including CNN and LSTM, alongside multiple Transformer-based models, to provide a comprehensive solution to issues such as tokenization, spatial inconsistencies, lack of case distinctions, and complex dialects inherent in Urdu. Through rigorous evaluation across diverse datasets, CuMeta has demonstrated its superiority over traditional and deep learning models, achieving an impressive F1-score of 97.52%. The innovative approach of CuMeta not only enhances text classification for Urdu but also offers significant implications for other low-resource, cursive-script languages like Arabic and Persian. This research marks a significant advancement in NLP, particularly for languages that have historically been underrepresented in digital and academic platforms.