Leveraging CuMeta for enhanced document classification in cursive languages with transformer stacking
摘要
Classifying textual data is crucial in the expanding digital landscape, especially for underrepresented cursive languages like Urdu, which pose unique challenges due to their intricate linguistic features and vast digital content. Despite the growing demand, research in effective Urdu text classification remains limited. This study introduces CuMeta, a novel meta-stacking classifier designed specifically to address these challenges. CuMeta integrates the strengths of various deep learning algorithms, including CNN and LSTM, alongside multiple Transformer-based models, to provide a comprehensive solution to issues such as tokenization, spatial inconsistencies, lack of case distinctions, and complex dialects inherent in Urdu. Through rigorous evaluation across diverse datasets, CuMeta has demonstrated its superiority over traditional and deep learning models, achieving an impressive F1-score of 97.52%. The innovative approach of CuMeta not only enhances text classification for Urdu but also offers significant implications for other low-resource, cursive-script languages like Arabic and Persian. This research marks a significant advancement in NLP, particularly for languages that have historically been underrepresented in digital and academic platforms.