Regional dialectal variation in Bangla language is one of the major challenges for natural language processing applications, as existing systems focus on standard Bangla while regional dialects remain underrepresented. With over 265 million speakers worldwide, automated dialect identification for Bangla remains an challenging area that limits the development of dialect-aware NLP systems. This study addresses this gap by proposing a deep learning method for the categorization of four famous Bangladeshi dialects: Chittagong, Sylhet, Barisal, and Standard Bangla based on social media and online news website text data. Due to the lack of standardized datasets for Bangladeshi dialect classification, a highly curated dataset of 980 labeled sentences of each class were preprocessed using tokenization, normalization, and sequence padding. To overcome the complexity of dialectal features and overfitting issues common with limited dialectal data, a combination architecture was implemented by stacking Bidirectional Long Short-Term Memory (BiLSTM) and Bidirectional Gated Recurrent Unit (BiGRU) layers. The architecture was further enhanced by incorporating L2 regularization and dropout techniques. Data augmentation methods were utilized to address class imbalance issues and improve model robustness. The proposed model achieved 87.57% accuracy in dialect classification, demonstrating the effectiveness of the hybrid deep learning approach for Bangladeshi dialect identification. This work contributes to developing inclusive NLP technologies that can effectively handle Bangla’s linguistic diversity, with applications in social media analysis and culturally-aware language technologies.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Hybrid BiLSTM-BiGRU Model for Classification of Bangladeshi Dialects

  • Md. Tahidul Islam,
  • Md. Abu Johab,
  • Md. Roton Ahmed,
  • Nakib Aman

摘要

Regional dialectal variation in Bangla language is one of the major challenges for natural language processing applications, as existing systems focus on standard Bangla while regional dialects remain underrepresented. With over 265 million speakers worldwide, automated dialect identification for Bangla remains an challenging area that limits the development of dialect-aware NLP systems. This study addresses this gap by proposing a deep learning method for the categorization of four famous Bangladeshi dialects: Chittagong, Sylhet, Barisal, and Standard Bangla based on social media and online news website text data. Due to the lack of standardized datasets for Bangladeshi dialect classification, a highly curated dataset of 980 labeled sentences of each class were preprocessed using tokenization, normalization, and sequence padding. To overcome the complexity of dialectal features and overfitting issues common with limited dialectal data, a combination architecture was implemented by stacking Bidirectional Long Short-Term Memory (BiLSTM) and Bidirectional Gated Recurrent Unit (BiGRU) layers. The architecture was further enhanced by incorporating L2 regularization and dropout techniques. Data augmentation methods were utilized to address class imbalance issues and improve model robustness. The proposed model achieved 87.57% accuracy in dialect classification, demonstrating the effectiveness of the hybrid deep learning approach for Bangladeshi dialect identification. This work contributes to developing inclusive NLP technologies that can effectively handle Bangla’s linguistic diversity, with applications in social media analysis and culturally-aware language technologies.