Enhancing Low-Resource Bangla Fake News Detection through Deep Convolutional Neural Networks
摘要
In today’s digital landscape, the detection of false information is of utmost importance, especially in languages like Bangla, which lack abundant natural language processing (NLP) resources. The rapid spread of misinformation through online platforms, particularly within Bangla-speaking communities, has become a pressing concern. However, the limited availability of NLP tools for Bangla has posed significant challenges in developing reliable models for identifying deceptive content. In response to these challenges, researchers have made notable progress in classifying Bangla news using deep learning techniques and language models like BERT. This study presents a detailed exploration of a deep convolutional neural network (CNN) model tailored for categorizing Bangla news articles as authentic or counterfeit. By integrating BERT (Bangla Electra) into the model’s architecture, an impressive accuracy rate of 94.33% was achieved, with our proprietary model surpassing this at 94.5%.To ensure the reliability of results, a range of NLP techniques were applied during data preprocessing, including data cleansing, tokenization, stop word and punctuation removal, and stemming. Feature extraction involved the combined use of TF-IDF and Bag of Words techniques. The dataset, obtained from Kaggle, comprised 7,000 genuine news texts and 1,000 counterfeit news texts. In summary, this research significantly contributes to Bangla news classification by showcasing the effectiveness of deep CNN models in accurately discerning between legitimate and fabricated news articles.