UCoD: Ensemble BERT for Hierarchical Classification of the Urdu Disinformation Corpus
摘要
Online disinformation poses a growing threat, requiring fact-checking and detection/prevention measures. To address this, we propose a hierarchical classification approach using the DistilBERT and XLM-RoBERTa ensemble architectures on the Urdu Corpus of Disinformation (UCoD). Our ensemble outperforms other models like RNNs, LSTMs, k-nearest neighbors, random forests, and quadratic discriminant analysis, achieving a weighted F1 of 68.7 on UCoD. These results confirm the advantage of ensembles for imbalanced corpora, supporting the use of deep learning techniques in combating disinformation.