<p>This paper presents an extensive review and experimental framework for hate speech detection using large language models (LLMs). We explored traditional machine learning approaches and illustrated how state-of-the-art LLMs, such as BERT variants (DistilBERT and ModernBERT) and recent developments, such as DeepSeek R1, Llama 3, Gemma 3, Mistral, and Phi 3.5, enhance text classification performance. In this study, we evaluated the effectiveness of various large language models (LLMs) for Hindi–English code-mixed hate speech detection. We constructed a custom dataset using Reddit posts on social media, which were auto-labeled and subsequently verified and validated by expert annotators. LoRa adapters were used for model fine-tuning and in-domain performance evaluation of the model, with precision, recall, F1-score, and accuracy as metrics. To ensure a fair comparison, all the models were trained for the same number of epochs under the same training conditions. According to our findings, DeepSeek R1 outperformed larger general-purpose LLMs, such as Llama 3 and Gemma 3, with smaller margins. DeepSeek R1 surpassed the other models with the highest accuracy of 79 percent, demonstrating a better understanding of the complexity of code-mixed hate speech detection. These results indicate that lightweight, responsibly fine-tuned LLMs can strengthen moderation for multilingual, code-mixed communities without prohibitive computing costs. In practice, the approach offers a clear path towards safer, more inclusive platforms by pairing accuracy with bias-aware development and human oversight.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine tuning large language models for hate speech detection in Hinglish and code mixed custom dataset through a socially responsible approach for safer digital platforms

  • Bhawani Singh Rathore,
  • Sandeep Chaurasia

摘要

This paper presents an extensive review and experimental framework for hate speech detection using large language models (LLMs). We explored traditional machine learning approaches and illustrated how state-of-the-art LLMs, such as BERT variants (DistilBERT and ModernBERT) and recent developments, such as DeepSeek R1, Llama 3, Gemma 3, Mistral, and Phi 3.5, enhance text classification performance. In this study, we evaluated the effectiveness of various large language models (LLMs) for Hindi–English code-mixed hate speech detection. We constructed a custom dataset using Reddit posts on social media, which were auto-labeled and subsequently verified and validated by expert annotators. LoRa adapters were used for model fine-tuning and in-domain performance evaluation of the model, with precision, recall, F1-score, and accuracy as metrics. To ensure a fair comparison, all the models were trained for the same number of epochs under the same training conditions. According to our findings, DeepSeek R1 outperformed larger general-purpose LLMs, such as Llama 3 and Gemma 3, with smaller margins. DeepSeek R1 surpassed the other models with the highest accuracy of 79 percent, demonstrating a better understanding of the complexity of code-mixed hate speech detection. These results indicate that lightweight, responsibly fine-tuned LLMs can strengthen moderation for multilingual, code-mixed communities without prohibitive computing costs. In practice, the approach offers a clear path towards safer, more inclusive platforms by pairing accuracy with bias-aware development and human oversight.