<p>The rapid proliferation of hate speech in digital spaces necessitates robust and language-specific detection frameworks, especially for underrepresented languages such as Malayalam. This study introduces a comprehensive deep learning framework for detecting hate speech in accented Malayalam speech, integrating advanced feature engineering, class balancing, and robustness evaluation. A diverse dataset was curated from Malayalam YouTube videos and movies to capture phonetic, dialectal, and prosodic variations. Distinct acoustic features-including Zero Crossing Rate (ZCR), Short-Time Fourier Transform (STFT), Mel-Frequency Cepstral Coefficients (MFCC), Root Mean Square (RMS), and Mel Spectrogram-were extracted, producing 162 feature vectors per utterance. Data augmentation techniques, including noise injection, time stretching, and pitch shifting, were applied to enhance diversity. A customized 1D Convolutional Neural Network (CNN) was developed for binary classification of hate and non-hate speech. Building upon the initial study, this extended research introduces SMOTE-based class balancing, speaker-independent validation, and robustness testing using Gaussian noise to simulate real-world acoustic distortions. Multiple models, including CNN, Support Vector Machine(SVM), Random Forest, XGBoost, and LightGBM, were evaluated under stratified, hold-out, and noise-perturbed conditions. The CNN achieved an accuracy of 95.3%, F1-score of 0.966, and AUC of 0.998 on the noisy hold-out set, demonstrating strong generalization and noise resilience. These results confirm that the integration of balanced learning, multi-level validation, and robustness evaluation significantly enhances the reliability of hate speech detection in accented and low-resource Malayalam speech, establishing a transferable framework for multilingual content moderation. To address potential overfitting and data leakage risks noted in the initial evaluations, this extended study introduces cross-dataset validation, detailed speaker-independent and hold-out testing, and rigorous noise robustness analysis. The findings are revalidated using augmented and balanced data under multiple evaluation settings to ensure reliability and generalizability.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A novel deep learning framework with advanced feature engineering for hate speech detection in accented Malayalam speech

  • Rizwana Kallooravi Thandil,
  • Muneer V.K,
  • Houlath K,
  • Sabique P.V

摘要

The rapid proliferation of hate speech in digital spaces necessitates robust and language-specific detection frameworks, especially for underrepresented languages such as Malayalam. This study introduces a comprehensive deep learning framework for detecting hate speech in accented Malayalam speech, integrating advanced feature engineering, class balancing, and robustness evaluation. A diverse dataset was curated from Malayalam YouTube videos and movies to capture phonetic, dialectal, and prosodic variations. Distinct acoustic features-including Zero Crossing Rate (ZCR), Short-Time Fourier Transform (STFT), Mel-Frequency Cepstral Coefficients (MFCC), Root Mean Square (RMS), and Mel Spectrogram-were extracted, producing 162 feature vectors per utterance. Data augmentation techniques, including noise injection, time stretching, and pitch shifting, were applied to enhance diversity. A customized 1D Convolutional Neural Network (CNN) was developed for binary classification of hate and non-hate speech. Building upon the initial study, this extended research introduces SMOTE-based class balancing, speaker-independent validation, and robustness testing using Gaussian noise to simulate real-world acoustic distortions. Multiple models, including CNN, Support Vector Machine(SVM), Random Forest, XGBoost, and LightGBM, were evaluated under stratified, hold-out, and noise-perturbed conditions. The CNN achieved an accuracy of 95.3%, F1-score of 0.966, and AUC of 0.998 on the noisy hold-out set, demonstrating strong generalization and noise resilience. These results confirm that the integration of balanced learning, multi-level validation, and robustness evaluation significantly enhances the reliability of hate speech detection in accented and low-resource Malayalam speech, establishing a transferable framework for multilingual content moderation. To address potential overfitting and data leakage risks noted in the initial evaluations, this extended study introduces cross-dataset validation, detailed speaker-independent and hold-out testing, and rigorous noise robustness analysis. The findings are revalidated using augmented and balanced data under multiple evaluation settings to ensure reliability and generalizability.