In the digital age, there’s a crucial necessity for developing machine learning models tailored for threat detection within audio data. This research bridges the gap between text and audio data analysis by examining the robustness and adaptability of various machine learning and deep learning methodologies when the data medium transitions from text to audio. The study is divided into two folds: first began by training models on text data, followed by adopting a multimodal strategy to classify audio data. The overarching goal is to formulate models proficient at pinpointing threats within audio data, thereby enhancing the safety of digital communication. The employed models include Logistic Regression, Gradient Boosting, Naive Bayes, Random Forest, SVM, XGBoost, LSTM, and RNN. Their efficacy was gauged using metrics such as accuracy, precision, recall, F1 score, ROC AUC, and Cohen’s Kappa. Notably, the XGBoost model showcased exemplary performance on both the text and audio datasets with average accuracies of 88% and 76% respectively, leveraging 10-fold cross-validation during its training phase. The outcomes of this research could offer invaluable insights for future endeavors focused on threat detection in audio data using models predominantly trained on text.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Conversational Risk Analysis: Identifying and Classifying Potential Threats in Interpersonal Communication

  • Ankit Bansal,
  • Srikanth Bharadwaj,
  • Prithvi Panchineni,
  • Swathi Suddala,
  • Shashikant Chaudhary

摘要

In the digital age, there’s a crucial necessity for developing machine learning models tailored for threat detection within audio data. This research bridges the gap between text and audio data analysis by examining the robustness and adaptability of various machine learning and deep learning methodologies when the data medium transitions from text to audio. The study is divided into two folds: first began by training models on text data, followed by adopting a multimodal strategy to classify audio data. The overarching goal is to formulate models proficient at pinpointing threats within audio data, thereby enhancing the safety of digital communication. The employed models include Logistic Regression, Gradient Boosting, Naive Bayes, Random Forest, SVM, XGBoost, LSTM, and RNN. Their efficacy was gauged using metrics such as accuracy, precision, recall, F1 score, ROC AUC, and Cohen’s Kappa. Notably, the XGBoost model showcased exemplary performance on both the text and audio datasets with average accuracies of 88% and 76% respectively, leveraging 10-fold cross-validation during its training phase. The outcomes of this research could offer invaluable insights for future endeavors focused on threat detection in audio data using models predominantly trained on text.