Recently, ransomware has grown and developed into one of the most terrifying threats against data security. Therefore, detection approaches are highly in demand in cybersecurity. In this regard, this paper discusses the impacts of balancing techniques on the performance of various machine learning classifiers for ransomware detection, namely, Random Forest, AdaBoost, Naive Bayes, Logistic Regression, Support Vector Classifier (SVC), Multi-layer Perceptron, and K-Nearest Neighbors. In this context, we perform an empirical study by applying advanced preprocessing techniques—namely, feature selection and data normalization—together with balancing techniques to handle class imbalance. The classifiers are also tested on imbalanced and balanced datasets to check their performance in ransomware detection. It is observed that Random Forest produces accuracy of 100% on both the unbalanced and the balanced datasets, while both Logistic Regression and Naive Bayes generate lower performance on the imbalanced dataset. However, when applying balancing techniques, classifier performance radically improves, reaching 99% for both AdaBoost and Multi-layer Perceptron, while Naive Bayes reaches 88%, and Logistic Regression rises as high as 94%, with large increases in AUC. It indicates that balancing techniques provide a great role in fine-tuning machine learning models when applied for the detection of ransomware. This research provides insight into how such balancing can enhance the effectiveness of a classifier, hence trying to develop more accurate mechanisms for the detection and robust defense mechanisms against such ransomware threats, which keep on evolving continuously.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Comprehensive Approach to Ransomware Detection Utilizing Diverse Machine Learning Models and Balancing Techniques

  • Md. Ehsanul Haque,
  • Amran Hossain,
  • Akash Barua,
  • Jeba Maliha,
  • Ahsan Habib Siam,
  • Md. Shafiqul Alam

摘要

Recently, ransomware has grown and developed into one of the most terrifying threats against data security. Therefore, detection approaches are highly in demand in cybersecurity. In this regard, this paper discusses the impacts of balancing techniques on the performance of various machine learning classifiers for ransomware detection, namely, Random Forest, AdaBoost, Naive Bayes, Logistic Regression, Support Vector Classifier (SVC), Multi-layer Perceptron, and K-Nearest Neighbors. In this context, we perform an empirical study by applying advanced preprocessing techniques—namely, feature selection and data normalization—together with balancing techniques to handle class imbalance. The classifiers are also tested on imbalanced and balanced datasets to check their performance in ransomware detection. It is observed that Random Forest produces accuracy of 100% on both the unbalanced and the balanced datasets, while both Logistic Regression and Naive Bayes generate lower performance on the imbalanced dataset. However, when applying balancing techniques, classifier performance radically improves, reaching 99% for both AdaBoost and Multi-layer Perceptron, while Naive Bayes reaches 88%, and Logistic Regression rises as high as 94%, with large increases in AUC. It indicates that balancing techniques provide a great role in fine-tuning machine learning models when applied for the detection of ransomware. This research provides insight into how such balancing can enhance the effectiveness of a classifier, hence trying to develop more accurate mechanisms for the detection and robust defense mechanisms against such ransomware threats, which keep on evolving continuously.