In the dynamic field of cybersecurity, the development of effective, scalable intrusion detection systems is crucial. Utilizing distributed computing frameworks like Apache Spark on Hadoop YARN offers significant scalability and performance benefits, essential for modern intrusion detection tasks. This study evaluates the efficacy of various machine learning classifiers—Decision Trees, Random Forests, SVMs, Logistic Regression, GBT, MLP, and Naive Bayes—within such an environment. Our findings highlight notable disparities in performance metrics like accuracy, precision, recall, F1-score, and AUC among these classifiers. Notably, the GBT and chi-squared optimized random forest classifiers excelled in accuracy, F1-scores, and AUC. We also examined the impact of chi-squared feature selection on model performance, observing marked improvements in accuracy and computational efficiency. Further, the study underscores the significance of balancing false negatives and positives in NIDS, emphasizing comprehensive pre-deployment attack evaluation. Real-world simulations on a Spark-Hadoop cluster affirmed the system’s scalability and robustness, showing consistent performance across various log sizes while maintaining high accuracy. Looking ahead, we aim to leverage GPU support in Spark 3 for enhancing MLP models and integrate a real-time firewall log parser into the Hadoop ecosystem to bolster threat detection capabilities. This research underscores the value of distributed computing in fortifying cybersecurity, offering insights and advancements to counteract cyberthreats effectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Toward Scalable Security: Intrusion Detection with Spark-Based Classifiers on Hadoop YARN

  • Omar Alhammadi,
  • Ashraf Elnagar

摘要

In the dynamic field of cybersecurity, the development of effective, scalable intrusion detection systems is crucial. Utilizing distributed computing frameworks like Apache Spark on Hadoop YARN offers significant scalability and performance benefits, essential for modern intrusion detection tasks. This study evaluates the efficacy of various machine learning classifiers—Decision Trees, Random Forests, SVMs, Logistic Regression, GBT, MLP, and Naive Bayes—within such an environment. Our findings highlight notable disparities in performance metrics like accuracy, precision, recall, F1-score, and AUC among these classifiers. Notably, the GBT and chi-squared optimized random forest classifiers excelled in accuracy, F1-scores, and AUC. We also examined the impact of chi-squared feature selection on model performance, observing marked improvements in accuracy and computational efficiency. Further, the study underscores the significance of balancing false negatives and positives in NIDS, emphasizing comprehensive pre-deployment attack evaluation. Real-world simulations on a Spark-Hadoop cluster affirmed the system’s scalability and robustness, showing consistent performance across various log sizes while maintaining high accuracy. Looking ahead, we aim to leverage GPU support in Spark 3 for enhancing MLP models and integrate a real-time firewall log parser into the Hadoop ecosystem to bolster threat detection capabilities. This research underscores the value of distributed computing in fortifying cybersecurity, offering insights and advancements to counteract cyberthreats effectively.