Comparative Analysis of Malware Classification Using Supervised Machine Learning Algorithms
摘要
Privacy is a myth, a statement persistently encountered when talking about the world of Internet. Malwares are a constant, ominous threat to data which cripples the cyberspace today. The myth of digital privacy began with the conception and subsequent proliferation of malwares. Any device connected through the Internet is a potential target and runs the risk of its security being breached and information being compromised. In this paper, a benchmarked dataset Big 2015 is used for the malware classification experiment. Seven different machine learning models namely Random Forest, Support Vector Machines, Logistic Regression, Naïve Bayes, AdaBoost, Gradient Boost and Bagging, are used to train and test the dataset and to establish the one that performs the best. The performance metrics put in place are Accuracy, Precision, Recall and F1-score. It is seen that ensemble machine learning approach, namely Random Forest, Bagging and Gradient Boost performed better in accordance to the performance parameters considered.