The arms race between malware authors and defenders is characterized by (1) mutations to the malware samples and (2) model retraining to detect those mutations. Due to an exponential growth in the number of new malware samples reported per day (1.5 million [4]), detection frameworks’ reliance on model retraining naturally increased. Model retraining is the de facto approach to counter malware mutations. In this paper, we question the efficacy of machine learning in the context of malware detection by exposing various limitations in the retraining approaches. We show that model retraining only provides a marginal performance improvement for malicious sample detection while simultaneously degrading the benign sample detection performance. To address various issues in malware detection, we investigate the efficiency of several model retraining approaches. Our proposed approaches allow the malware detectors to retrain models in time to enable malware family emergence detection while concurrently monitoring the evolving patterns of malware family mutations.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exposing the Limitations of Machine Learning for Malware Detection Under Concept Drift

  • Ahmed Abusnaina,
  • Afsah Anwar,
  • Muhammad Saad,
  • Abdulrahman Alabduljabbar,
  • Rhongho Jang,
  • Saeed Salem,
  • David Mohaisen

摘要

The arms race between malware authors and defenders is characterized by (1) mutations to the malware samples and (2) model retraining to detect those mutations. Due to an exponential growth in the number of new malware samples reported per day (1.5 million [4]), detection frameworks’ reliance on model retraining naturally increased. Model retraining is the de facto approach to counter malware mutations. In this paper, we question the efficacy of machine learning in the context of malware detection by exposing various limitations in the retraining approaches. We show that model retraining only provides a marginal performance improvement for malicious sample detection while simultaneously degrading the benign sample detection performance. To address various issues in malware detection, we investigate the efficiency of several model retraining approaches. Our proposed approaches allow the malware detectors to retrain models in time to enable malware family emergence detection while concurrently monitoring the evolving patterns of malware family mutations.