One step forward, two steps back: ML-based malware detection under concept drift
摘要
The arms race between malware authors and detection frameworks is marked by continuous malware mutations and corresponding model retraining efforts. With over 1.5 million new samples reported daily (VirusTotal Statistics, 2025 https://www.virustotal.com/en/statistics/), retraining has become the de facto response to evolving threats. In this paper, we question the effectiveness of this approach by exposing key limitations: while retraining offers only marginal improvements in detecting malicious samples, it often degrades performance on benign samples. To address these challenges, we evaluate multiple retraining strategies that enable the timely detection of emerging malware families while tracking mutation patterns. Our analysis reveals that retraining can unintentionally aid adversaries by allowing the reuse of old malware samples, which online detectors often discard. Additionally, we uncover labeling inconsistencies−such as family renaming−across online detection engines, which obscure shared malicious capabilities and weaken family-based detection efforts.