EM-BERT: A Language Model Based Method to Detect Encrypted Malicious Network Traffic
摘要
The evolving evasion and encryption techniques equip malware with lethal weapons and pose substantial serious risk to society. In order to detect encrypted malware, various detection methods have been proposed that rely on traffic characteristics to identify malicious behaviors. However, encryption poses a significant challenge to existing detection methods that rely on a single type of characteristic. To address this issue, we propose a language model-based detection method called EM-BERT that leverages representation learning algorithms to explore the spatial and temporal characteristics of network behaviors. We also developed a pre-processing method to eliminate irrelevant information and improve detection accuracy. To demonstrate the effectiveness of our approach, we evaluated it on a large dataset consisting of multiple open-access datasets and a self-built one. Our experimental results show that EM-BERT, combined with our pre-processing method, outperforms baseline malware detection systems and exhibits strong generalizability and robustness, achieving precision and recall rates exceeding 99.9%.