With the increasing diversity and complexity of malicious code, the threats posed by malicious code are continuously growing, making malicious code detection increasingly challenging. Currently, researchers typically extract malware features and use deep-learning models for malware detection. Using image-based features for classification can improve the accuracy and efficiency of malware detection. However, traditional deep learning-based malware detection faces challenges such as data imbalance, insufficient data volume, and malware obfuscation. We propose to augment the malicious code dataset using Generative Adversarial Network (GAN) technology to improve the accuracy and efficiency of malicious code detection. First, we converted malicious software binary files into grayscale images. We used Generative Adversarial Network (GAN), Wasserstein Generative Adversarial Network (WGAN), and Wasserstein Generative Adversarial Networks-Gradient Penalty (WGAN-GP) models to augment the grayscale image dataset of malicious code. Subsequently, we obtained three augmented datasets: w1, w2, and w3. These datasets were then fed into the ResNet50 neural network model for training. Experimental results indicate that using the GAN model for data augmentation improved ResNet’s recognition accuracy by approximately 10%. When using the WGAN model for data augmentation, ResNet’s recognition accuracy increased by 2%. Similarly, using the WGAN-GP model for data augmentation also increased ResNet’s recognition accuracy by 2%. Among the three models, the GAN model demonstrated the strongest improvement effect. We demonstrate that the ResNet50 network with GAN models can significantly improve malware code detection.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Malicious Code Detection Based on Generative Adversarial Model

  • Jinzhihao Zhang,
  • Jia Yang,
  • Weiqi Zhou

摘要

With the increasing diversity and complexity of malicious code, the threats posed by malicious code are continuously growing, making malicious code detection increasingly challenging. Currently, researchers typically extract malware features and use deep-learning models for malware detection. Using image-based features for classification can improve the accuracy and efficiency of malware detection. However, traditional deep learning-based malware detection faces challenges such as data imbalance, insufficient data volume, and malware obfuscation. We propose to augment the malicious code dataset using Generative Adversarial Network (GAN) technology to improve the accuracy and efficiency of malicious code detection. First, we converted malicious software binary files into grayscale images. We used Generative Adversarial Network (GAN), Wasserstein Generative Adversarial Network (WGAN), and Wasserstein Generative Adversarial Networks-Gradient Penalty (WGAN-GP) models to augment the grayscale image dataset of malicious code. Subsequently, we obtained three augmented datasets: w1, w2, and w3. These datasets were then fed into the ResNet50 neural network model for training. Experimental results indicate that using the GAN model for data augmentation improved ResNet’s recognition accuracy by approximately 10%. When using the WGAN model for data augmentation, ResNet’s recognition accuracy increased by 2%. Similarly, using the WGAN-GP model for data augmentation also increased ResNet’s recognition accuracy by 2%. Among the three models, the GAN model demonstrated the strongest improvement effect. We demonstrate that the ResNet50 network with GAN models can significantly improve malware code detection.