In recent years, malicious SMS creators have frequently employed Chinese camouflaging strategies such as homophone replacement, homograph replacement, and symbol substitution to encode malicious sensitive information in a non-sensitive form, aiming to evade the interception rules of regulatory platforms. However, most current malicious SMS detection models primarily focus on enhancing the performance of deep learning models and have failed to effectively address the adversarial risks posed by such Chinese camouflaged malicious SMS. In response to this situation, this paper first conducts a comprehensive investigation and summary of the commonly used Chinese text camouflaging techniques by malicious SMS creators. Based on this foundation, this paper innovatively proposes a malicious SMS detection method based on multi-modal fusion. Specifically, for traditional malicious SMS, we utilize the pre-trained model BERT-WWW to construct a detection model based on semantics; for visually similar characters and symbol substitution, we develop an image-based malicious SMS detection model; and to counter homophone character circumvention techniques, we also train a pinyin-based detection model. Finally, we effectively fuse these three modal detection models (semantics, pinyin, and image) at both the feature level and the decision level, significantly enhancing the robustness of the overall detection algorithm.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Malicious SMS Detection Based on Multi-Modal Fusion

  • Jiayin Liu,
  • Wensheng Huang,
  • Chenyang Qian,
  • Minxi Zhang

摘要

In recent years, malicious SMS creators have frequently employed Chinese camouflaging strategies such as homophone replacement, homograph replacement, and symbol substitution to encode malicious sensitive information in a non-sensitive form, aiming to evade the interception rules of regulatory platforms. However, most current malicious SMS detection models primarily focus on enhancing the performance of deep learning models and have failed to effectively address the adversarial risks posed by such Chinese camouflaged malicious SMS. In response to this situation, this paper first conducts a comprehensive investigation and summary of the commonly used Chinese text camouflaging techniques by malicious SMS creators. Based on this foundation, this paper innovatively proposes a malicious SMS detection method based on multi-modal fusion. Specifically, for traditional malicious SMS, we utilize the pre-trained model BERT-WWW to construct a detection model based on semantics; for visually similar characters and symbol substitution, we develop an image-based malicious SMS detection model; and to counter homophone character circumvention techniques, we also train a pinyin-based detection model. Finally, we effectively fuse these three modal detection models (semantics, pinyin, and image) at both the feature level and the decision level, significantly enhancing the robustness of the overall detection algorithm.