<p>Malicious URL identification remains a core cybersecurity challenge due to the ever-changing tactics of phishing, defacement, and malware hosting. Conventional methods that rely on lexical features or pre-trained language models often miss the subtle structural and semantic signals in obfuscated URLs. We present URLMoE, a multi-expert architecture that uses a mixture-of-experts design to fuse three complementary perspectives: (i) a text expert built on BERT with layer-wise semantic fusion; (ii) a structural expert that employs dilated convolutional networks to model dependencies among URL tokens; and (iii) a handcrafted-metadata expert that captures entropy, digit ratio, and semantic cues. The expert representations are integrated with an improved multi-head attention mechanism and a gating module, enabling adaptive focus across modalities. Auxiliary classification heads for each expert further encourage specialized learning during training. Extensive experiments on the malicious phish benchmark dataset show that URLMoE achieves a combined accuracy and F1-score of 98.99%, significantly outperforming single-expert and conventional transformer-based baselines. The proposed model offers a robust and interpretable framework for detecting malware URLs and phishing in the real world.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

URLMoE: leveraging multi-head attention and expert specialization for robust malicious URL classification

  • Fatin A. Elhaj,
  • Taqwa A. Alhaj,
  • Shamiel H. Ibrahim,
  • Omayma Husain,
  • Tariq M. Khan,
  • Tasneem Darwish

摘要

Malicious URL identification remains a core cybersecurity challenge due to the ever-changing tactics of phishing, defacement, and malware hosting. Conventional methods that rely on lexical features or pre-trained language models often miss the subtle structural and semantic signals in obfuscated URLs. We present URLMoE, a multi-expert architecture that uses a mixture-of-experts design to fuse three complementary perspectives: (i) a text expert built on BERT with layer-wise semantic fusion; (ii) a structural expert that employs dilated convolutional networks to model dependencies among URL tokens; and (iii) a handcrafted-metadata expert that captures entropy, digit ratio, and semantic cues. The expert representations are integrated with an improved multi-head attention mechanism and a gating module, enabling adaptive focus across modalities. Auxiliary classification heads for each expert further encourage specialized learning during training. Extensive experiments on the malicious phish benchmark dataset show that URLMoE achieves a combined accuracy and F1-score of 98.99%, significantly outperforming single-expert and conventional transformer-based baselines. The proposed model offers a robust and interpretable framework for detecting malware URLs and phishing in the real world.