This article focuses on the management of DNS system regarding to the introduction and promotion of Internationalized Domain Names (IDNs). In fact, Internationalized Domain Names (IDNs) allow users to use non-Latin characters to register domain names, enabling the inclusion of various languages. However IDN has triggered more potential DNS abuses. To address these additional risks of abuses initiatives have been developed focusing primarily on two types of attacks: semantic and homograph attacks. Nonetheless, the main drawbacks of these solutions are their complexity and resources consuming characteristics. The aim of this paper is to address these challenges by proposing an approach that is resource-efficient and easy to implement. Our approach is based on the calculation of visual similarity using Siamese neural networks to identify homographs that can be used for malicous purposes to deviate people from the correct domain. We defined and trained our model on 5000 domain names of the .SN ccTLD. The model is promising since it helps detecting potential risks of abuse at the domain name creation level, allowing possibilities to use mitigation mechanisms. Despite the various difficulties of Overfitting due to the size of the dataset and the length of the domain names, the model was able to achieve an accuracy of 98.7%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Preventing Typosquatting with IDN Conversion: A Siamese Neural Network Approach

  • Evrard Cabrel Nguemeyou Tchouangang,
  • Idrissa Sarr,
  • Alex Corenthin,
  • Bassirou Kassé

摘要

This article focuses on the management of DNS system regarding to the introduction and promotion of Internationalized Domain Names (IDNs). In fact, Internationalized Domain Names (IDNs) allow users to use non-Latin characters to register domain names, enabling the inclusion of various languages. However IDN has triggered more potential DNS abuses. To address these additional risks of abuses initiatives have been developed focusing primarily on two types of attacks: semantic and homograph attacks. Nonetheless, the main drawbacks of these solutions are their complexity and resources consuming characteristics. The aim of this paper is to address these challenges by proposing an approach that is resource-efficient and easy to implement. Our approach is based on the calculation of visual similarity using Siamese neural networks to identify homographs that can be used for malicous purposes to deviate people from the correct domain. We defined and trained our model on 5000 domain names of the .SN ccTLD. The model is promising since it helps detecting potential risks of abuse at the domain name creation level, allowing possibilities to use mitigation mechanisms. Despite the various difficulties of Overfitting due to the size of the dataset and the length of the domain names, the model was able to achieve an accuracy of 98.7%.