<p>The rapid growth of domain registrations and Internet use has led to an increase in abusive activities such as gambling, fraudulence, phishing, and malware. This circumstance occurs in most countries where the Internet is used as a common platform for communications. For example, in Vietnam, where there are over 600,000 registered domains, these abusive activities have severely caused financial and reputational damages and are extremely difficult to monitor. The existing manual approach to monitoring and identifying abusive domains is inefficient and labor-intensive. In this paper, we propose hybrid learning models that take advantage of domain lexical and metadata analysis to evaluate the credibility of registered domains and associated websites. Specifically, our hybrid learning models are designed as a fusion of joint features of the lexical domain encoder and RoBERTa-based encoders. To enhance the robustness of the models in practice, we applied a three-stage pipeline with a web-based interface. Throughout an extensive evaluation with realistic datasets, the proposed models demonstrate a substantial improvement over conventional baselines, offering an efficient and scalable solution to monitor abusive web domains in general.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Hybrid Learning-Based Approach with Metadata Analysis to Credibility Classification of Web Domains

  • Tran The Son,
  • Bao D. Nguyen,
  • Dung Q. T. Truong,
  • Sang A. Phung,
  • Ron T. Ton,
  • Dung D. Truong,
  • Nam V. Pham,
  • Hang Dinh,
  • Minh N. H. Nguyen

摘要

The rapid growth of domain registrations and Internet use has led to an increase in abusive activities such as gambling, fraudulence, phishing, and malware. This circumstance occurs in most countries where the Internet is used as a common platform for communications. For example, in Vietnam, where there are over 600,000 registered domains, these abusive activities have severely caused financial and reputational damages and are extremely difficult to monitor. The existing manual approach to monitoring and identifying abusive domains is inefficient and labor-intensive. In this paper, we propose hybrid learning models that take advantage of domain lexical and metadata analysis to evaluate the credibility of registered domains and associated websites. Specifically, our hybrid learning models are designed as a fusion of joint features of the lexical domain encoder and RoBERTa-based encoders. To enhance the robustness of the models in practice, we applied a three-stage pipeline with a web-based interface. Throughout an extensive evaluation with realistic datasets, the proposed models demonstrate a substantial improvement over conventional baselines, offering an efficient and scalable solution to monitor abusive web domains in general.