Identifying Duplicate Hidden Services on the Dark Web
摘要
Cyber-attacks have become increasingly sophisticated, making it challenging to prevent them. Attackers often leverage the Dark Web to coordinate and execute these attacks, exploiting its anonymity. This paper addresses to generate cyber threat intelligence by proposing a novel approach for identifying duplicated hidden services on the Dark Web. Our system collects information using two crawlers: a keyword-based hidden services crawler and a seed-based hidden services crawler. The collected HTML code through crawlers is analyzed to remove banners, and the banner removed index images are analyzed using deep learning techniques. Experimental results indicate that this method is effective in identifying duplicated hidden services.