Using Natural Language Processing and Machine Learning to Detect Online Grooming Attacks
摘要
A Natural Language Processing solution which incorporates Online Grooming phases has been developed within this paper. This solution coded each phrase within a transcript between a Honeypot profile and an Online Groomer on a chatroom to one of these phases. This was then compared to a human reviewed coding of each of these phrases to check for accuracy. The paper found that this coding identified the Initiation phase (with underaged declaration detection) within 75% of the transcripts with a 3% false positive rate. Most detections were incorrect for the Risk Assessment and Sexual phases. From analysis of this some words/phrase used in the Sexual phase detection had significantly more ‘incorrect’ human reviews than ‘correct’ (21%). It is likely that through filtration of these words/phrases an effective solution could be established, as 38% of these words/phrases had significantly more ‘correct’ human reviews than ‘incorrect’.