<p>Online recruitment facilitates the automatic hiring process for recruiters and provides convenience to job seekers via online job platforms. Parallelly, it has given rise to malicious use of such platforms by fraudsters who post fake jobs and steal money and personal information from innocent job seekers. It is difficult to detect fake jobs manually, as these are meticulously crafted to mimic legitimate ones. Previously, various machine learning approaches have employed Bag-of-Words (BoW) and Term Frequency-Inverse Document Frequency (TF-IDF) feature extraction methods for this task. However, these methods are non-contextual and show skewness in results due to imbalances in data distribution. This paper presents Fraud-BERT, a transformer-based contextual framework leveraging Bidirectional Encoder Representations from Transformers (BERT) via transfer learning approach and evaluates it on a highly imbalanced fake job dataset, which is popularly named as Employment Scam Aegean Dataset (EMSCAD). The dataset is available on Kaggle website. The superiority of the proposed method is demonstrated by a comparative analysis with conventional methods. The results of the study conclude that the proposed method is more robust in tackling imbalanced data and, it has significantly out-performed existing state-of-the-art studies with F1 score of 0.93 and 99% accuracy.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fraud-BERT: transformer based context aware online recruitment fraud detection

  • Khushboo Taneja,
  • Jyoti Vashishtha,
  • Saroj Ratnoo

摘要

Online recruitment facilitates the automatic hiring process for recruiters and provides convenience to job seekers via online job platforms. Parallelly, it has given rise to malicious use of such platforms by fraudsters who post fake jobs and steal money and personal information from innocent job seekers. It is difficult to detect fake jobs manually, as these are meticulously crafted to mimic legitimate ones. Previously, various machine learning approaches have employed Bag-of-Words (BoW) and Term Frequency-Inverse Document Frequency (TF-IDF) feature extraction methods for this task. However, these methods are non-contextual and show skewness in results due to imbalances in data distribution. This paper presents Fraud-BERT, a transformer-based contextual framework leveraging Bidirectional Encoder Representations from Transformers (BERT) via transfer learning approach and evaluates it on a highly imbalanced fake job dataset, which is popularly named as Employment Scam Aegean Dataset (EMSCAD). The dataset is available on Kaggle website. The superiority of the proposed method is demonstrated by a comparative analysis with conventional methods. The results of the study conclude that the proposed method is more robust in tackling imbalanced data and, it has significantly out-performed existing state-of-the-art studies with F1 score of 0.93 and 99% accuracy.