In today’s digital era, the job market is increasingly plagued by misleading job postings, creating obstacles for both job seekers and legitimate employers. According to the FBI’s Internet Crime Complaint Center (IC3), 16,012 people reported being victims of employment scams in 2020, with losses totaling more than $59 million, which in turn is increasing every. This research initiates an endeavor to improve the identification of such deceptive listings using a cutting-edge machine learning strategy. By harnessing the comprehensive dataset from EMSCAD, we implement sophisticated natural language processing methods to distill the text data into a more refined state. Following this, we utilize advanced feature extraction techniques, such as TF-IDF and Count Vectorization, to convert the text into a numerical format that captures the characteristics of authentic and fraudulent advertisements. To combat the dataset’s inherent skewness, we examine a range of oversampling strategies, from the basic Random Oversampling to the more complex SMOTE and ADASYN techniques. After this we have trained our preprocessed dataset on six machine learning models, namely RF, ETC, SVC, LR, MNB, and GBC. Our process is further refined through hyperparameter optimization using Grid Search CV, which ensures the fine-tuning of our varied machine learning models. Our findings indicate that the integration of SMOTE for oversampling with TF-IDF for vectorization, in conjunction with the GBC, has yielded an accuracy of 98% and AUC-ROC score of 97% across different evaluation metrics. The research underscores the potential of these methods to improve fraud detection in online job markets.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing the Detection of Deceptive Job Advertisements: A Machine Learning Approach with Advanced Oversampling and Vectorization Techniques

  • Prakriti Lohumi,
  • Kushagra Gupta,
  • Varsha Sharma,
  • Rachna Jain

摘要

In today’s digital era, the job market is increasingly plagued by misleading job postings, creating obstacles for both job seekers and legitimate employers. According to the FBI’s Internet Crime Complaint Center (IC3), 16,012 people reported being victims of employment scams in 2020, with losses totaling more than $59 million, which in turn is increasing every. This research initiates an endeavor to improve the identification of such deceptive listings using a cutting-edge machine learning strategy. By harnessing the comprehensive dataset from EMSCAD, we implement sophisticated natural language processing methods to distill the text data into a more refined state. Following this, we utilize advanced feature extraction techniques, such as TF-IDF and Count Vectorization, to convert the text into a numerical format that captures the characteristics of authentic and fraudulent advertisements. To combat the dataset’s inherent skewness, we examine a range of oversampling strategies, from the basic Random Oversampling to the more complex SMOTE and ADASYN techniques. After this we have trained our preprocessed dataset on six machine learning models, namely RF, ETC, SVC, LR, MNB, and GBC. Our process is further refined through hyperparameter optimization using Grid Search CV, which ensures the fine-tuning of our varied machine learning models. Our findings indicate that the integration of SMOTE for oversampling with TF-IDF for vectorization, in conjunction with the GBC, has yielded an accuracy of 98% and AUC-ROC score of 97% across different evaluation metrics. The research underscores the potential of these methods to improve fraud detection in online job markets.