Currently, text message fraud, which is commonly known as “smoothing”, has become an increasingly serious problem, and the lack of adequate mechanisms to detect fraudulent content in messages has led many citizens to become victims of scams when receiving such messages. Therefore, the objective of this study was to detect fraud in text messages using Recurrent Neural Networks (RNN). The applied methodology was based on 5 phases: Dataset creation; Data preparation; Preprocessing; Feature extraction (Word2Vec and FastText); Training model (RNN, LSTM, GRU, CNN-RNN, CNN-LSTM and CNN-GRU) and Testing model (Accuracy, Precision, Recall, F1-Score and AUC). The top result was obtained with the model with a combination of the 3-layer RNN network (200 (relu), 160 (relu), and 1 (sigmoid)) and the 300-dimensional FastText-NewL embedding: with 85.62% Accuracy, 84.49% Precision, 88.77% Recall, 86.57% F1-score, and 93.14% AUC. The results demonstrate a robust classification model for fraudulent text messages.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Proposal of a Predictive Model for Fraud Detection in Text Messages Using Recurrent Neural Networks

  • Yair Andrey Salinas Bolaños,
  • Wilfredo Ticona

摘要

Currently, text message fraud, which is commonly known as “smoothing”, has become an increasingly serious problem, and the lack of adequate mechanisms to detect fraudulent content in messages has led many citizens to become victims of scams when receiving such messages. Therefore, the objective of this study was to detect fraud in text messages using Recurrent Neural Networks (RNN). The applied methodology was based on 5 phases: Dataset creation; Data preparation; Preprocessing; Feature extraction (Word2Vec and FastText); Training model (RNN, LSTM, GRU, CNN-RNN, CNN-LSTM and CNN-GRU) and Testing model (Accuracy, Precision, Recall, F1-Score and AUC). The top result was obtained with the model with a combination of the 3-layer RNN network (200 (relu), 160 (relu), and 1 (sigmoid)) and the 300-dimensional FastText-NewL embedding: with 85.62% Accuracy, 84.49% Precision, 88.77% Recall, 86.57% F1-score, and 93.14% AUC. The results demonstrate a robust classification model for fraudulent text messages.