Fraudulent conduct occurs in a variety of contexts, including ecommerce, hospital assistance, business accounts, and financing. Over $100 billion in annual revenue is earned via deception. Although deception is perilous for businesses, it can be recognized if businesses employ advanced technologies like standard algorithms and machine learning. The 6 million rows of data’s extremely imbalanced allocation of both positive and negative classifications is the greatest technological obstacle to deception forecasting. The apparent anomalies in this data’s characterization present another impediment to its utilization. The purpose of this investigation is to resolve both of these problems by thoroughly exploring and analyzing the data, then selecting an appropriate machine learning technique to handle the distortion. The author’s argument is that by using feature engineering and extreme gradient-boosted decision trees, an ideal solution can be achieved with a significantly improved predictive power of 0.997, as demonstrated by the area under the precision-recall curve. It is important to emphasize that this methodology is acceptable for real-world applications because the conclusions were produced without the data being artificially balanced.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Strategies for Detecting and Analyzing Money-Laundering Activity in Online Social Networks

  • K. Ashwitha,
  • R. Swathi

摘要

Fraudulent conduct occurs in a variety of contexts, including ecommerce, hospital assistance, business accounts, and financing. Over $100 billion in annual revenue is earned via deception. Although deception is perilous for businesses, it can be recognized if businesses employ advanced technologies like standard algorithms and machine learning. The 6 million rows of data’s extremely imbalanced allocation of both positive and negative classifications is the greatest technological obstacle to deception forecasting. The apparent anomalies in this data’s characterization present another impediment to its utilization. The purpose of this investigation is to resolve both of these problems by thoroughly exploring and analyzing the data, then selecting an appropriate machine learning technique to handle the distortion. The author’s argument is that by using feature engineering and extreme gradient-boosted decision trees, an ideal solution can be achieved with a significantly improved predictive power of 0.997, as demonstrated by the area under the precision-recall curve. It is important to emphasize that this methodology is acceptable for real-world applications because the conclusions were produced without the data being artificially balanced.