In the rapidly evolving financial sector, efficient data management and processing is critical. With the advent of big data, the ability of organizations to analyze large amounts of information with unprecedented efficiency has increased significantly. This paper focuses on using Apache Spark and Apache Hive within the extensive data framework to process and transform semi-structured data extracted from the bank's core banking system into a structured format. In our newly designed data migration system, anomaly detection is used, and the ELT process and data lake model are implemented throughout the data processing stages. In this study, using a big data environment, a gradient-boosted tree (GBT) classifier and logistic regression are used as machine learning-based methods to detect the anomalies. Logistic regression produced an accuracy of 97.6%, and GBT classifier achieved an accuracy of 91.4%. Leveraging Spark's powerful processing capabilities, we detect anomalies in streaming data, which is critical for fraud detection, risk management and data integrity. The methodology aims to accelerate data analysis processes, improve operational efficiency and enhance customer satisfaction. The new system will provide a more robust data processing framework and reduce the number of processing steps. This project is an innovative effort that contributes to strengthening the bank's data management strategy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Semi-Structured Data Streaming and Anomaly Detection in a Big Data Environment with Spark

  • Alp Revanbahs,
  • Behiye Alak,
  • Aysegul Senol Calim

摘要

In the rapidly evolving financial sector, efficient data management and processing is critical. With the advent of big data, the ability of organizations to analyze large amounts of information with unprecedented efficiency has increased significantly. This paper focuses on using Apache Spark and Apache Hive within the extensive data framework to process and transform semi-structured data extracted from the bank's core banking system into a structured format. In our newly designed data migration system, anomaly detection is used, and the ELT process and data lake model are implemented throughout the data processing stages. In this study, using a big data environment, a gradient-boosted tree (GBT) classifier and logistic regression are used as machine learning-based methods to detect the anomalies. Logistic regression produced an accuracy of 97.6%, and GBT classifier achieved an accuracy of 91.4%. Leveraging Spark's powerful processing capabilities, we detect anomalies in streaming data, which is critical for fraud detection, risk management and data integrity. The methodology aims to accelerate data analysis processes, improve operational efficiency and enhance customer satisfaction. The new system will provide a more robust data processing framework and reduce the number of processing steps. This project is an innovative effort that contributes to strengthening the bank's data management strategy.