Semi-Structured Data Streaming and Anomaly Detection in a Big Data Environment with Spark
摘要
In the rapidly evolving financial sector, efficient data management and processing is critical. With the advent of big data, the ability of organizations to analyze large amounts of information with unprecedented efficiency has increased significantly. This paper focuses on using Apache Spark and Apache Hive within the extensive data framework to process and transform semi-structured data extracted from the bank's core banking system into a structured format. In our newly designed data migration system, anomaly detection is used, and the ELT process and data lake model are implemented throughout the data processing stages. In this study, using a big data environment, a gradient-boosted tree (GBT) classifier and logistic regression are used as machine learning-based methods to detect the anomalies. Logistic regression produced an accuracy of 97.6%, and GBT classifier achieved an accuracy of 91.4%. Leveraging Spark's powerful processing capabilities, we detect anomalies in streaming data, which is critical for fraud detection, risk management and data integrity. The methodology aims to accelerate data analysis processes, improve operational efficiency and enhance customer satisfaction. The new system will provide a more robust data processing framework and reduce the number of processing steps. This project is an innovative effort that contributes to strengthening the bank's data management strategy.