Data Pipelines Implementation and Management for Data Engineering: A Case Study Applied to the Public Sector
摘要
This article explores the implementation and management of automated Data Pipeline (DP) in the Attorney General’s Office (AGU), a public institution in Brazil, resulting in process simplification, increased efficiency, and best practices in data management used for big data. Challenges such as integrating heterogeneous data, governance, updates, and data transformations were identified. As a result of modifying the current infrastructure and processes supported by best practices that significantly contributed to scalability and operational efficiency, the results are demonstrated with better performance. The results also show a significant reduction in the process’s complexity and improvements in management efficiency, operational scalability, and data governance, where information becomes available with fewer efforts from the data engineering team.