This article explores the implementation and management of automated Data Pipeline (DP) in the Attorney General’s Office (AGU), a public institution in Brazil, resulting in process simplification, increased efficiency, and best practices in data management used for big data. Challenges such as integrating heterogeneous data, governance, updates, and data transformations were identified. As a result of modifying the current infrastructure and processes supported by best practices that significantly contributed to scalability and operational efficiency, the results are demonstrated with better performance. The results also show a significant reduction in the process’s complexity and improvements in management efficiency, operational scalability, and data governance, where information becomes available with fewer efforts from the data engineering team.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Pipelines Implementation and Management for Data Engineering: A Case Study Applied to the Public Sector

  • Elon Oliveira Albuquerque,
  • Wesley Gongora de Almeida,
  • Bruno Justino Garcia Praciano,
  • Márcio Bastos de Medeiros,
  • Fábio Lúcio Lopes de Mendonça,
  • Robson de Oliveira Albuquerque

摘要

This article explores the implementation and management of automated Data Pipeline (DP) in the Attorney General’s Office (AGU), a public institution in Brazil, resulting in process simplification, increased efficiency, and best practices in data management used for big data. Challenges such as integrating heterogeneous data, governance, updates, and data transformations were identified. As a result of modifying the current infrastructure and processes supported by best practices that significantly contributed to scalability and operational efficiency, the results are demonstrated with better performance. The results also show a significant reduction in the process’s complexity and improvements in management efficiency, operational scalability, and data governance, where information becomes available with fewer efforts from the data engineering team.