Data Pipeline Approaches in Serverless Computing
摘要
This paper provides a comprehensive overview of serverless approaches in the design and implementation of data pipelines for cloud-based environments. Serverless computing, with its event-driven architecture, optimizes infrastructure management, enabling effective scaling and cost optimization (HoseinyFarahabady et al. in Journal of Cloud Infrastructure 21:55–67, 2018; Mahmoudi et al. in Journal of Distributed Systems 14:122–134, 2019). We explore different techniques to enhance the effectiveness, scalability, and flexibility of data pipelines using serverless technologies such as AWS Lambda, Google Cloud Functions, and Azure Functions (Pogiatzis and Samakovitis in Appl Sci 11:191–202, 2021; McGrath, T., Brenner, K. Incorporating BaaS elements within serverless data processing workflows. In: International Workshop on Cloud Computing, IEEE, New York, pp. 203–215, 2017). The paper categorizes these methods into ingestion, processing, and storage stages, highlighting their benefits and limitations (Son and Boyd in Journal of Cloud Solutions 22:50–63, 2021; McDermott and Tan in Edge Computing Review 11:150–160, 2020). Eventually, we discuss recent advancements and future trends in serverless data pipeline infrastructures (Zhang, Y., Zhou, L. “Real-time optimization in serverless IoT applications.” In: 15th International Conference on IoT Systems, Elsevier, Amsterdam, pp. 78–85, 2020.; Rahman, M. M., Rausch, T. “Orchestrating containers for data-heavy pipelines in serverless environments.” In: Cloud Computing Innovations, Springer, Heidelberg, pp. 101–113, 2019.).