Nowadays, the ability to process, generate and collect various data in real-time has become more relevant for faster reactions to important events in technical enterprises. The numerous distributed stream processing engines is developed to provide the advance large streaming applications. Apache Kafka and Flink also become more popular in the processing and analyzing of large amount of data. Apache Flink is best suited for data analytics, event-driven and data pipeline applications. Applying Apache Kafka to Flink provides low idle time and low latency, and the most demanding stream processing applications in the world. This paper investigates the effective pipeline architecture for developing data driven analytics. Confirming the experimental results, the system proves to optimize the maximum latency, average latency, and idle time of the pipeline architecture. Additionally, further experiments were performed by implementing different levels of parallelism while streaming data, allowing for a comprehensive analysis of maximum latency and average latency of the pipeline.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Pipeline Architecture with Apache Kafka and Apache Flink: Data-Driven Architecture

  • Chan Min Kyaw,
  • Nwe Nwe Myint Thein

摘要

Nowadays, the ability to process, generate and collect various data in real-time has become more relevant for faster reactions to important events in technical enterprises. The numerous distributed stream processing engines is developed to provide the advance large streaming applications. Apache Kafka and Flink also become more popular in the processing and analyzing of large amount of data. Apache Flink is best suited for data analytics, event-driven and data pipeline applications. Applying Apache Kafka to Flink provides low idle time and low latency, and the most demanding stream processing applications in the world. This paper investigates the effective pipeline architecture for developing data driven analytics. Confirming the experimental results, the system proves to optimize the maximum latency, average latency, and idle time of the pipeline architecture. Additionally, further experiments were performed by implementing different levels of parallelism while streaming data, allowing for a comprehensive analysis of maximum latency and average latency of the pipeline.