错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Designing a Data Pipeline Architecture for Intelligent Analysis of Streaming Data

  • Iryna Mysiuk,
  • Roman Mysiuk,
  • Roman Shuvar,
  • Volodymyr Yuzevych,
  • Anatolii Pavlenchyk,
  • Volodymyr Dalyk

摘要

The paper describes the process of creating a data pipeline architecture. Based on multiple data sources and processing through intermediate data warehouses and end-to-end business intelligence. This study is a continuation of the previously considered approaches in the context of the generalization of the entire chain of data interaction. The analysis used pre-collected data from social media news sites, which can be considered as one of the most popular streaming data. The information collection consists of automated work with a large amount of information from several web pages using the Selenium tool, and the data warehouse is implemented on the Elasticsearch, Kibana, and Logstash (EKL) technology stack. The analysis and classification of the results are based on working with a machine learning model and data visualization methods. The developed data pipeline architecture can be useful for developing data processing processes, highlighting vulnerabilities and advantages of methodologies of data selection, processing, and analysis.