This paper explores the background of constructing a railway data lake and presents a comprehensive framework for its architecture. The core components of this framework include three key elements: data ingestion standards, data ingestion strategies, and data storage. The study investigates two methods of data ingestion: physical ingestion and logical ingestion. Physical ingestion is particularly suitable for scenarios involving smaller data volumes and lower real-time requirements, while logical ingestion is designed for handling large datasets with high real-time demands and elevated security levels.The paper clearly defines six standards for data ingestion and outlines the specific workflows associated with each of the two ingestion methods. Establishing a railway data lake plays a crucial role in addressing issues such as data silos and weak data governance. Furthermore, the implementation process of the data lake is designed to integrate various technologies and management strategies, thereby enhancing the overall optimization of railway data management.By creating a data lake, the railway sector can not only improve data accessibility and interoperability but also foster a data-driven culture that promotes informed decision-making. This initiative is expected to streamline operations, enhance data quality, and ultimately contribute to the efficiency and safety of railway systems. The insights gained from this research provide a valuable reference for the future development of data management frameworks within the railway industry and other sectors facing similar challenges.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Research on the Construction and Data Ingestion Methodology of the Railway Data Lake

  • Zou Dan,
  • Wang Peiran,
  • Wu Jiang,
  • Sun Siqi,
  • Guohua Li,
  • Xiaoning Ma

摘要

This paper explores the background of constructing a railway data lake and presents a comprehensive framework for its architecture. The core components of this framework include three key elements: data ingestion standards, data ingestion strategies, and data storage. The study investigates two methods of data ingestion: physical ingestion and logical ingestion. Physical ingestion is particularly suitable for scenarios involving smaller data volumes and lower real-time requirements, while logical ingestion is designed for handling large datasets with high real-time demands and elevated security levels.The paper clearly defines six standards for data ingestion and outlines the specific workflows associated with each of the two ingestion methods. Establishing a railway data lake plays a crucial role in addressing issues such as data silos and weak data governance. Furthermore, the implementation process of the data lake is designed to integrate various technologies and management strategies, thereby enhancing the overall optimization of railway data management.By creating a data lake, the railway sector can not only improve data accessibility and interoperability but also foster a data-driven culture that promotes informed decision-making. This initiative is expected to streamline operations, enhance data quality, and ultimately contribute to the efficiency and safety of railway systems. The insights gained from this research provide a valuable reference for the future development of data management frameworks within the railway industry and other sectors facing similar challenges.