Modern application scenarios in the dynamic industrial, financial, and economic sectors increasingly require quick and agile machine learning solutions. Instead of waiting hours for batch processing systems to deliver results, these systems should ideally adapt and make decisions as soon as new data comes in. As the demand for real-time machine learning solutions using streaming data is steadily increasing, this paper explores a software architecture that efficiently combines the Apache Kafka ecosystem with Microsoft’s machine learning framework ML.NET for reliable data processing and model adaptation. The research addresses the complexity of deciding when to re-train these models in an unbounded data stream context. Various update strategies, including periodic, and performance-based model training, are evaluated for effectiveness under different conditions. The goal of this paper is to propose a completely autonomous machine learning pipeline that is capable of keeping models updated while minimizing computational costs required for re-training and ensuring prediction accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Machine Learning Update Strategies for Real-Time Production Environments

  • Philipp Neuhauser,
  • Philipp Fleck,
  • Sebastian Leitner,
  • Stefan Wagner

摘要

Modern application scenarios in the dynamic industrial, financial, and economic sectors increasingly require quick and agile machine learning solutions. Instead of waiting hours for batch processing systems to deliver results, these systems should ideally adapt and make decisions as soon as new data comes in. As the demand for real-time machine learning solutions using streaming data is steadily increasing, this paper explores a software architecture that efficiently combines the Apache Kafka ecosystem with Microsoft’s machine learning framework ML.NET for reliable data processing and model adaptation. The research addresses the complexity of deciding when to re-train these models in an unbounded data stream context. Various update strategies, including periodic, and performance-based model training, are evaluated for effectiveness under different conditions. The goal of this paper is to propose a completely autonomous machine learning pipeline that is capable of keeping models updated while minimizing computational costs required for re-training and ensuring prediction accuracy.