Planning and managing projects, especially those involving machine learning (ML), can be challenging without clear guidance. To address this, various general and specialized ML methodologies have been developed. However, people are still needed to follow these methods: junior team members might not know what to do or in what order, while senior members, though experienced, may skip steps under stress. To make project planning and execution of ML projects more systematic and reliable, we created an automatic interactive chat agent by applying Retrieval-Augmented Generation (RAG) to a state-of-the-art foundational large language model (LLM). This chat agent acts like a “virtual team member”, answering questions about projects and/or a chosen ML methodology, and actioning some limited delegated tasks. We also describe MLProjBench, a novel dataset of independently created questions and model answers, intended to serve as a benchmark for future similar tools aimed at this novel task, and evaluated our chat agent using it. We assess a combination of background documents via RAG to model answers, and we also compare the outputs of different model variants. Our results show that a pre-trained LLM based on neural transformers, equipped with a methodology manual for additional background knowledge via a RAG regime can provide valuable assistance. Among a set of six chat agent variants automatically evaluated against a set of human-authored model answers, the variant that uses MANDALA.ML performed best in our automatic evaluation. We also introduce a new evaluation metric, \(\omega \mathrm {LP\text {-} ROUGE}\) to avoid possible metric gaming that ROUGE suffers from.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Welcome to the Team: Chat Agents for Machine Learning Project Management Support

  • Michael Reiche,
  • Jochen L. Leidner

摘要

Planning and managing projects, especially those involving machine learning (ML), can be challenging without clear guidance. To address this, various general and specialized ML methodologies have been developed. However, people are still needed to follow these methods: junior team members might not know what to do or in what order, while senior members, though experienced, may skip steps under stress. To make project planning and execution of ML projects more systematic and reliable, we created an automatic interactive chat agent by applying Retrieval-Augmented Generation (RAG) to a state-of-the-art foundational large language model (LLM). This chat agent acts like a “virtual team member”, answering questions about projects and/or a chosen ML methodology, and actioning some limited delegated tasks. We also describe MLProjBench, a novel dataset of independently created questions and model answers, intended to serve as a benchmark for future similar tools aimed at this novel task, and evaluated our chat agent using it. We assess a combination of background documents via RAG to model answers, and we also compare the outputs of different model variants. Our results show that a pre-trained LLM based on neural transformers, equipped with a methodology manual for additional background knowledge via a RAG regime can provide valuable assistance. Among a set of six chat agent variants automatically evaluated against a set of human-authored model answers, the variant that uses MANDALA.ML performed best in our automatic evaluation. We also introduce a new evaluation metric, \(\omega \mathrm {LP\text {-} ROUGE}\) to avoid possible metric gaming that ROUGE suffers from.