This paper explores the limitations of traditional prototyping methods for evaluating generative AI products and proposes live prototyping as a solution. Recognizing the critical role of Large Language Model (LLM) output quality in user experience, the authors argue that static or clickable prototypes fail to capture the dynamic and personalized nature of GenAI interactions. To address this, the authors developed live prototypes connected to real LLMs and employed them in user studies. The findings demonstrate that live prototypes enable users to explore their own use cases, leading to more accurate and relevant feedback. Additionally, live prototypes provide valuable insights into users’ perceptions of LLM output quality, facilitating improved model evaluation and stakeholder communication. This paper discusses the benefits of live prototypes for empathy building and aligning users’ expectations with product capabilities. The paper acknowledges the limitations of this approach, including the technical expertise required for development and maintenance. This work contributes to a growing understanding of effective evaluation methodologies for GenAI products in the field of human-computer interaction.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Live Prototyping for Evaluating Generative AI: Exploring the Potential and the Pitfalls

  • Yinni Guo,
  • Jesse Sliter,
  • Cameron Oelsen

摘要

This paper explores the limitations of traditional prototyping methods for evaluating generative AI products and proposes live prototyping as a solution. Recognizing the critical role of Large Language Model (LLM) output quality in user experience, the authors argue that static or clickable prototypes fail to capture the dynamic and personalized nature of GenAI interactions. To address this, the authors developed live prototypes connected to real LLMs and employed them in user studies. The findings demonstrate that live prototypes enable users to explore their own use cases, leading to more accurate and relevant feedback. Additionally, live prototypes provide valuable insights into users’ perceptions of LLM output quality, facilitating improved model evaluation and stakeholder communication. This paper discusses the benefits of live prototypes for empathy building and aligning users’ expectations with product capabilities. The paper acknowledges the limitations of this approach, including the technical expertise required for development and maintenance. This work contributes to a growing understanding of effective evaluation methodologies for GenAI products in the field of human-computer interaction.