The integration of chatbots and generative AI in e-commerce enhances engagement and efficiency but introduces ethical risks such as bias, toxicity, and inconsistencies. This paper presents an LLM-as-a-Judge framework to evaluate chatbot responses across five key dimensions, derived from interviews with stakeholders of one of the world’s largest online retailers. Our analysis of models like GPT-4o and Prometheus shows that while larger models offer more reliable assessments, optimized prompting enables cost-effective alternatives. Strong metric correlations suggest a streamlined evaluation approach, and log probability-based scoring improves robustness. These findings provide a foundation for deploying ethical AI in e-commerce while balancing accuracy, efficiency, and trustworthiness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Toward Safer and Trustworthy Chatbots in E-Commerce: An LLM-as-a-Judge Approach for Ethical Evaluation

  • Tamás Janusko,
  • Ricardo Bochnia,
  • Moritz Harnisch,
  • Anna-Magdalena Krauß,
  • Daniel Richter,
  • Gunnar Hempel,
  • Steffen Tomschke,
  • Jürgen Anke,
  • Maik Thiele

摘要

The integration of chatbots and generative AI in e-commerce enhances engagement and efficiency but introduces ethical risks such as bias, toxicity, and inconsistencies. This paper presents an LLM-as-a-Judge framework to evaluate chatbot responses across five key dimensions, derived from interviews with stakeholders of one of the world’s largest online retailers. Our analysis of models like GPT-4o and Prometheus shows that while larger models offer more reliable assessments, optimized prompting enables cost-effective alternatives. Strong metric correlations suggest a streamlined evaluation approach, and log probability-based scoring improves robustness. These findings provide a foundation for deploying ethical AI in e-commerce while balancing accuracy, efficiency, and trustworthiness.