<p>AI-based conversational agents (CA) have emerged as popular tools for addressing mental health concerns. While research has primarily focused on usability or outcome measures, there is a scarcity of studies examining the quality of CA-delivered conversations. This article emphasizes the importance of evaluating CA content by drawing parallels with the rigorous training and evaluation process for human therapists. The Thera-Turing Test (TTT) proposes a comprehensive evaluation model that scrutinizes CA conversations independently, reducing potential bias. The TTT recommends that judges rate user conversations with the CA without knowing that AI-CAs delivered those conversations. The components of the TTT include timing, conversation creation, specific vs. full assessment, selection criteria for conversations and the judges, and selecting the appropriate measures. The proposed TTT framework addresses diverse user needs by encompassing safety, common factors, specific factors, cultural awareness, referrals, and readiness for “practice” (or full-scale launch). Alternative approaches to CA evaluations are also explored, and potential critiques are addressed, highlighting the unique challenges of text-based therapy interactions. The Thera-Turing Test offers a promising avenue for ensuring the quality and effectiveness of mental health CAs, contributing to the delivery of high-quality mental health care at scale.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Framework for Evaluating Mental Health Artificial Intelligence-based Conversational Agents

  • Eduardo L. Bunge,
  • Christina Desage

摘要

AI-based conversational agents (CA) have emerged as popular tools for addressing mental health concerns. While research has primarily focused on usability or outcome measures, there is a scarcity of studies examining the quality of CA-delivered conversations. This article emphasizes the importance of evaluating CA content by drawing parallels with the rigorous training and evaluation process for human therapists. The Thera-Turing Test (TTT) proposes a comprehensive evaluation model that scrutinizes CA conversations independently, reducing potential bias. The TTT recommends that judges rate user conversations with the CA without knowing that AI-CAs delivered those conversations. The components of the TTT include timing, conversation creation, specific vs. full assessment, selection criteria for conversations and the judges, and selecting the appropriate measures. The proposed TTT framework addresses diverse user needs by encompassing safety, common factors, specific factors, cultural awareness, referrals, and readiness for “practice” (or full-scale launch). Alternative approaches to CA evaluations are also explored, and potential critiques are addressed, highlighting the unique challenges of text-based therapy interactions. The Thera-Turing Test offers a promising avenue for ensuring the quality and effectiveness of mental health CAs, contributing to the delivery of high-quality mental health care at scale.