Evaluating the Feasibility and Acceptability of a GPT-Based Chatbot for Depression Screening: A Mixed-Methods Study
摘要
Depression is a significant challenge to public health, exhibiting escalating prevalence and socio-economic burdens. Traditional screening methods are often hampered by inaccessibility and expense. Large language models (LLMs), like Generative Pre-trained Transformers (GPT), offer a novel solution. This study assesses a GPT-powered chatbot’s effectiveness for mental health interactions and its potential for initial depression screening. Using OpenAI’s GPT-3.5 Turbo, we developed ’HopeBot,’ a chatbot based on the Patient Health Questionnaire-9 (PHQ-9). It was engineered to conduct voice-based interactions for initial depression screening through prompt design and audio processing techniques. 10 simulated diverse personas engaged with HopeBot, and 20 participants analyzed these sessions and provided structured feedback. We used statistical analysis, including the Mann–Whitney U test, to evaluate the chatbot’s performance. The study’s findings indicate that GPT-based chatbots, exemplified by HopeBot, provide cost-effective support for managing depression, garnering substantial user satisfaction with an impressive average rating of 3.8 out of 5. Users highly praised HopeBot’s skills in comprehension, adaptability, and role clarity. However, the study identified challenges in HopeBot’s ability to handle extreme emotional responses. Furthermore, there was a significant gap between user-reported satisfaction and their actual willingness to embrace artificial intelligence (AI)-driven chatbots for personal mental health assessments or recommendations to their surroundings, highlighting widespread reservations about incorporating AI into healthcare practices.