This study focuses on developing and evaluating a medical assistant that utilizes Retrieval-Augmented Generation (RAG) combined with GPT-4 to answer questions related to Congenital Diaphragmatic Hernia (CDH). We aimed to assess the assistant’s response accuracy, identify its limitations, and determine if performance would improve with an expanded knowledge base, also known as its vector database. Initially, the knowledge base contained 230 papers, and the assistant demonstrated mixed results, correctly answering only 19 out of 100 expected CDH questions during the internal evaluation. However, it effectively established a baseline behavior for ignoring out-of-scope questions through prompt engineering, avoiding responses to medically inappropriate inquiries, general pediatric questions, and non-medically related questions. In the external evaluation conducted by healthcare professionals, the assistant correctly answered 67 out of 258 CDH-related questions. Following the expansion of the knowledge base with 770 additional papers, the assistant’s performance significantly improved, achieving an accuracy of 91 out of 100 questions in the internal evaluation. During the external reevaluation after deployment, it answered 224 out of 258 questions correctly. Despite these enhancements, we emphasize the ongoing need to continually update and refine the assistant’s knowledge base to enhance its reliability and accuracy in a healthcare setting.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Development of a CDH-Specific Conversational Assistant Using RAG

  • Vanessa Klotzman,
  • Cristina V. Lopes,
  • John Schomberg,
  • Yostina S. Armanyous,
  • Danielle Linden,
  • Iris Ma,
  • Andreina Giron,
  • Peter Yu,
  • Hira Ahmad,
  • Laura F. Goodman,
  • Mustafa Kabeer,
  • Yigit Guner

摘要

This study focuses on developing and evaluating a medical assistant that utilizes Retrieval-Augmented Generation (RAG) combined with GPT-4 to answer questions related to Congenital Diaphragmatic Hernia (CDH). We aimed to assess the assistant’s response accuracy, identify its limitations, and determine if performance would improve with an expanded knowledge base, also known as its vector database. Initially, the knowledge base contained 230 papers, and the assistant demonstrated mixed results, correctly answering only 19 out of 100 expected CDH questions during the internal evaluation. However, it effectively established a baseline behavior for ignoring out-of-scope questions through prompt engineering, avoiding responses to medically inappropriate inquiries, general pediatric questions, and non-medically related questions. In the external evaluation conducted by healthcare professionals, the assistant correctly answered 67 out of 258 CDH-related questions. Following the expansion of the knowledge base with 770 additional papers, the assistant’s performance significantly improved, achieving an accuracy of 91 out of 100 questions in the internal evaluation. During the external reevaluation after deployment, it answered 224 out of 258 questions correctly. Despite these enhancements, we emphasize the ongoing need to continually update and refine the assistant’s knowledge base to enhance its reliability and accuracy in a healthcare setting.