Retrieval-Augmented Generation (RAG) is a method used to optimize the output of large language models (LLMs). This study investigates the feasibility of using an LLM within a RAG framework to generate recommendations for Traditional Chinese Medicine (TCM) formulations. The study employs the mixtral-8x7b model as the LLM within the RAG architecture, utilizing clinical records from outpatient TCM visits as external data sources for generating TCM formulation recommendations. The recommendations from the RAG-based LLM are compared with those generated by the ChatGPT 3.5 model, evaluating their consistency with actual clinical prescriptions. Results indicate that the RAG-based LLM achieved an average score of 74, demonstrating a high level of alignment with clinical prescriptions across the cases studied. In contrast, the ChatGPT 3.5 model only achieved an average score of 25, primarily due to inconsistencies in the generated recommendations, which rendered them clinically unusable. The study concludes that while the RAG-based LLM shows potential in generating TCM formulation recommendations, there remains a need for improvement in the model’s accuracy.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Pilot Study of Retrieval-Augmented Generation Model in Recommending Traditional Chinese Medicine Formulations

  • Ya-Chuan Chan,
  • Po-Yu Huang,
  • Zhi-Liang Chen,
  • Chih-Nung Wang,
  • Wen-Chen Lin,
  • Jung-Peng Chiu,
  • Yi-Chun Chiu,
  • Yang-Hsien Lin,
  • Eddie T. C. Huang,
  • Simon See,
  • Kang-Ping Lin

摘要

Retrieval-Augmented Generation (RAG) is a method used to optimize the output of large language models (LLMs). This study investigates the feasibility of using an LLM within a RAG framework to generate recommendations for Traditional Chinese Medicine (TCM) formulations. The study employs the mixtral-8x7b model as the LLM within the RAG architecture, utilizing clinical records from outpatient TCM visits as external data sources for generating TCM formulation recommendations. The recommendations from the RAG-based LLM are compared with those generated by the ChatGPT 3.5 model, evaluating their consistency with actual clinical prescriptions. Results indicate that the RAG-based LLM achieved an average score of 74, demonstrating a high level of alignment with clinical prescriptions across the cases studied. In contrast, the ChatGPT 3.5 model only achieved an average score of 25, primarily due to inconsistencies in the generated recommendations, which rendered them clinically unusable. The study concludes that while the RAG-based LLM shows potential in generating TCM formulation recommendations, there remains a need for improvement in the model’s accuracy.