Recent advances in Large Language Models (LLMs) have shown remarkable capabilities in various Natural Language Processing (NLP) tasks. However, their performance in specialized domains, particularly healthcare, is still limited by the lack of domain expertise and output reliability. In this paper, we present a retrieval-augmented framework for the CHIP 2024 Diagnostic Consistency Task in Typical Medical Case Records. Our approach addresses these limitations by combining the generative power of LLMs with the retrieval of relevant medical knowledge. Specifically, we introduce a hybrid retrieval mechanism that integrates traditional BM25 with dense retrieval methods, along with a context compression strategy and self-consistency verification module. This framework enables more reliable and accurate diagnostic predictions by leveraging both contextual similarities and domain-specific knowledge. Our method achieves an average F1 score of 0.9517 on the test set, demonstrating the effectiveness of our approach for specialized medical diagnosis tasks.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reliable Typical Case Diagnosis via Optimized Retrieval-Augmented Generation Techniques

  • Kaiyuan Zhang,
  • Bo Wang,
  • Changsen Yuan,
  • Chong Feng,
  • Ge Shi

摘要

Recent advances in Large Language Models (LLMs) have shown remarkable capabilities in various Natural Language Processing (NLP) tasks. However, their performance in specialized domains, particularly healthcare, is still limited by the lack of domain expertise and output reliability. In this paper, we present a retrieval-augmented framework for the CHIP 2024 Diagnostic Consistency Task in Typical Medical Case Records. Our approach addresses these limitations by combining the generative power of LLMs with the retrieval of relevant medical knowledge. Specifically, we introduce a hybrid retrieval mechanism that integrates traditional BM25 with dense retrieval methods, along with a context compression strategy and self-consistency verification module. This framework enables more reliable and accurate diagnostic predictions by leveraging both contextual similarities and domain-specific knowledge. Our method achieves an average F1 score of 0.9517 on the test set, demonstrating the effectiveness of our approach for specialized medical diagnosis tasks.