This study addresses the challenges of information extraction and complete tumor registry coding for lymphoma [1], a malignancy originating from the lymphatic system. Lymphoma encompasses various subtypes, primarily classified into Hodgkin lymphoma and non-Hodgkin lymphoma, both of which exhibit a high incidence globally. Accurate coding of lymphoma diagnoses is particularly challenging due to the complexity and diversity of its subtypes, each possessing unique clinical features and diagnostic criteria. The research explores the potential of large language models (LLMs) in automating the coding process for lymphoma clinical diagnoses by leveraging extensive medical literature, case reports, and clinical guidelines. By simulating the cognitive processes of human coders, these models can offer personalized coding suggestions, thereby enhancing the accuracy and efficiency of diagnostic coding. Through a comparative analysis of coding results generated by LLMs and human coders, this study aims to evaluate the strengths and limitations of these models, ultimately promoting their application in clinical practice. We employed a pipeline with lymphoma info extraction, GraphRAG, and mutual mapping for accuracy enhancement. Our UniGPT model, trained on bilingual and medical data, was assessed for effectiveness.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Lymphoma Tumor Coding and Information Extraction: A Comparative Analysis of Large Language Model-Based Methods

  • Jian Hou,
  • Fen Yang,
  • Weihua Chen,
  • Yining Wang,
  • Jing Feng,
  • Jin Shi,
  • Jiahui Fan,
  • Jianlin Li,
  • Jinmin Gu,
  • Siyu Lv,
  • Yingying Cen

摘要

This study addresses the challenges of information extraction and complete tumor registry coding for lymphoma [1], a malignancy originating from the lymphatic system. Lymphoma encompasses various subtypes, primarily classified into Hodgkin lymphoma and non-Hodgkin lymphoma, both of which exhibit a high incidence globally. Accurate coding of lymphoma diagnoses is particularly challenging due to the complexity and diversity of its subtypes, each possessing unique clinical features and diagnostic criteria. The research explores the potential of large language models (LLMs) in automating the coding process for lymphoma clinical diagnoses by leveraging extensive medical literature, case reports, and clinical guidelines. By simulating the cognitive processes of human coders, these models can offer personalized coding suggestions, thereby enhancing the accuracy and efficiency of diagnostic coding. Through a comparative analysis of coding results generated by LLMs and human coders, this study aims to evaluate the strengths and limitations of these models, ultimately promoting their application in clinical practice. We employed a pipeline with lymphoma info extraction, GraphRAG, and mutual mapping for accuracy enhancement. Our UniGPT model, trained on bilingual and medical data, was assessed for effectiveness.