Lymphoma Tumor Coding and Information Extraction: A Comparative Analysis of Large Language Model-Based Methods
摘要
This study addresses the challenges of information extraction and complete tumor registry coding for lymphoma [1], a malignancy originating from the lymphatic system. Lymphoma encompasses various subtypes, primarily classified into Hodgkin lymphoma and non-Hodgkin lymphoma, both of which exhibit a high incidence globally. Accurate coding of lymphoma diagnoses is particularly challenging due to the complexity and diversity of its subtypes, each possessing unique clinical features and diagnostic criteria. The research explores the potential of large language models (LLMs) in automating the coding process for lymphoma clinical diagnoses by leveraging extensive medical literature, case reports, and clinical guidelines. By simulating the cognitive processes of human coders, these models can offer personalized coding suggestions, thereby enhancing the accuracy and efficiency of diagnostic coding. Through a comparative analysis of coding results generated by LLMs and human coders, this study aims to evaluate the strengths and limitations of these models, ultimately promoting their application in clinical practice. We employed a pipeline with lymphoma info extraction, GraphRAG, and mutual mapping for accuracy enhancement. Our UniGPT model, trained on bilingual and medical data, was assessed for effectiveness.