Overview of the Lymphoma Information Extraction and Automatic Coding Evaluation Task in CHIP 2024
摘要
Lymphoma diagnostic coding is vital for streamlining clinical documentation and supporting tumor registries. This paper introduces the “Lymphoma Information Extraction and Automatic Coding Task,” organized at CHIP 2024, to evaluate the performance of large language models (LLMs) in generating tumor registry codes based on ICD-10 and ICD-O-3 standards. The dataset, comprising 162 clinical case reports, was divided into training set, test set A and test set B. The task assessed accuracy across three coding types, with leaderboard evaluations highlighting innovative approaches. Accuracy is used as evaluation metric. Eighteen participating teams submitted their solutions, with the top-ranked team achieving a score of 0.9296. The top-ranked teams employed diverse methodologies, including LLMs fine-tuning, structured pipelines, Chain of Thought reasoning, and retrieval-augmented generation frameworks, achieving impressive results. This paper provides insights into the dataset, evaluation metrics, and top-performing solutions, offering a foundation for advancing automated medical coding with LLMs.