Large Language Models (LLMs) have been widely employed in various text processing tasks. In computer vision, these models have found application in generating captions and text from natural images, as well as in Visual Question Answering (VQA) systems. In the field of medical imaging, there are studies based on text generation proposing automated diagnoses of X-rays, magnetic resonance imaging scans, computed tomography scans, and other modalities. Few initiatives seek to apply and harness the potential of LLMs in medical text generation; they use models with tens of billions of parameters and are thus computationally expensive. This work addresses this gap by evaluating the use of frozen pre-trained models (CXAS U-Net and BioGPT) for chest X-ray report generation. We adapt the BLIP-2 modular architecture where only a cross-modal alignment module must be trained in order to generate text from images. We were able to achieve competitive scores over Clinical Efficacy (CE) metrics compared to some state-of-the-art (SOTA) methods, while obtaining lower scores for Natural Language Generation (NLG) metrics. Our findings suggest that NLG metrics may not serve as suitable proxies for evaluating models in the chest X-ray generation task.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

LLM-Driven Chest X-Ray Report Generation With a Modular, Reduced-Size Architecture

  • Talles Viana Vargas,
  • Helio Pedrini,
  • André Santanchè

摘要

Large Language Models (LLMs) have been widely employed in various text processing tasks. In computer vision, these models have found application in generating captions and text from natural images, as well as in Visual Question Answering (VQA) systems. In the field of medical imaging, there are studies based on text generation proposing automated diagnoses of X-rays, magnetic resonance imaging scans, computed tomography scans, and other modalities. Few initiatives seek to apply and harness the potential of LLMs in medical text generation; they use models with tens of billions of parameters and are thus computationally expensive. This work addresses this gap by evaluating the use of frozen pre-trained models (CXAS U-Net and BioGPT) for chest X-ray report generation. We adapt the BLIP-2 modular architecture where only a cross-modal alignment module must be trained in order to generate text from images. We were able to achieve competitive scores over Clinical Efficacy (CE) metrics compared to some state-of-the-art (SOTA) methods, while obtaining lower scores for Natural Language Generation (NLG) metrics. Our findings suggest that NLG metrics may not serve as suitable proxies for evaluating models in the chest X-ray generation task.