<p>In this study, as a proof-of-concept, we aim to initiate the development of <b>Rad</b>iology <b>F</b>oundation <b>M</b>odel, termed as <b>RadFM</b>. We consider three perspectives: dataset construction, model design, and thorough evaluation, concluded as follows: (<i>i</i>), we contribute 4 multimodal datasets with 13M 2D images and 615K 3D scans. When combined with a vast collection of existing datasets, this forms our training dataset, termed as <b>Med</b>ical <b>M</b>ulti-modal <b>D</b>ataset, <b>MedMD</b>. (<i>ii</i>), we propose an architecture that enables to integrate text input with 2D or 3D medical scans, and generates responses for diverse radiologic tasks, including diagnosis, visual question answering, report generation, and rationale diagnosis; (<i>iii</i>), beyond evaluation on 9 existing datasets, we propose a new benchmark, <b>RadBench</b>, comprising three tasks aiming to assess foundation models comprehensively. We conduct both automatic and human evaluations on RadBench. RadFM outperforms former accessible multi-modal foundation models, including GPT-4V. Additionally, we adapt RadFM for diverse public benchmarks, surpassing various existing SOTAs.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Towards generalist foundation model for radiology by leveraging web-scale 2D&3D medical data

  • Chaoyi Wu,
  • Xiaoman Zhang,
  • Ya Zhang,
  • Hui Hui,
  • Yanfeng Wang,
  • Weidi Xie

摘要

In this study, as a proof-of-concept, we aim to initiate the development of Radiology Foundation Model, termed as RadFM. We consider three perspectives: dataset construction, model design, and thorough evaluation, concluded as follows: (i), we contribute 4 multimodal datasets with 13M 2D images and 615K 3D scans. When combined with a vast collection of existing datasets, this forms our training dataset, termed as Medical Multi-modal Dataset, MedMD. (ii), we propose an architecture that enables to integrate text input with 2D or 3D medical scans, and generates responses for diverse radiologic tasks, including diagnosis, visual question answering, report generation, and rationale diagnosis; (iii), beyond evaluation on 9 existing datasets, we propose a new benchmark, RadBench, comprising three tasks aiming to assess foundation models comprehensively. We conduct both automatic and human evaluations on RadBench. RadFM outperforms former accessible multi-modal foundation models, including GPT-4V. Additionally, we adapt RadFM for diverse public benchmarks, surpassing various existing SOTAs.