Medical Visual Question Answering (MedVQA) is crucial for medical data analysis and patient diagnosis, aiding medical practitioners with fast and accurate answers. The recent deep learning model requires a huge amount of data to train; however, collecting large samples and annotations in medical domain is challenging. Therefore, training the model from scratch with small samples easily leads to overfitting. To overcome this problem, pre-trained models can be leveraged and transfer prior knowledge to the medical domain. Efficiently transferring knowledge from pre-trained models with limited data to different domains remains challenging. To address this issue, an efficient convolution-based adapter (EC-Adapter) is introduced, which is versatile and applicable to any pre-trained architecture. The proposed adapter leverages the depth-wise and point-wise convolution operation and add this as parallel layer to the model. The proposed EC-Adapter is simple, lightweight and effective as compared to state-of-the-art low-rank adapters, potentially benefiting large language or vision models. It achieves superior performance while requiring significantly fewer parameters than existing complex methods. In the era of increasingly large and diverse medical datasets, EC-Adapter offers a promising solution to enhance the adaptability and efficiency of pre-trained models in medical applications. The efficacy of the model is demonstrated through extensive experiments and analysis on two publicly available MedVQA datasets: SLAKE and PathVQA.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Efficient Adapter on Pre-trained Visual Feature Reliance in Medical Visual Question Answering

  • Aakansha Mishra,
  • Prateek Keserwani,
  • Vikram N. Rajendiran,
  • Ashok K. Senapati

摘要

Medical Visual Question Answering (MedVQA) is crucial for medical data analysis and patient diagnosis, aiding medical practitioners with fast and accurate answers. The recent deep learning model requires a huge amount of data to train; however, collecting large samples and annotations in medical domain is challenging. Therefore, training the model from scratch with small samples easily leads to overfitting. To overcome this problem, pre-trained models can be leveraged and transfer prior knowledge to the medical domain. Efficiently transferring knowledge from pre-trained models with limited data to different domains remains challenging. To address this issue, an efficient convolution-based adapter (EC-Adapter) is introduced, which is versatile and applicable to any pre-trained architecture. The proposed adapter leverages the depth-wise and point-wise convolution operation and add this as parallel layer to the model. The proposed EC-Adapter is simple, lightweight and effective as compared to state-of-the-art low-rank adapters, potentially benefiting large language or vision models. It achieves superior performance while requiring significantly fewer parameters than existing complex methods. In the era of increasingly large and diverse medical datasets, EC-Adapter offers a promising solution to enhance the adaptability and efficiency of pre-trained models in medical applications. The efficacy of the model is demonstrated through extensive experiments and analysis on two publicly available MedVQA datasets: SLAKE and PathVQA.