CalorieVoL: Integrating Volumetric Context Into Multimodal Large Language Models for Image-Based Calorie Estimation
摘要
Multimodal Large Language Models (MLLMs) can perform various food-related tasks with high quality. Notably, high-performance MLLMs, such as GPT-4V, can even estimate caloric content from food images. However, these MLLMs often struggle to accurately recognize volume information, which often leads to errors in calorie estimation. To address this issue, we propose a new MLLM framework called CalorieVoL, designed to enhance the recognition of volume information in food items. By integrating this framework into MLLMs like GPT-4V, we achieved higher scores in terms of MAE and correlation coefficients on Nutrition5k compared to simple MLLMs. Our experiments also showed that the volume-aware recognition improved responses in scenarios where accurate volume estimation is critical.