Food recognition, as a fundamental task in food computing, plays a crucial role in downstream tasks such as nutritional assessment and dietary management. Recognizing multiple food items in a single image aligns more closely with real dietary environments. However, there is a relative scarcity of publicly available Chinese cuisine multi-food image datasets in this field. Despite the numerous studies focused on single-food recognition, those approaches do not apply to multi-food recognition due to significant intra-category variations and high inter-category similarities among multiple food items present in a single image. To address the first issue, we create a large-scale Chinese cuisine multi-food image dataset, JNU FoodNet162, consisting of 35,365 images across 162 categories. In response to the second question, we propose a Multi-Receptive-field Fusion Network (MRFNet) for multi-food recognition, which captures unique fine-grained features of Chinese food images using the multi-receptive-field pyramid network, fuses feature information from different receptive fields through the detailed and semantic information fusion network and finally adds a small food image prediction head to achieve recognition of small food. The experimental results indicate that MRFNet achieves state-of-the-art mean mAP values on JNU FoodNet162, JNU FoodNet, UEC Food-100, and UEC Food-256.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

MRFNet: A Multi-receptive-Field Fusion Network for Multi-food Recognition

  • Xingda Shang,
  • Jing Chen,
  • Xintong Liu

摘要

Food recognition, as a fundamental task in food computing, plays a crucial role in downstream tasks such as nutritional assessment and dietary management. Recognizing multiple food items in a single image aligns more closely with real dietary environments. However, there is a relative scarcity of publicly available Chinese cuisine multi-food image datasets in this field. Despite the numerous studies focused on single-food recognition, those approaches do not apply to multi-food recognition due to significant intra-category variations and high inter-category similarities among multiple food items present in a single image. To address the first issue, we create a large-scale Chinese cuisine multi-food image dataset, JNU FoodNet162, consisting of 35,365 images across 162 categories. In response to the second question, we propose a Multi-Receptive-field Fusion Network (MRFNet) for multi-food recognition, which captures unique fine-grained features of Chinese food images using the multi-receptive-field pyramid network, fuses feature information from different receptive fields through the detailed and semantic information fusion network and finally adds a small food image prediction head to achieve recognition of small food. The experimental results indicate that MRFNet achieves state-of-the-art mean mAP values on JNU FoodNet162, JNU FoodNet, UEC Food-100, and UEC Food-256.