Cross-modal recipe retrieval, which is a system for finding cooking recipes based on a user-submitted food image or finding food pictures based on recipes, has garnered significant attention in recent years. Many advanced techniques aim to enhance the performance of cross-modal recipe retrieval on general benchmarks. However, leveraging the fine-grained modalities interaction for enhancing multi-modal representation is still limited. We embed a Text-Contextualized Visual Enhancing (TCVE) module into an intermediate layer of the image encoder (e.g., ResNet-50) to enrich the visual encoder. Utilizing the similarity between image local features and intermediate recipe representations, TCVE enhances visual representation by deeper model relearning. We performed comprehensive experiments to validate the proposed approach, FMI (Fine-grained Modalities Interaction for Cross-Modal Recipe Retrieval). As a result, our approach achieves a competitive advantage on the Recipe1M dataset. Specifically, improvements of +3.9 R@1 and +4.9 R@1 are achieved on the 1k and 10k test setup, respectively.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Fine-Grained Modalities Interaction for Cross-Modal Recipe Retrieval

  • Fangying Qu,
  • Yuqing Lu,
  • Zhuo Yao,
  • Fan Zhao

摘要

Cross-modal recipe retrieval, which is a system for finding cooking recipes based on a user-submitted food image or finding food pictures based on recipes, has garnered significant attention in recent years. Many advanced techniques aim to enhance the performance of cross-modal recipe retrieval on general benchmarks. However, leveraging the fine-grained modalities interaction for enhancing multi-modal representation is still limited. We embed a Text-Contextualized Visual Enhancing (TCVE) module into an intermediate layer of the image encoder (e.g., ResNet-50) to enrich the visual encoder. Utilizing the similarity between image local features and intermediate recipe representations, TCVE enhances visual representation by deeper model relearning. We performed comprehensive experiments to validate the proposed approach, FMI (Fine-grained Modalities Interaction for Cross-Modal Recipe Retrieval). As a result, our approach achieves a competitive advantage on the Recipe1M dataset. Specifically, improvements of +3.9 R@1 and +4.9 R@1 are achieved on the 1k and 10k test setup, respectively.