错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Cross-modal Recipe Retrieval with Hierarchical Transformers and Pretrained Food Image Encoder

  • Hanyan Qin,
  • Xiankun Zhang,
  • Chen Song

摘要

Social media platforms have seen a surge in sharing food experiences and recipes. To manage the vast amount of data generated by this trend, a cross-modal retrieval system has been developed to retrieve recipes and images. Existing literature has mainly focused on textual aspects of recipes, except for cross-modal pre-trained models, disregarding the features of food images themselves. To address this limitation, we comprehensively analyzed the characteristics of food images and proposed a new cross-modal retrieval framework. Our approach uses a pre-trained network based on food images as the image encoder and a hierarchical Transformer as the text encoder. Our research showed that food images have unique features such as color, texture, and shape that can be utilized to enhance cross-modal retrieval. We also discovered that the current state-of-the-art cross-modal retrieval models that rely solely on textual information are limited in their capacity to retrieve images. Therefore, we developed a new cross-modal retrieval framework that combines both textual and visual information. The image encoder, based on a pre-trained network, extracts visual features from food images, while the text encoder, based on a hierarchical Transformer, extracts textual features from recipes. Our experiments demonstrated that our enhanced dual encoders significantly outperformed the existing baseline models on the dataset. Our proposed framework, which incorporates both textual and visual information, can improve the accuracy of cross-modal retrieval systems and enhance the user experience in searching for food-related information.