E-commerce platforms frequently update food images to improve user experience, but manual labeling is costly and time-consuming. This study presents an automated labeling system combining Content-Based Image Retrieval (CBIR) and Vision Transformer (ViT) for feature extraction. ViT-Large (Patch 16) achieved superior performance with 0.8334 mAP and 0.6910 Precision, outperforming CNNs like EfficientNet and ResNet. A vector database enables fast retrieval, and a temporary labeling mechanism assigns high-confidence labels from top-k similar images. These labels are later used to retrain ViT models, ensuring continuous improvement. The system reduces labeling costs, maintains data quality, and enhances update efficiency for food-related e-commerce applications.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Automated Food Image Labeling for E-Commerce Websites: Combining Content-Based Image Retrieval and Majority Labeling

  • Khang Nguyen Hoang,
  • Huynh Vu Nhu Nguyen,
  • Hoang Ngoc Tran

摘要

E-commerce platforms frequently update food images to improve user experience, but manual labeling is costly and time-consuming. This study presents an automated labeling system combining Content-Based Image Retrieval (CBIR) and Vision Transformer (ViT) for feature extraction. ViT-Large (Patch 16) achieved superior performance with 0.8334 mAP and 0.6910 Precision, outperforming CNNs like EfficientNet and ResNet. A vector database enables fast retrieval, and a temporary labeling mechanism assigns high-confidence labels from top-k similar images. These labels are later used to retrain ViT models, ensuring continuous improvement. The system reduces labeling costs, maintains data quality, and enhances update efficiency for food-related e-commerce applications.