The rapid growth in digital imagery across various fields, such as material analysis, medical imaging, and e-commerce, has motivated the research of efficient content-based image retrieval (CBIR) systems (Swain and Ballard in Int J Comput Vis 7:11–32, 1991; Rui et al. in J Vis Commun Image Represent 10:39–62, 1999). However, most of the existing CBIR systems are still based on large annotated datasets, and they are not able to generalize well to new categories, which severely limits their robustness in real applications. This study explores a new strategy that uses few-shot learning to overcome these constraints, with a specific application to texture-based image retrieval (Swain and Ballard in Int J Comput Vis 7:11–32, 1991). We propose a CBIR framework with a ResNet18 backbone for feature extraction, a 128-dimensional embedding layer for compact representation and a triplet loss optimization strategy to ensure the embedding spaces are discriminative (He et al. in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2016:770–778, 2016; Schroff et al. in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2015:815–823). The system is tested on three diverse datasets: KTH-TIPS (Fritz et al. in THE KTH-TIPS Database), MINC (Bell et al. in Material Recognition in the Wild with the Materials in Context Database, 2015), and FMD (Papers with Code—FMD (Materials) Dataset), and achieves state-of-the-art retrieval performances with Precision@5 scores of 1.00, 0.96, and 0.77, respectively. Some techniques, such as data augmentation and feature normalization, are further applied to make the system more robust and generalizable. The study identifies the major existing research gaps in CBIR, including the under-exploration of few-shot learning in texture-specific domains, inefficiency in multimodal feature fusion, and scalability issues when dealing with large-scale datasets. Addressing these gaps, this research presents a scalable, robust, and efficient framework for CBIR that significantly advances the field of texture-based image retrieval.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Texture Content-Based Image Retrieval Using Few-Shot Learning

  • Prabhdrisht Kaur,
  • M. Ravinder

摘要

The rapid growth in digital imagery across various fields, such as material analysis, medical imaging, and e-commerce, has motivated the research of efficient content-based image retrieval (CBIR) systems (Swain and Ballard in Int J Comput Vis 7:11–32, 1991; Rui et al. in J Vis Commun Image Represent 10:39–62, 1999). However, most of the existing CBIR systems are still based on large annotated datasets, and they are not able to generalize well to new categories, which severely limits their robustness in real applications. This study explores a new strategy that uses few-shot learning to overcome these constraints, with a specific application to texture-based image retrieval (Swain and Ballard in Int J Comput Vis 7:11–32, 1991). We propose a CBIR framework with a ResNet18 backbone for feature extraction, a 128-dimensional embedding layer for compact representation and a triplet loss optimization strategy to ensure the embedding spaces are discriminative (He et al. in IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2016:770–778, 2016; Schroff et al. in 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR) 2015:815–823). The system is tested on three diverse datasets: KTH-TIPS (Fritz et al. in THE KTH-TIPS Database), MINC (Bell et al. in Material Recognition in the Wild with the Materials in Context Database, 2015), and FMD (Papers with Code—FMD (Materials) Dataset), and achieves state-of-the-art retrieval performances with Precision@5 scores of 1.00, 0.96, and 0.77, respectively. Some techniques, such as data augmentation and feature normalization, are further applied to make the system more robust and generalizable. The study identifies the major existing research gaps in CBIR, including the under-exploration of few-shot learning in texture-specific domains, inefficiency in multimodal feature fusion, and scalability issues when dealing with large-scale datasets. Addressing these gaps, this research presents a scalable, robust, and efficient framework for CBIR that significantly advances the field of texture-based image retrieval.