Long-tailed hashing is to learn hash functions in unbalanced distribution datasets to represent images as binary hash codes for fast and accurate image retrieval. In contrast to balanced distribution datasets, unbalanced distributions are more common in the real world. However, Existing long-tailed hashing methods only focus on how to better learn from unbalanced datasets to improve performance, without giving good consideration to quantization error, which is very crucial in hash learning. In this paper, we propose a simple but efficient quantization method for long-tailed hashing. Specifically, to address the lack of samples in the tail classes, we take a uniform discrete distribution as the optimal target distribution. We use the Sliced Wasserstein distance as a measure of distribution distance. It makes good use of the discrete nature of hash functions and has low computational complexity. Then we formulate the optimization objective of the quantization error as minimizing the distance between the output of the learned hash function and this objective distribution, which can be added as an additional term of the loss function to existing long-tailed hashing methods. We conduct experiments on two long-tailed datasets, and the results show that our proposed method greatly improves the performance of existing long-tailed hashing methods.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Long-Tailed Hashing with Wasserstein Quantization

  • Zujun Fu,
  • Hanjiang Lai,
  • Yan Pan

摘要

Long-tailed hashing is to learn hash functions in unbalanced distribution datasets to represent images as binary hash codes for fast and accurate image retrieval. In contrast to balanced distribution datasets, unbalanced distributions are more common in the real world. However, Existing long-tailed hashing methods only focus on how to better learn from unbalanced datasets to improve performance, without giving good consideration to quantization error, which is very crucial in hash learning. In this paper, we propose a simple but efficient quantization method for long-tailed hashing. Specifically, to address the lack of samples in the tail classes, we take a uniform discrete distribution as the optimal target distribution. We use the Sliced Wasserstein distance as a measure of distribution distance. It makes good use of the discrete nature of hash functions and has low computational complexity. Then we formulate the optimization objective of the quantization error as minimizing the distance between the output of the learned hash function and this objective distribution, which can be added as an additional term of the loss function to existing long-tailed hashing methods. We conduct experiments on two long-tailed datasets, and the results show that our proposed method greatly improves the performance of existing long-tailed hashing methods.