RESET: Relational Similarity Extension for V3C1 Video Dataset
摘要
Effective content-based information retrieval (IR) is crucial across multimedia platforms, especially in the realm of videos. Whether navigating a personal home video collection or browsing a vast streaming service like YouTube, users often find that a simple metadata search falls short of meeting their information needs. Achieving a reliable estimation of visual similarity holds paramount significance for various IR applications, such as query-by-example, results clustering, and relevance feedback. While many pre-trained models exist for this purpose, they often mismatch with human-perceived similarity leading to biased retrieval results. Up until now, the practicality of fine-tuning such models has been hindered by the absence of suitable datasets. This paper introduces RESET: RElational Similarity Evaluation dataseT. RESET contains over 17,000 similarity annotations for query-candidate-candidate triples of video keyframes taken from the publicly available V3C1 video collection. RESET addresses both close and distant triplets within the realm of unconstrained V3C1 imagery and two of its compact sub-domains: wedding and diving. Offering fine-grained similarity annotations along with their context, re-annotations by multiple users, and similarity estimations from 30 pre-trained models, RESET serves dual purposes. It facilitates the evaluation of novel visual embedding models w.r.t. similarity preservation and provides a resource for fine-tuning visual embeddings to better align with human-perceived similarity. The dataset is available from https://osf.io/ruh5k .