More Realistic Data Environment
摘要
This chapter delves into the paradigm shift from model-centric to data-centric AI, underscoring the pivotal role of realistic datasets in advancing dynamic vision tasks. Early benchmarks like OTB50 and ALOV++ provided foundational methodologies but lacked the complexity to emulate real-world conditions. Modern datasets such as LaSOT, GOT-10k, and VideoCube address this gap, incorporating spatiotemporal coherence, multimodal integration, and real-world scenario modeling. The introduction of SOTVerse, with its modular 3E paradigm—Environment, Evaluation, and Executor—offers a dynamic framework to assess model robustness under challenging conditions. Additionally, DTVLT redefines visual-language tracking by leveraging Large Language Models (LLMs) for generating multi-granular annotations, enhancing both semantic richness and temporal granularity. These benchmarks exemplify the transformative potential of data-centric AI in fostering robust, generalizable, and adaptive systems, aligning algorithmic capabilities with practical demands.