Visual Turing Test: Human-To-Machine Comparisons
摘要
This chapter introduces the Human-to-Machine (H2M) evaluation paradigm, emphasizing the Visual Turing Test (VTT) as a transformative framework for benchmarking machine vision systems against human cognitive abilities. Building on the legacy of the classical Turing Test, the VTT integrates human-centric metrics, focusing on robustness, adaptability, and semantic understanding in dynamic and ambiguous scenarios. Through case studies on classification robustness, visual distortions, and global instance tracking, this chapter reveals the significant performance gaps between humans and machines in handling occlusions, ambiguities, and unpredictable conditions. The findings underscore the need for interdisciplinary approaches, combining cognitive science and machine learning, to advance machine vision toward human-like intelligence and real-world applicability.