Introduction
摘要
In this chapter, we begin by introducing the basic concepts and technical advancements for Video Grounding. We highlight its importance within the research community, discuss its relationships with other Vision-Language learning tasks, and share our insights on the potential extensions towards more generalized settings.