Deep Learning for Video Captioning
摘要
Among various video understanding tasks, video captioning is a classic and challenging task aimed at learning from human intelligence to establish a connection between vision and language. The goal of video captioning is to automatically generate a natural language to describe the visual content of a video. This can have a significant impact on video indexing and retrieval, applicable in helping visually impaired people. In this chapter, we introduce video captioning with surveys of the state-of-the-art methods. Specifically, we categorize and review modern methods and summarize popular benchmark datasets and evaluation metrics.