Efficient Video Understanding
摘要
Video understanding requires substantially more computing resources compared to their image counterparts due to the additional temporal dimension. As a result, the development of efficient deep video models and training strategies is necessary for practical video understanding applications. In this chapter, we will delve into the design choices for creating compact video understanding models, such as CNNs and Transformers. Furthermore, we will explore training strategies aimed at achieving efficient video understanding, minimizing both time and memory consumption. Finally, we will discuss dynamic inference techniques that adaptively allocate computation resources to different video frames to further accelerate video analysis without sacrificing performance. Through this chapter, we aim to provide a comprehensive overview of efficient deep learning methods for video understanding.