错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Deep Learning Basics for Video Understanding

  • Zuxuan Wu,
  • Yu-Gang Jiang

摘要

In the last few years, deep learning has revolutionized the field of video understanding, enabling the development of highly accurate and efficient models for tasks such as video action recognition, temporal action localization, and video captioning. This chapter provides an introduction to the key concepts and techniques in deep learning for video understanding, including convolutional neural networks (CNNs), Recurrent Neural Networks (RNNs), and Transformers, which are the backbone of many state-of-the-art video models. Throughout the chapter, we also discuss the benefits as well as limitations of these backbones. By the end of the chapter, readers will have a solid understanding of the basics of deep learning for video understanding and be well-equipped to explore more advanced topics in this exciting field.