Transformer-Driven Models for Language, Vision, and Multimodality
摘要
In this chapter, we will learn about the modeling and learning techniques that drive multimodal applications. We will focus specifically on the recent advances in transformer-based modeling for natural language understanding, and image understanding, and how these approaches connect for jointly understanding combinations of language and image.