错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Multimodal Learning

  • Blaž Škrlj

摘要

In this chapter, we proceed with the notion of multimodal machine learning. As seen in the previous chapter, machine learning approaches that enable learning from different modalities, from texts and graphs to images and temporal signals such as sound-based data and time series, have been developed. Some of the methods discussed already considered a multimodal scenario; for example, transcribing speech to text considers sound modality as the input while outputting text. The remainder of this book focuses on this branch of approaches. One of the reasons multimodal machine learning is becoming a regular research topic on many top-tier conferences is the increasing amount of readily available data. Nowadays, benchmark data sets for, e.g., multimodal classification and regression, are freely available, pushing the boundaries of systems’ capabilities each year. This chapter is organized as follows. In Sect. 6.1, we offer an overview of the existing state-of-the-art and a brief history of multimodal machine learning. We proceed with an overview of multimodal machine learning.