Latent Challenges of Multimodal Deep Learning Models: Taxonomy and Survey
摘要
In the evolution of artificial intelligence technology, multimodal deep learning models are a crucial and evolutionary step. Algorithms will soon be widely employed to process multimodal information in every domain. However, researchers often require in-depth exploration of methods and models when venturing into a new subdomain of studied approaches and technologies. Various classification features support prior research on the categorization of multimodal deep learning models. Typically, such taxonomies classify these models as early, intermediate, and late fusion models. This article presents an in-depth review of multimodal deep learning models. It proposes a classification of obstacles that are rarely addressed in literature, but which hold significant implications for future advancements in the field. These challenges comprise computational complexity, generalizability, scalability, security, and a new data processing philosophy. Various aspects of each obstacle are taken into account. This survey can aid researchers in anticipating possible future issues and generating ideas for extended experiments.