Multimodal Data Fusion Architectures in Audiovisual Speech Recognition
摘要
In the big data era, we are facing a diversity of datasets from different sources in different domains that describe a single life event. These datasets consist of multiple modalities, each of which has a different representation, distribution, scale, and density. Multimodal fusion is the concept of integrating information from multiple modalities in a joint representation with the goal of predicting an outcome through a classification or regression task. In this paper, the strategies of multimodal data fusion were reviewed. The challenges of multimodal data fusion were expressed. A full assessment of the recent studies in multimodal data fusion was presented. In addition, a comparative study of the recent research on this point was performed, and the advantages and disadvantages of each were illustrated. Furthermore, the audiovisual speech recognition task was expressed as a case study of multimodal data fusion techniques. Particularly, this paper can be considered a powerful guide for interested researchers in the fields of multimodal data fusion and audiovisual speech recognition.