Multimodal Behavior Modeling
摘要
This chapter explains the complexities of multimodal behavior modeling. While humans recognize social intentions unconsciously, the automation of such recognition involves significant challenges, particularly in spontaneous interactions. We explore the variety of multimodal behavior modeling through two specific examples—ASD severity estimation and social skill estimation—to showcase its potential. Using datasets from clinical populations and machine learning techniques, the chapter demonstrates the potential analytic power of automatic behavior modeling with multimodal features like text, audio, and visual cues. We also discuss expected challenges in this field in the era of LLMs, such as the need for sequential modeling and the limited amount of data. Accordingly, we hope to show the future direction of multimodal behavior modeling.