Multi-view Intention Recognition in Face-to-Face Communication
摘要
In this paper, we propose an intention recognition method based on generative dataset. Addressing the lack of intention datasets and the recognition methods in face-to-face communication scenarios, we analyze the motions corresponding to intentions and generate a motion dataset using a diffusion model. We then employ a Transformer-based method to map video to intention. In addition, we introduce a joint intention processing method that effectively handles the differences in motion semantics across different camera views, resulting in more accurate recognition outcomes in the case of multi-view data. Overall, this article summarizes a unified framework from acquisition to recognition.