In this study, our goal is to create interactive avatar agents that can autonomously animate nuanced facial movements realistically, from both visual and behavioral perspectives. Given high-level inputs about the environment and agent profile, our framework harnesses LLMs to produce a series of detailed text descriptions of the avatar agents’ facial motions. These descriptions are then processed by our task-agnostic driving engine into motion token sequences, which are subsequently converted into continuous motion embeddings that are further consumed by our standalone neural-based renderer to generate the final photorealistic avatar animations. To our knowledge, we are the first to utilize the planning and reasoning ability of LLMs together with neural rendering for generalized non-verbal prediction and photo-realistic rendering of avatar agents.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Disentangling Planning, Driving and Rendering for Photorealistic Avatar Agents

  • Duomin Wang,
  • Bin Dai,
  • Yu Deng,
  • Baoyuan Wang

摘要

In this study, our goal is to create interactive avatar agents that can autonomously animate nuanced facial movements realistically, from both visual and behavioral perspectives. Given high-level inputs about the environment and agent profile, our framework harnesses LLMs to produce a series of detailed text descriptions of the avatar agents’ facial motions. These descriptions are then processed by our task-agnostic driving engine into motion token sequences, which are subsequently converted into continuous motion embeddings that are further consumed by our standalone neural-based renderer to generate the final photorealistic avatar animations. To our knowledge, we are the first to utilize the planning and reasoning ability of LLMs together with neural rendering for generalized non-verbal prediction and photo-realistic rendering of avatar agents.