错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Language Models and Audio-Driven Dynamic Facial Motion Synthesis: A New Paradigm in AI-Driven Interview Training

  • Aakash Garg,
  • Rohan Chaudhury,
  • Mihir Godbole,
  • Jinsil Hwaryoung Seo

摘要

The paper introduces an innovative conversational AI chatbot equipped with a visual avatar, specifically tailored for enhancing skills for nursing interviews in real-time. The chatbot utilizes large language models to simulate realistic interview scenarios, enabling nurses to practice and refine their techniques in a dynamic environment. We present a unique integration of avatar animation into the chatbot system adding a significant layer of realism to these interactions, fostering more natural and engaging interactions. We utilize SadTalker, an AI model that uses 3D motion coefficients to animate still images with audio input. The incorporation of certain rendering techniques, detailed in the paper, helps us reach real-time audio-visual generation. The paper emphasizes the potential of such AI-driven tools in revolutionizing nursing education, particularly in developing critical interview skills for Courtroom Trials. Future developments will aim to explore further and expand the capabilities of this framework, investigating its potential applications across a wider spectrum of educational and professional training contexts.