Mobile virtual assistants have come a long way, but they still stumble when faced with complex requests and adapting to different app interfaces. This is largely due to their reliance on pre-defined commands and limited contextual awareness. Our research introduces a novel approach to overcome these limitations by integrating Multimodal Large Language Models (MLLMs) into mobile virtual assistants. By harnessing the power of MLLMs, our proposed virtual assistant can process complex requests and seamlessly interact with diverse app interfaces. Traditional systems struggle with maintaining context and understanding dynamic app interfaces, but our approach tackles this by enabling the assistant to analyze and interpret visual elements within these interfaces. Our methodology centers on a comprehensive learning process where the virtual assistant learns to interact with various applications and understand their functionalities. This is further enhanced by the MLLM’s ability to capture and interpret visual cues, allowing for more precise handling of complex queries and task executions. The source code is available at: https://github.com/ELO-Lab/LLM-MobileVA .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Developing a Mobile Virtual Assistant Using Large Language Models for Task Automation

  • Dinh Minh Hieu Phan,
  • Ngoc Hoang Luong

摘要

Mobile virtual assistants have come a long way, but they still stumble when faced with complex requests and adapting to different app interfaces. This is largely due to their reliance on pre-defined commands and limited contextual awareness. Our research introduces a novel approach to overcome these limitations by integrating Multimodal Large Language Models (MLLMs) into mobile virtual assistants. By harnessing the power of MLLMs, our proposed virtual assistant can process complex requests and seamlessly interact with diverse app interfaces. Traditional systems struggle with maintaining context and understanding dynamic app interfaces, but our approach tackles this by enabling the assistant to analyze and interpret visual elements within these interfaces. Our methodology centers on a comprehensive learning process where the virtual assistant learns to interact with various applications and understand their functionalities. This is further enhanced by the MLLM’s ability to capture and interpret visual cues, allowing for more precise handling of complex queries and task executions. The source code is available at: https://github.com/ELO-Lab/LLM-MobileVA .