The Design of an Embodied Intelligence Interaction System for the Blind or Low Vision Population Based on Large Multimodal Models
摘要
This research introduces a groundbreaking theoretical framework for assistive technologies through the integration of Embodied Intelligence (EI) with Large Multimodal Models (LMMs), specifically designed to address the complex challenges faced by individuals with blindness or low vision (BLV). The study establishes a novel paradigm that redefines human-environment interaction through three foundational principles: 1) Cognitive embodiment through distributed sensory processing, 2) Anticipatory reasoning via spatiotemporal attention mechanisms, 3) Ethical co-adaptation between users and intelligent systems. We propose a human-centered architecture that synergizes multimodal perception with LMM-based cognitive modeling, enabling contextually aware environmental engagement that transcends conventional reactive assistance paradigms.