Multi-modal multi-task federated foundation models for next-generation extended reality systems: towards privacy-preserving distributed intelligence in AR/VR/MR
摘要
Extended reality (XR) systems, encompassing virtual reality (VR), augmented reality (AR), and mixed reality (MR), offer a transformative interface for immersive, multi-modal, and embodied human-computer interaction. In this paper, we envision that multi-modal multi-task (M3T) federated foundation models (FedFMs) can offer transformative capabilities for XR systems through integrating the representational strength of M3T foundation models (FMs) with the privacy-preserving and personalized model training principles of federated learning (FL). To this end, we first present the modular architecture of FedFMs, which entails different coordination paradigms for model training and aggregations. Afterward, we codify XR challenges that affect the implementation of FedFMs under the SHIFT dimensions: (1)