Multimodal Transformer Training in Personalized Federated Learning
摘要
In contemporary artificial intelligence paradigms, processing multimodal data effectively is paramount, especially given the diversity of data modalities from visual to auditory signals, text, and sensory information prevalent across numerous domains. Building on the Transformer architecture’s success in natural language processing and computer vision, handling complex multimodal data has become a feasible endeavor. Nonetheless, the deployment of these models faces significant privacy concerns. Personalized federated learning (PFL) emerges as a resolution, merging model personalization with data privacy preservation. This paper presents a sophisticated multimodal Transformer framework augmented by PFL, enabling distributed learning across heterogeneous data while providing customization for individual client needs. Our novel approach substantially elevates model performance and privacy adherence, demonstrating an improvement 15% in accuracy over conventional multimodal learning approaches, thereby marking a leap forward in domain-agnostic, personalized multimodal machine learning.