FedEPA: Enhancing Personalization and Modality Alignment in Multimodal Federated Learning
摘要
Federated Learning (FL) enables decentralized training across multiple participants with privacy preservation. However, most FL systems consider clients hold only unimodal data, limiting their real-world applicability, as institutions often possess multimodal data. Moreover, the label scarcity further constrains the behavior of most FL methods. In this work, we propose FedEPA for multimodal learning. FedEPA employs a personalized local model aggregation strategy, leveraging client-side labeled data to derive personalized weights and mitigate data heterogeneity. Furthermore, it introduces an unsupervised modality alignment strategy effective with limited labeled data, which decomposes features into aligned and context components and uses contrastive learning for cross-modal alignment and intra-modal disentanglement. A multimodal feature fusion strategy creates a joint embedding. Experimental results confirm that FedEPA consistently achieves superior performance over compared FL methods on multimodal classification tasks under limited labeled data conditions.