In the context of data-intensive society, cross-modal data streams are generated, collected and processed at an exponential rate. These multimodal data originating from different perceptual terminals not only present differentiated statistical distribution characteristics, but also pose a serious challenge to the parameter isomorphism assumption of the traditional federated learning framework. The traditional federated learning paradigm rigidly requires that the terminals adopt a unified neural network architecture, and when encountering modal heterogeneous features, the global model generalization effectiveness will face a double attenuation of convergence efficiency and prediction accuracy. The multi-client architecture in federated learning makes the emergence of multimodal data in the model unavoidable, and the traditional federated learning based on parameter aggregation performs poorly in the current application scenario. We propose a knowledge distillation framework for federated learning. The framework is compatible with unimodal and multimodal client participation, supports local model heterogeneous design, and effectively overcomes the performance limitations of traditional methods in scenarios of data modality differences and model architecture diversity. The core approach is to design a weighted aggregation mechanism based on modal complementarity, which quantifies the information contribution of different modalities by calculating the cosine similarity between client modal feature vectors, and then adaptively assigns aggregation weights to the high correlation between high-resolution images and textual semantic features of the visual modality, replacing the problem of inhibition of the key modalities by traditional average aggregation. Experiments show that this strategy improves the generalization performance of local models in heterogeneous data environments by guiding cross-modal knowledge fusion between clients.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distillation of Knowledge for Federated Learning Based on Multimodal Fusion

  • Feiyang Wei,
  • Yifan Liu,
  • Fan Feng,
  • Yi Liu,
  • Zhenpeng Liu

摘要

In the context of data-intensive society, cross-modal data streams are generated, collected and processed at an exponential rate. These multimodal data originating from different perceptual terminals not only present differentiated statistical distribution characteristics, but also pose a serious challenge to the parameter isomorphism assumption of the traditional federated learning framework. The traditional federated learning paradigm rigidly requires that the terminals adopt a unified neural network architecture, and when encountering modal heterogeneous features, the global model generalization effectiveness will face a double attenuation of convergence efficiency and prediction accuracy. The multi-client architecture in federated learning makes the emergence of multimodal data in the model unavoidable, and the traditional federated learning based on parameter aggregation performs poorly in the current application scenario. We propose a knowledge distillation framework for federated learning. The framework is compatible with unimodal and multimodal client participation, supports local model heterogeneous design, and effectively overcomes the performance limitations of traditional methods in scenarios of data modality differences and model architecture diversity. The core approach is to design a weighted aggregation mechanism based on modal complementarity, which quantifies the information contribution of different modalities by calculating the cosine similarity between client modal feature vectors, and then adaptively assigns aggregation weights to the high correlation between high-resolution images and textual semantic features of the visual modality, replacing the problem of inhibition of the key modalities by traditional average aggregation. Experiments show that this strategy improves the generalization performance of local models in heterogeneous data environments by guiding cross-modal knowledge fusion between clients.