Robust Aggregation on Federated Distillation in Adversarial Environments via Entropy and Semantic-Aware Logit Filtering
摘要
Federated Distillation (FD), as a parameter-free federated learning paradigm, aggregates client-generated logits on a shared public dataset and demonstrates significant advantages in handling model heterogeneity. However, compared to model or gradient parameter aggregation, the low dimensionality and fixed semantic structure of logits make attacks more covert and harder to be detected by existing robust aggregation methods because of the lack of effective defense mechanisms designed for this paradigm. To address this challenge, we propose a robust federated aggregation method that enhances the robustness of FD by filtering out excessive deviations of client updates based on the entropy of local logits and their cosine distance to the global consensus. Specifically, we first compute the Entropy Alignment Score (EAS) to evaluate the confidence of each client’s output and exclude highly uncertain predictions. We then apply the Cosine Distance Score (CDS) to measure semantic consistency with the global consensus, further identifying anomalous clients. We conduct extensive experiments on the MNIST dataset under various types of adversarial attacks. The experimental results show that our enhanced FD framework achieves an average performance improvement of 6.9% over the original FD method.