错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Practical Robustness

  • Rachid Guerraoui,
  • Nirupam Gupta,
  • Rafael Pinot

摘要

In Chaps. 4 and 5 , we introduced a basic robustification of the distributed mini-batch gradient-descent (DMGD) method by essentially replacing the averaging operation at the server with a robust aggregation rule, and then with pre-aggregation schemes. These modifications to DMGD yield an order-optimal asymptotic training error, when using sufficiently large batch-sizes, despite the presence of some adversarial nodes in the system. The resulting Robust DMGD method, however, might be impractical in some cases, as large batch-sizes inflate significantly the gradient complexity. In this chapter, we introduce an advanced technique for reducing the gradient complexity of Robust DMGD, while preserving the order-optimal nature of the asymptotic training error. Thereby, yielding a practical solution to robustness. The technique consists in using Polyak’s momentum on the local mini-batch gradients in order to diminish the effect of local gradient noise in the training error of the learning algorithm. This provides an order-optimal robustness against adversarial nodes with a reasonable gradient complexity.