<p>Distributed learning is a well-established method for estimation tasks over extensively distributed datasets. However, non-randomly stored data can introduce bias into local parameter estimates, leading to significant performance degradation in classical distributed algorithms. In this paper, the authors propose a novel Distributed Quasi-Newton Pilot (DQNP) method for distributed learning with non-randomly distributed data. The proposed approach accommodates both randomly and non-randomly distributed data settings and imposes no constraints on the uniformity of local sample sizes. Additionally, it avoids the need to transfer the Hessian matrix or compute its inversion, thereby greatly reducing computational and communication complexity. The authors theoretically demonstrate that the resulting estimator achieves statistical efficiency under mild conditions. Extensive numerical experiments on synthetic and real-world data validate the theoretical findings and illustrate the effectiveness of the proposed method.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distributed Quasi-Newton Algorithm for Non-Randomly Stored Data

  • Xirui Liu,
  • Mixia Wu,
  • Bangshu Liu

摘要

Distributed learning is a well-established method for estimation tasks over extensively distributed datasets. However, non-randomly stored data can introduce bias into local parameter estimates, leading to significant performance degradation in classical distributed algorithms. In this paper, the authors propose a novel Distributed Quasi-Newton Pilot (DQNP) method for distributed learning with non-randomly distributed data. The proposed approach accommodates both randomly and non-randomly distributed data settings and imposes no constraints on the uniformity of local sample sizes. Additionally, it avoids the need to transfer the Hessian matrix or compute its inversion, thereby greatly reducing computational and communication complexity. The authors theoretically demonstrate that the resulting estimator achieves statistical efficiency under mild conditions. Extensive numerical experiments on synthetic and real-world data validate the theoretical findings and illustrate the effectiveness of the proposed method.