Abstract <p>The paper explores the possibility of applying the idea of knowledge distillation to the boosting method. The rationale for this approach is that, in many cases, the best forecast quality is achieved in ensembles using trees of excess depth. In these cases, it may be worthwhile to train an ensemble of shallower trees using a deeper model as a “teacher.” This makes it possible, in particular, to assess the real “depth” of dependences between variables in a problem, as well as to obtain more visual visualizations of solutions. The study also provides material for understanding the mechanisms of the effectiveness of the knowledge distillation procedure.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Distillation of Knowledge in Boosting Models

  • V. M. Nedel’ko

摘要

Abstract

The paper explores the possibility of applying the idea of knowledge distillation to the boosting method. The rationale for this approach is that, in many cases, the best forecast quality is achieved in ensembles using trees of excess depth. In these cases, it may be worthwhile to train an ensemble of shallower trees using a deeper model as a “teacher.” This makes it possible, in particular, to assess the real “depth” of dependences between variables in a problem, as well as to obtain more visual visualizations of solutions. The study also provides material for understanding the mechanisms of the effectiveness of the knowledge distillation procedure.