<p>Sparse neural networks are neural models with pruned weight matrices, which can be both more efficient and more interpretable than dense models. The structures that underlie effective sparse architectures, however, are poorly understood. In this paper, we propose a new technique for sparsification of recurrent neural networks (RNNs), called <i>moduli regularization</i>. Moduli regularization imposes a geometric relationship between neurons in the hidden state of the RNN parameterized by a manifold. We further provide an explicit end-to-end moduli learning mechanism, in which optimal geometry is inferred during training. We verify the effectiveness of our scheme in three settings, testing in navigation, natural language processing, and synthetic long-term recall tasks. While past work has found some evidence of local topology positively affecting network quality, we show that the quality of trained sparse models also heavily depends on the global topological characteristics of the network.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Geometric sparsification in recurrent neural networks

  • Wyatt Mackey,
  • Ioannis Schizas,
  • Jared Deighton,
  • David L. Boothe Jr,
  • Vasileios Maroulas

摘要

Sparse neural networks are neural models with pruned weight matrices, which can be both more efficient and more interpretable than dense models. The structures that underlie effective sparse architectures, however, are poorly understood. In this paper, we propose a new technique for sparsification of recurrent neural networks (RNNs), called moduli regularization. Moduli regularization imposes a geometric relationship between neurons in the hidden state of the RNN parameterized by a manifold. We further provide an explicit end-to-end moduli learning mechanism, in which optimal geometry is inferred during training. We verify the effectiveness of our scheme in three settings, testing in navigation, natural language processing, and synthetic long-term recall tasks. While past work has found some evidence of local topology positively affecting network quality, we show that the quality of trained sparse models also heavily depends on the global topological characteristics of the network.