<p>In this work we investigate an observation made by Kipf and Welling (5th International Conference on Learning Representations, 2017), who suggested that untrained Graph Convolutional Networks (GCNs) can generate meaningful node embeddings. In particular, we investigate the effect of training only a single layer of a GCN or a GAT (Graph Attention Network), while keeping the rest of the layers frozen. We propose a basis on which the effect of the untrained layers and their contribution to the generation of embeddings can be predicted. Moreover, we show that network width influences the dissimilarity of node embeddings produced after the initial node features pass through the untrained part of the model. Additionally, we establish a connection between partially trained GCNs and oversmoothing, showing that they are capable of reducing it. We verify our theoretical results experimentally and show the benefits of using deep networks that resist oversmoothing, in a “cold start” scenario, where there is a lack of feature information for unlabeled nodes.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Partially trained graph convolutional networks resist oversmoothing

  • Dimitrios Kelesis,
  • Dimitris Fotakis,
  • Georgios Paliouras

摘要

In this work we investigate an observation made by Kipf and Welling (5th International Conference on Learning Representations, 2017), who suggested that untrained Graph Convolutional Networks (GCNs) can generate meaningful node embeddings. In particular, we investigate the effect of training only a single layer of a GCN or a GAT (Graph Attention Network), while keeping the rest of the layers frozen. We propose a basis on which the effect of the untrained layers and their contribution to the generation of embeddings can be predicted. Moreover, we show that network width influences the dissimilarity of node embeddings produced after the initial node features pass through the untrained part of the model. Additionally, we establish a connection between partially trained GCNs and oversmoothing, showing that they are capable of reducing it. We verify our theoretical results experimentally and show the benefits of using deep networks that resist oversmoothing, in a “cold start” scenario, where there is a lack of feature information for unlabeled nodes.