With the rapid growth of academic literature, traditional single-label classification methods struggle to meet hierarchical and multidisciplinary classification needs. Existing hierarchical text classification methods face limitations in modeling label structures, capturing implicit semantic relationships, and addressing class imbalance. To tackle these challenges, this paper proposes HPMG-ATC, a hierarchical multi-label classification model that integrates Graph Attention Networks (GAT), adaptive hierarchical Mixup, and hierarchy-aware prompt templates. GAT models label dependencies explicitly, while Mixup generates diverse samples based on hierarchical relations to enhance semantic learning. A zero-boundary multi-label cross-entropy loss further improves hierarchical consistency. Experiments on the WOS and FoRC4CL datasets show that HPMG-ATC achieves Micro-F1 and Macro-F1 scores of 87.41% and 82.23%, respectively, significantly outperforming baseline models. Ablation and few-shot experiments confirm the effectiveness of each module and the model’s robustness in low-resource scenarios. This study provides an efficient solution for academic literature classification and contributes new insights to hierarchical multi-label classification research.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

A Hierarchical Prompt-Enhanced Mix-Up Model with Graph Attention Networks for Academic Text Classification

  • Weixing Yuan,
  • Fanjun Meng,
  • Yan Gou,
  • Yingqi Wang,
  • Xingjian Xu

摘要

With the rapid growth of academic literature, traditional single-label classification methods struggle to meet hierarchical and multidisciplinary classification needs. Existing hierarchical text classification methods face limitations in modeling label structures, capturing implicit semantic relationships, and addressing class imbalance. To tackle these challenges, this paper proposes HPMG-ATC, a hierarchical multi-label classification model that integrates Graph Attention Networks (GAT), adaptive hierarchical Mixup, and hierarchy-aware prompt templates. GAT models label dependencies explicitly, while Mixup generates diverse samples based on hierarchical relations to enhance semantic learning. A zero-boundary multi-label cross-entropy loss further improves hierarchical consistency. Experiments on the WOS and FoRC4CL datasets show that HPMG-ATC achieves Micro-F1 and Macro-F1 scores of 87.41% and 82.23%, respectively, significantly outperforming baseline models. Ablation and few-shot experiments confirm the effectiveness of each module and the model’s robustness in low-resource scenarios. This study provides an efficient solution for academic literature classification and contributes new insights to hierarchical multi-label classification research.