<p>Ancient scripts embody valuable historical and cultural heritage, emphasizing the importance of their digital preservation. However, open-set recognition of such scripts remains challenging due to limited samples for novel classes, high morphological similarity among characters, and the inadequacy of traditional closed-set methods to accommodate emerging categories. To address these challenges, particularly the large intra-class variance and small inter-class differences caused by visually similar or variant characters, we propose FFSSC, a novel framework for discovering new classes of ancient texts via feature fusion and semi-supervised clustering. FFSSC consists of three main components: (i) a multi-level feature fusion Vision Transformer (FF-ViT) enhanced with an Isolation Score mechanism for key token selection, (ii) a Self-Supervised Ranking Contrastive Loss (SRC) that improves feature discrimination by training the model to group similar characters together while separating dissimilar ones, and (iii) a Semi-Supervised Gaussian Mixture Model (Semi-GMM) that clusters features to discover new classes. Experiments on multiple open-set ancient script datasets show that FFSSC achieves superior performance in novel class detection and incremental category expansion, while maintaining stable accuracy on known classes. This work lays a foundation for intelligent recognition and digital preservation of ancient scripts, supporting practical applications of open-set recognition in cultural heritage studies.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

FFSSC: a framework for discovering new classes of ancient texts based on feature fusion and semi-supervised clustering

  • Yueran Wang,
  • Shanxiong Chen,
  • Qiuyue Ruan,
  • Chunming Wu,
  • Yuqi Ma,
  • Fei Deng,
  • Fa Li,
  • Youxin Liao,
  • Jiahao Zhang,
  • Xun Pu

摘要

Ancient scripts embody valuable historical and cultural heritage, emphasizing the importance of their digital preservation. However, open-set recognition of such scripts remains challenging due to limited samples for novel classes, high morphological similarity among characters, and the inadequacy of traditional closed-set methods to accommodate emerging categories. To address these challenges, particularly the large intra-class variance and small inter-class differences caused by visually similar or variant characters, we propose FFSSC, a novel framework for discovering new classes of ancient texts via feature fusion and semi-supervised clustering. FFSSC consists of three main components: (i) a multi-level feature fusion Vision Transformer (FF-ViT) enhanced with an Isolation Score mechanism for key token selection, (ii) a Self-Supervised Ranking Contrastive Loss (SRC) that improves feature discrimination by training the model to group similar characters together while separating dissimilar ones, and (iii) a Semi-Supervised Gaussian Mixture Model (Semi-GMM) that clusters features to discover new classes. Experiments on multiple open-set ancient script datasets show that FFSSC achieves superior performance in novel class detection and incremental category expansion, while maintaining stable accuracy on known classes. This work lays a foundation for intelligent recognition and digital preservation of ancient scripts, supporting practical applications of open-set recognition in cultural heritage studies.