Oracle bone characters (OBCs) serve as vital resources for the in-depth study of Chinese history and the development of writing systems. The recognition of OBCs holds immense significance in the realm of oracle research. Despite the growing adoption of deep learning techniques for OBC recognition, their widespread implementation has been hindered by challenges such as category imbalance and inaccurate labeling in existing datasets. We construct a radical-level oracle bone character dataset (ROBC) in response to these challenges. To mitigate the issue of inaccurate labeling, we rigorously clean the existing dataset based solely on glyph criteria. Moreover, we pioneer the annotation of radical-level information in the oracle bone character dataset. Through statistical analysis, we demonstrate the efficacy of radical-level annotations in alleviating the class imbalance issue prevalent in existing OBC datasets. In addition, we conduct closed and open set recognition tasks on the ROBC using multiple baseline models and achieve considerable results, demonstrating the versatility and robustness of the ROBC dataset and laying the foundation for future research in the OBC recognition. The ROBC dataset is temporarily available at https://github.com/ycfang-lab/ROBC .

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

ROBC: A Radical-Level Oracle Bone Character Dataset

  • Zhengchen Li,
  • Xintong Li,
  • Kaiwen Qian,
  • Yuchun Fang

摘要

Oracle bone characters (OBCs) serve as vital resources for the in-depth study of Chinese history and the development of writing systems. The recognition of OBCs holds immense significance in the realm of oracle research. Despite the growing adoption of deep learning techniques for OBC recognition, their widespread implementation has been hindered by challenges such as category imbalance and inaccurate labeling in existing datasets. We construct a radical-level oracle bone character dataset (ROBC) in response to these challenges. To mitigate the issue of inaccurate labeling, we rigorously clean the existing dataset based solely on glyph criteria. Moreover, we pioneer the annotation of radical-level information in the oracle bone character dataset. Through statistical analysis, we demonstrate the efficacy of radical-level annotations in alleviating the class imbalance issue prevalent in existing OBC datasets. In addition, we conduct closed and open set recognition tasks on the ROBC using multiple baseline models and achieve considerable results, demonstrating the versatility and robustness of the ROBC dataset and laying the foundation for future research in the OBC recognition. The ROBC dataset is temporarily available at https://github.com/ycfang-lab/ROBC .