错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Reference Classification Using BERT Models to Support Scientific-Document Writing

  • Ryoma Hosokawa,
  • Junji Yamato,
  • Ryuichiro Higashinaka,
  • Genichiro Kikui,
  • Hiroaki Sugiyama

摘要

We are presently developing a document clustering method to group reference papers based on semantic similarity to support the writing of scientific papers by organizing the citations more appropriately. Currently, no dataset can be used for clustering experiments and evaluations for this purpose. In this study, we created two datasets of papers and corresponding references from PMC, an online medical paper archive. Then we performed clustering of the reference papers based on their abstracts using BERT, BioBERT, SciBERT and PubMedBERT. In this case, we input the number of clusters for clustering. Clustering by BERT-based models trained on the similarity of pairs of references was more accurate than clustering by embedding the abstracts of references in each BERT model. Moreover, the trained BERT-based models had a clustering accuracy better or comparable to human experts. In addition, we predicted the number of clusters that used information from the references. The prediction accuracy for the number of clusters was about 40%. Evaluation measures for the clustering results are also discussed.