This paper presents a novel approach to extract relations from materials science texts by leveraging domain-specific databases. We introduce a method that combines sequence-to-sequence (seq2seq) language models with knowledge graph embeddings derived from the Materials Project database. Our approach first constructs a comprehensive materials knowledge graph incorporating various properties such as element composition, magnetic ordering, and related materials. We then train knowledge graph embeddings using the translation-based methods. The resulting embeddings are integrated into a seq2seq-based relation extraction model through special knowledge graph tokens. When evaluated on the Materials Science Procedural Text Corpus, our method achieves state-of-the-art performance with a micro-averaged F1-score of 61.87%, representing a 2.63-point improvement over the baseline Flan T5-large model. This work demonstrates the effectiveness of incorporating domain-specific database information for enhancing relation extraction from materials science literature.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Leveraging Domain-Specific Databases for Seq2Seq-Based Relation Extraction from Materials Science Texts

  • Masaki Asada,
  • Ken Fukuda

摘要

This paper presents a novel approach to extract relations from materials science texts by leveraging domain-specific databases. We introduce a method that combines sequence-to-sequence (seq2seq) language models with knowledge graph embeddings derived from the Materials Project database. Our approach first constructs a comprehensive materials knowledge graph incorporating various properties such as element composition, magnetic ordering, and related materials. We then train knowledge graph embeddings using the translation-based methods. The resulting embeddings are integrated into a seq2seq-based relation extraction model through special knowledge graph tokens. When evaluated on the Materials Science Procedural Text Corpus, our method achieves state-of-the-art performance with a micro-averaged F1-score of 61.87%, representing a 2.63-point improvement over the baseline Flan T5-large model. This work demonstrates the effectiveness of incorporating domain-specific database information for enhancing relation extraction from materials science literature.