<p>Keyword extraction is crucial in natural language processing (NLP) tasks, aiding in information retrieval, document summarization, and content categorization. While many studies have discussed keyword extraction for different languages, the Turkish language presents unique challenges due to its rich morphology, complex syntax, and agglutinative nature. This paper proposes a keyword extraction model for Turkish based on the deep learning model of bidirectional encoder representation transformers (BERT) and NLP. The proposed model has been trained using a novel Turkish dataset specifically collected for this task. The dataset was fetched from over 128,000 theses published in the National Thesis Center of Türkiye. 90% of the dataset used for training the model, and 10% of the dataset used for testing. Our experimental results indicate that the proposed model outperforms similar existing methods highlighting a significant advancement in Turkish text keyword extraction. The performance of the proposed model achieved values of 97.77% F1-score, 97.84% precision, and 97.71% recall.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

BERT-based keyword extraction model for the Turkish language

  • Bilal Babayigit,
  • Hamza Sattuf

摘要

Keyword extraction is crucial in natural language processing (NLP) tasks, aiding in information retrieval, document summarization, and content categorization. While many studies have discussed keyword extraction for different languages, the Turkish language presents unique challenges due to its rich morphology, complex syntax, and agglutinative nature. This paper proposes a keyword extraction model for Turkish based on the deep learning model of bidirectional encoder representation transformers (BERT) and NLP. The proposed model has been trained using a novel Turkish dataset specifically collected for this task. The dataset was fetched from over 128,000 theses published in the National Thesis Center of Türkiye. 90% of the dataset used for training the model, and 10% of the dataset used for testing. Our experimental results indicate that the proposed model outperforms similar existing methods highlighting a significant advancement in Turkish text keyword extraction. The performance of the proposed model achieved values of 97.77% F1-score, 97.84% precision, and 97.71% recall.