This study introduces KazakhSum, a curated dataset comprising 202k open-access articles from major Kazakh news websites, designed to support abstractive summarization tasks in low-resource languages. Using this dataset, we conducted experiments with various Transformer-based models, including mBART, mT5-small, mT5-base, and GPT-2, and evaluated the quality of generated summaries using standard metrics such as ROUGE and BLEU. The results demonstrate the strong capability of these models in handling text summarization in Kazakh, providing valuable insights into their application in low-resource linguistic contexts. Furthermore, the study explores the practical application of these models in automating Search Engine Optimization (SEO) metadata generation, such as descriptions, keywords, and titles, within Content Management Systems (CMS), laying a foundation for their broader adoption in information retrieval and semantic analysis.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Kazakh Abstractive Summarization: Dataset, Model Evaluation, and Applications in Automated SEO Metadata Generation

  • Mamyr Altaibek,
  • Altanbek Zulhazhav,
  • Gulmira Bekmanova,
  • Banu Yergesh,
  • Assel Omarbekova,
  • Sharipbay Altynbek

摘要

This study introduces KazakhSum, a curated dataset comprising 202k open-access articles from major Kazakh news websites, designed to support abstractive summarization tasks in low-resource languages. Using this dataset, we conducted experiments with various Transformer-based models, including mBART, mT5-small, mT5-base, and GPT-2, and evaluated the quality of generated summaries using standard metrics such as ROUGE and BLEU. The results demonstrate the strong capability of these models in handling text summarization in Kazakh, providing valuable insights into their application in low-resource linguistic contexts. Furthermore, the study explores the practical application of these models in automating Search Engine Optimization (SEO) metadata generation, such as descriptions, keywords, and titles, within Content Management Systems (CMS), laying a foundation for their broader adoption in information retrieval and semantic analysis.