High-quality datasets are crucial for advancing research in automatic text summarization. At present, summarization models for resource-rich languages like Chinese and English have made significant progress. However, for low-resource languages such as Tibetan, the lack of large-scale publicly available summarization datasets means that related research is still in its early stages. To address this gap, this paper constructs an open Tibetan summarization dataset, TiLTS. By collecting extensive texts from Tibetan news websites and leveraging resources from other languages, it obtains 36,507 (document, summary) pairs. Compared to the only publicly available Tibetan summarization dataset, Ti-SUM, TiLTS has clear advantages in both data size and challenge. This paper also conducts experiments and analyses using several summarization algorithms on this dataset.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

TiLTS:Tibetan Long Text Summarization Dataset

  • Yanrong Hao,
  • Bo Chen,
  • Xiaobing Zhao

摘要

High-quality datasets are crucial for advancing research in automatic text summarization. At present, summarization models for resource-rich languages like Chinese and English have made significant progress. However, for low-resource languages such as Tibetan, the lack of large-scale publicly available summarization datasets means that related research is still in its early stages. To address this gap, this paper constructs an open Tibetan summarization dataset, TiLTS. By collecting extensive texts from Tibetan news websites and leveraging resources from other languages, it obtains 36,507 (document, summary) pairs. Compared to the only publicly available Tibetan summarization dataset, Ti-SUM, TiLTS has clear advantages in both data size and challenge. This paper also conducts experiments and analyses using several summarization algorithms on this dataset.