Ti-MSUM: Tibetan Image and Text Summarization Dataset
摘要
Automatic image-text summarization is a key research direction in natural language processing technology, which is significant for addressing information overload and enhancing the usability and comprehensibility of text data. Tibetan, as a minority language in China, falls within the category of relatively scarce languages, possessing a unique writing system and grammatical rules. Compared to mainstream languages such as Chinese and English, research progress in the field of automatic image-text summarization for Tibetan has been relatively slow, primarily due to the lack of high-quality available datasets. To fill this gap, we utilized web crawling technology to collect 4,500 authentic Tibetan articles from various Tibetan WeChat public accounts, and innovatively used each article’s title as a summary to construct a rich and diverse Tibetan image-text summarization dataset, Ti-MSUM. The launch of this dataset aims to meet the needs of researchers and further promote the rapid development of Tibetan in the field of automatic image-text summarization.