Text Summarisation for Low-Resourced Languages, A Review
摘要
Text summarisation is becoming increasingly important for humans to more quickly understand and analyse documents with large amounts of text. In this paper, we review and discuss approaches and methods used in the development of text summarisation models for low-resourced languages, specifically South African languages. We compare approaches and results to give guidance on what may be the best approach to building a sophisticated text summarisation model for South African languages. The results showed that there is one text summarisation model created for isiXhosa out of 11 South African languages, and only a few studies were done for African languages. We recommend future work to focus on developing necessary datasets for South African languages, developing language-specific preprocessing tools such as stemmers and stop-word lists, and finally, using the developed data to build or use more sophisticated language models.