错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

The Application of Chinese Word Segmentation to Less-Resourced Language Processing

  • Meng-hsien Shih

摘要

It has been more than half a century since the first million-word Brown Corpus was constructed. However, the lack of resources has posed a challenge to corpus construction and language processing for low-resource languages. Therefore, the corpus construction of most local languages such as Hokkien in Taiwan is still under development. In this paper, a Chinese segmenter based on syntactic analysis is adapted with a Hokkien dictionary (from Taiwan’s Ministry of Education) to segment words in Hokkien texts. The proposed approach reports an accuracy rate of 87.50% in correctly separating the lemmas from the example sentences in the dictionary, and an average performance of 92.78% in testing with additional six unseen Hokkien e-paper articles. The source code of the proposed Hokkien segmenter is released with an online corpus of Hokkien word segmentation. This segmenter can be used to construct larger Hokkien corpus with segmented words, and apply to the segmentation of other less-resourced languages in the Chinese language family including Hakka.