Mongolian-Chinese Cross-Lingual Topic Detection Based on Knowledge Distillation and Contrastive Learning Methods
摘要
Most topic detection research focuses on cross-language information processing in languages with abundant resources, such as English-Chinese, English-German, and others. There are also some studies targeting low-resource languages such as Vietnamese-Chinese and Tibetan-Chinese cross-language topic detection research. However, there is very little research on Mongolian-Chinese cross-language studies. The primary reasons are the lack of corpora in the Mongolian language and the limitations of conventional topic detection methods in text representation, clustering, and topic representation within the Mongolian language domain. Therefore, this paper proposes a Mongolian-Chinese cross-lingual topic detection method based on knowledge distillation and contrastive learning. Firstly, the pre-trained language model is fine-tuned through knowledge distillation tasks to enable the model to integrate the semantic representation of Mongolian-Chinese cross-lingual texts. Then, the effect of cross-lingual topic detection is improved by combining contrastive learning training between news of different topics. Experimental results show that the proposed method combining knowledge distillation and contrastive learning outperforms other baseline models, achieving at least 2, 0.8, and 3% points higher in F1 score, topic diversity, and topic consistency, respectively. This approach effectively enhances the accuracy of Mongolian-Chinese cross-lingual topic detection.