Stance in Translation: English-to-Turkish Cross-Lingual Stance Detection
摘要
Stance detection is the task of understanding the position of an author of a text with respect to a specific target, and research on this task has been commonly performed under mono-lingual settings. Especially for languages other than English, related work has been limited, partly due to the lack of annotated data. Nonetheless, recent studies have sought to capitalize on labeled data available in languages abundant with resources to facilitate stance detection in low-resource languages. In this study, we target at this cross-lingual stance detection problem and propose an approach based on translating existing stance-annotated tweet datasets in English to generate new datasets in a target low-resource languages, and subsequently employing them for stance detection in the target language. We have conducted experiments with different learning approaches and high-performance machine translation tools in order to demonstrate the contribution of cross-lingual stance detection for stance detection in the low-resource language of Turkish. Additionally, we have conducted an ablation study on data pre-processing methods and incorporated a data augmentation technique to observe the effects of these approaches on the overall stance detection performance. We also present the results of detecting stance on native datasets using translated datasets to showcase the contributions of our method. With minimal loss of information across English to Turkish dataset, the results obtained are very promising for cross-lingual stance detection regarding low-resource languages including Turkish.