Empirical Study on Translating Toxic Sentences with Large Language Model vs. Machine Translation
摘要
In this research, we explore how large language model (LLM) and machine translation (MT) perform in translating toxic sentences, providing insights into their capabilities in handling toxicity in translation. We employ ChatGPT as our LLM and leverage Google Translate as the MT model in our research. The empirical study on toxic sentence translation is conducted by mainly focused on translating three toxic sentence datasets in English, Chinese, and Indonesian languages. Our results indicate that LLM may reject translating toxic sentences in less than 8.67% of cases. We also investigate the comparison of the toxicity score of translations yielded by LLM vs. MT. Among the successful translations, we observe that LLM-generated translations on average have lower toxicity scores compared to the machine translation model. However, in further analysis, we find that number of LLM-generated translations varies from 18–42% display higher toxicity scores when compared to those generated by the machine translation model. We further evaluate several LLMs on the task of translating toxic sentences and perform an interpretability analysis by examining the activations associated with refusing and generating translations. Warning: This paper contains harmful and abusive contents.