Small Languages and Big Models: Using ML to Generate Norwegian Language Social Media Content for Training Purposes
摘要
The advancement of language models has showcased their tremendous potential for both good purposes, and harmful misuse. However, the majority of research have been concentrated on high-resource languages, leaving much to be desired in low-resource languages. This article focuses on exploring the use of language models in Norwegian, a low-resource language. Addressing the threats these models pose in the context of influence operations in social media. The methodology uses a mixed-methods approach, combining quantitative analysis and qualitative investigations. The quantitative analysis entails evaluating the performance of language models across various contexts, assessing their ability to generate perceived authentic content, and analyzing user responses to such generated content. The qualitative investigations involve conducting interviews and surveys to gather insights from participants, aiming to understand their experiences, perceptions, and concerns regarding the use of language models. By investigating the use of language models in a low-resource language, this thesis aims to contribute to the advancement of natural language processing research in an underrepresented linguistic context. As well as exploring the use of these language models for training purposes in isolated social networks.