Valence Ratings of Lemmas Generated by a Large Language Model Simulating Multiple Human Responses
摘要
Humans can report their opinions based on physical and emotional experiences, as well as different social and cultural roles. In this work, we investigate the ability of a Large Language Model (LLM) to answer about the valence of words based on the simulation of different respondents. The LLM GPT-3.5 was operated in chat mode, through its public application programming interface, with prompts developed to demand simulation of five people of various ages and backgrounds answering about the valence of 10 words on a scale of 1 to 9. This prompt was repeated 14 times, totaling 140 words, with and without the modifiers “average” and “overly positive”. The median of the five responses for each word was compared with the average valence previously reported by humans and median values of multiple prompts demanding the LLM to answer as a single simulated human. The valences estimated by the LLM when simulating an “average person” had a correlation coefficient of 0.8 with the human evaluations. A paired t-test indicate no significant difference between modes of simulation in the LLM, and between the distributions of LLM and human responses. Issues remain, such as robustness to prompt variations.