Detecting and managing participants’ large language model use for open–ended questions in online research
摘要
Recent studies have shown that a substantial share of participants use large language models (LLMs) to generate responses in open–ended online research settings, reducing the validity and authenticity of research data. While approaches for detecting the use of LLMs by participants exist, researchers seldom apply such methods. To investigate the consequences of LLM use for research data, we conducted an online study of a creative brainstorming task on Prolific (N = 450) and MTurk (N = 442). We employed two different indicators of LLM use: self–reports and user activity monitoring (the number of words generated per minute, the act of leaving a web page, and keystroke analysis). The keystroke analysis provided the most thorough indication of LLM use, especially in combination with other indicators. Self–reports indicated potential LLM use on secondary devices. Participants with indications of LLM use demonstrated higher creative fluency (number of generated ideas) and longer responses (number of words per response), strongly influencing the study results produced on MTurk but not on Prolific. In both studies, the data derived from participants with indications of LLM use exhibited lower construct reliability. In addition, asking the participants not to use LLMs reduced the frequency of LLM indications on Prolific but not on MTurk. We conclude that the consequences of LLM use depend on the research context, the type and extent of LLM use, and sample characteristics. For research practice, we recommend a combination of LLM detection measures, preventive study designs, and testing schemes to assess the consequences of LLM use for individual research settings.