<p>Several prior studies have used advanced methodological techniques to demonstrate that there is an issue with the quality of data that can be collected on Amazon’s Mechanical Turk (MTurk). The goal of the present project was to provide an accessible demonstration of this issue. We administered 27 semantic antonyms—pairs of items that assess clearly contradictory content (e.g., “I talk a lot” and “I rarely talk”)—to samples drawn from Connect (<i>N</i><sub>1</sub> = 100), Prolific (<i>N</i><sub>2</sub> = 100), and MTurk (<i>N</i><sub>3</sub> = 400; <i>N</i><sub>4</sub> = 600). Despite most of these item pairs being negatively correlated on Connect and Prolific, over 96% were <i>positively</i> correlated on MTurk. This issue could not be remedied by screening the data using common attention check measures nor by recruiting only “high-productivity” and “high-reputation” participants. These findings provide clear evidence that data collected on MTurk simply cannot be trusted.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Why you shouldn’t trust data collected on MTurk

  • Cameron S. Kay

摘要

Several prior studies have used advanced methodological techniques to demonstrate that there is an issue with the quality of data that can be collected on Amazon’s Mechanical Turk (MTurk). The goal of the present project was to provide an accessible demonstration of this issue. We administered 27 semantic antonyms—pairs of items that assess clearly contradictory content (e.g., “I talk a lot” and “I rarely talk”)—to samples drawn from Connect (N1 = 100), Prolific (N2 = 100), and MTurk (N3 = 400; N4 = 600). Despite most of these item pairs being negatively correlated on Connect and Prolific, over 96% were positively correlated on MTurk. This issue could not be remedied by screening the data using common attention check measures nor by recruiting only “high-productivity” and “high-reputation” participants. These findings provide clear evidence that data collected on MTurk simply cannot be trusted.