<p>A recent trend in computational cognitive science is to use large-scale naturalistic data to evaluate cognitive theories (Griffiths, 2015; Jones, 2017). This article extends this work by examining the connection between social feedback and language usage. Social feedback, operationalized here as the difference between the number of upvotes and downvotes on comments on the social media forum Reddit, provides a unique opportunity to investigate how social approval interacts with linguistic behavior at a large scale. To accomplish this, social feedback ratings from over 400 billion words of comments from the social media forum Reddit were mined. Two types of corpora were formed: 1) large (approximately 1.7 billion words) combined corpora of positive, negative, and neutral comments, and 2) individual corpora from over 1,500 high-level commenters on Reddit (following recent work by Johns, 2021a, 2023, 2024a,b using individualized language models). Both were used to examine social feedback and language usage across multiple sets of lexical organization and lexical semantic data with distributional modeling techniques. It was found that comments that receive positive social feedback tended to align more closely with normative language usage, suggesting that the acceptability of a communicative message may lie in the distributional form of that utterance and not just the content of it. Additionally, this trend was replicated at the individual level, where commenters who had more normative (or typical) usage of language generally received more positive ratings on their comments.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Large-Scale Social Feedback and Language Usage: a Distributional Analysis

  • Brendan T. Johns

摘要

A recent trend in computational cognitive science is to use large-scale naturalistic data to evaluate cognitive theories (Griffiths, 2015; Jones, 2017). This article extends this work by examining the connection between social feedback and language usage. Social feedback, operationalized here as the difference between the number of upvotes and downvotes on comments on the social media forum Reddit, provides a unique opportunity to investigate how social approval interacts with linguistic behavior at a large scale. To accomplish this, social feedback ratings from over 400 billion words of comments from the social media forum Reddit were mined. Two types of corpora were formed: 1) large (approximately 1.7 billion words) combined corpora of positive, negative, and neutral comments, and 2) individual corpora from over 1,500 high-level commenters on Reddit (following recent work by Johns, 2021a, 2023, 2024a,b using individualized language models). Both were used to examine social feedback and language usage across multiple sets of lexical organization and lexical semantic data with distributional modeling techniques. It was found that comments that receive positive social feedback tended to align more closely with normative language usage, suggesting that the acceptability of a communicative message may lie in the distributional form of that utterance and not just the content of it. Additionally, this trend was replicated at the individual level, where commenters who had more normative (or typical) usage of language generally received more positive ratings on their comments.