错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

On measurement of distances between texts in dictionary-based content analysis

  • Anton Oleinik

摘要

The article discusses the measurement of distances between heterogeneous texts. Some limitations of WordStat, a popular off-the-shelf software package for content analysis, when measuring distances between texts, are identified and investigated with the help of an experiment. A corpus of texts (c. 4 million words) composed of political leaders’ speeches and news items about Russia’s invasion of Ukraine in three languages was analyzed twice, using WordStat and an algorithm with explicitly set parameters. The same custom-built dictionary was used in both cases. A larger corpus of texts (c. 16 million words) was also analyzed using an extended version of the dictionary and the proposed metrics, Sigma (the standard deviation of observed frequencies from expected frequencies) and Cohen’s d. Some remedies are discussed, including the additional processing of output generated by WordStat and adding Sigma to the list of (dis)similarity measures.