<p>While qualitative research plays a vital role in understanding complex phenomena, it lends itself poorly to testing formal hypotheses due to its inability to fit statistical models to text data. Approaches that are traditionally used to quantify text data (e.g., content analysis) are generally time-consuming, prone to researcher bias, and neglect a substantial amount of potentially important semantic context. Although novel approaches have been proposed, these typically require large amounts of text data and tend to be inductive in nature. To enable researchers to ask hypothesis-based and open-ended questions from one’s text data, the current study proposes a novel retrieval augmented generation (RAG)-based approach (called text embedding similarity analysis, TESA) that transforms a hypothesis into two specific search terms: a population (or sample) and a variable of interest. Using pretrained large language models (LLM), we extract the semantic embedding of the search terms and text data and use cosine similarity to match search terms. This allows hypothesis testing by assessing the alignment between the distribution of similarity scores for a variable of interest with the expectation for the population.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Talk to your data: Introducing text embedding similarity analysis (TESA) in psychological research

  • Juul Vossen,
  • Evy Kuijpers,
  • Joeri Hofmans

摘要

While qualitative research plays a vital role in understanding complex phenomena, it lends itself poorly to testing formal hypotheses due to its inability to fit statistical models to text data. Approaches that are traditionally used to quantify text data (e.g., content analysis) are generally time-consuming, prone to researcher bias, and neglect a substantial amount of potentially important semantic context. Although novel approaches have been proposed, these typically require large amounts of text data and tend to be inductive in nature. To enable researchers to ask hypothesis-based and open-ended questions from one’s text data, the current study proposes a novel retrieval augmented generation (RAG)-based approach (called text embedding similarity analysis, TESA) that transforms a hypothesis into two specific search terms: a population (or sample) and a variable of interest. Using pretrained large language models (LLM), we extract the semantic embedding of the search terms and text data and use cosine similarity to match search terms. This allows hypothesis testing by assessing the alignment between the distribution of similarity scores for a variable of interest with the expectation for the population.