This study evaluates the performance of large language models (LLMs) in the qualitative analysis of an interview transcript. Four state-of-the-art LLMs—GPT-4, LLAMA 3.1, Gemini 2.0 flash, and DeepSeek-V3—were compared with human qualitative analysis using both inductive and deductive approaches. The deductive analysis employed the socio-ecological model as a framework to examine the complexities of supporting neurodivergent individuals in rural communities. Results indicate that Gemini achieved the highest cosine similarity to human coding in the inductive approach, while DeepSeek performed best in the deductive approach. Graph neural network (GNN) visualizations revealed that certain models struggled to capture the holistic context of the interview, demonstrating limitations in comprehension and contextual analysis compared to human coders. The models used in this study were freely available, and participant privacy was protected by anonymizing the transcript. The study was approved by the institutional review board. These findings highlight both the potential and the challenges of employing LLMs to augment qualitative research, particularly in nuanced and context-dependent data analyses. We discuss the implications of using LLMs to enhance qualitative research processes and propose future directions for improving their alignment with human analytical processes.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

From Code to Insight: How LLMs Help and Hinder Qualitative Research

  • Chaewon Kim,
  • Fengfeng Ke,
  • Alex Barrett,
  • Nuodi Zhang

摘要

This study evaluates the performance of large language models (LLMs) in the qualitative analysis of an interview transcript. Four state-of-the-art LLMs—GPT-4, LLAMA 3.1, Gemini 2.0 flash, and DeepSeek-V3—were compared with human qualitative analysis using both inductive and deductive approaches. The deductive analysis employed the socio-ecological model as a framework to examine the complexities of supporting neurodivergent individuals in rural communities. Results indicate that Gemini achieved the highest cosine similarity to human coding in the inductive approach, while DeepSeek performed best in the deductive approach. Graph neural network (GNN) visualizations revealed that certain models struggled to capture the holistic context of the interview, demonstrating limitations in comprehension and contextual analysis compared to human coders. The models used in this study were freely available, and participant privacy was protected by anonymizing the transcript. The study was approved by the institutional review board. These findings highlight both the potential and the challenges of employing LLMs to augment qualitative research, particularly in nuanced and context-dependent data analyses. We discuss the implications of using LLMs to enhance qualitative research processes and propose future directions for improving their alignment with human analytical processes.