Assessing the Potential and Limits of Large Language Models in Qualitative Coding
摘要
This paper examines the advantages and limitations of conducting automated coding of virtual tutoring session transcripts using the GPT-4 Turbo model via the OpenAI API. We compare three coding methods: (1) zero-shot, which relies solely on construct definitions; (2) few-shot, which includes annotated examples; and (3) coding with context, which provides GPT-4 with surrounding dialogue and study context. We used these approaches to code ten constructs from an existing codebook. We then had a set of experienced qualitative researchers rate the set of constructs across multiple dimensions. The results show that while zero-shot coding is effective for constructs with clear definitions, it tends to miss cases and struggles with constructs requiring contextual understanding. Few-shot coding works well for constructs that are seen as objective by experts, and those for which experts feel examples are needed to fully understand. However, it tends to overgeneralize based on the examples included in the prompt. Coding with context is particularly effective for constructs that often appear as part of sequences, but can also lead to the model coding more based on the context rather than the current line. This investigation highlights the potential of GPT-4 Turbo for efficient auto-coding of large datasets but emphasizes that specific prompting decisions impact quality and that the optimal decisions vary based on the characteristics of what is being coded.