Long live fine-tuning: task-specific transformers outperform zero-shot LLMs for misinformation response classification on Reddit
摘要
As large language models (LLMs) become widely used tools for online information access and verification, it is important to determine when their zero-shot flexibility is sufficient and when task-specific supervision remains valuable. We examine this question in a controlled misinformation-response classification setting comprising 900 Reddit comments associated with three PolitiFact-verified claims in environment, health, and immigration. Comments are labelled as belief (propagates the claim), fact-check (corrects it), or other. We compare nine models across three paradigms—BART-MNLI, three Llama variants, three commercial LLMs (Claude Haiku 4.5, Gemini Flash Lite 2.5, and Claude Sonnet 4.6), and fine-tuned DistilBERT and RoBERTa—under universal and topic-specific label schemas. Under the evaluated prompting and supervision conditions, fine-tuned RoBERTa reaches 0.62 macro-