<p>As large language models (LLMs) become widely used tools for online information access and verification, it is important to determine when their zero-shot flexibility is sufficient and when task-specific supervision remains valuable. We examine this question in a controlled misinformation-response classification setting comprising 900 Reddit comments associated with three PolitiFact-verified claims in environment, health, and immigration. Comments are labelled as <i>belief</i> (propagates the claim), <i>fact-check</i> (corrects it), or <i>other</i>. We compare nine models across three paradigms—BART-MNLI, three Llama variants, three commercial LLMs (Claude Haiku&#xa0;4.5, Gemini Flash Lite&#xa0;2.5, and Claude Sonnet&#xa0;4.6), and fine-tuned DistilBERT and RoBERTa—under universal and topic-specific label schemas. Under the evaluated prompting and supervision conditions, fine-tuned RoBERTa reaches 0.62 macro-<InlineEquation ID="IEq1"><EquationSource Format="TEX">\(F_1\)</EquationSource></InlineEquation>, compared with 0.50 for the best zero-shot configuration (Claude Haiku&#xa0;4.5), at a fraction of the per-query cost. The supervised advantage is concentrated on the <i>belief</i> class, which every evaluated zero-shot model under-detects. Increasing model size does not produce consistent improvements: Llama-3-8B matches Llama-3-70B in several settings, while Claude Sonnet&#xa0;4.6 underperforms the smaller Haiku model under generic labels. Sonnet’s belief <InlineEquation ID="IEq2"><EquationSource Format="TEX">\(F_1\)</EquationSource></InlineEquation> increases substantially when generic labels are replaced with claim-specific descriptions, indicating that label formulation and claim context interact with model behaviour. Performance also varies by more than 0.13 macro-<InlineEquation ID="IEq3"><EquationSource Format="TEX">\(F_1\)</EquationSource></InlineEquation> across the three claim datasets under matched label conditions. For the claims, label schemas, and zero-shot configurations evaluated here, task-specific fine-tuning provides the more reliable classification strategy when labelled examples are available.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Long live fine-tuning: task-specific transformers outperform zero-shot LLMs for misinformation response classification on Reddit

  • JooYoung Lee,
  • Lin Tian,
  • Angela Brillantes,
  • Adriana-Simona Mihăiță,
  • Marian-Andrei Rizoiu

摘要

As large language models (LLMs) become widely used tools for online information access and verification, it is important to determine when their zero-shot flexibility is sufficient and when task-specific supervision remains valuable. We examine this question in a controlled misinformation-response classification setting comprising 900 Reddit comments associated with three PolitiFact-verified claims in environment, health, and immigration. Comments are labelled as belief (propagates the claim), fact-check (corrects it), or other. We compare nine models across three paradigms—BART-MNLI, three Llama variants, three commercial LLMs (Claude Haiku 4.5, Gemini Flash Lite 2.5, and Claude Sonnet 4.6), and fine-tuned DistilBERT and RoBERTa—under universal and topic-specific label schemas. Under the evaluated prompting and supervision conditions, fine-tuned RoBERTa reaches 0.62 macro-\(F_1\), compared with 0.50 for the best zero-shot configuration (Claude Haiku 4.5), at a fraction of the per-query cost. The supervised advantage is concentrated on the belief class, which every evaluated zero-shot model under-detects. Increasing model size does not produce consistent improvements: Llama-3-8B matches Llama-3-70B in several settings, while Claude Sonnet 4.6 underperforms the smaller Haiku model under generic labels. Sonnet’s belief \(F_1\) increases substantially when generic labels are replaced with claim-specific descriptions, indicating that label formulation and claim context interact with model behaviour. Performance also varies by more than 0.13 macro-\(F_1\) across the three claim datasets under matched label conditions. For the claims, label schemas, and zero-shot configurations evaluated here, task-specific fine-tuning provides the more reliable classification strategy when labelled examples are available.