Did You Tell a Deadly Lie? Evaluating Large Language Models for Health Misinformation Identification
摘要
The rapid spread of health misinformation online poses significant challenges to public health, potentially leading to confusion, undermining trust in health authorities, and hindering effective health interventions. Large Language Models (LLMs) have shown promise in various natural language processing tasks, including misinformation detection. However, their effectiveness in identifying health-specific misinformation has not been extensively benchmarked. This study evaluates the performance of seven state-of-the-art LLMs - GPT-3.5, GPT-4, Gemini, Flan-T5 XL, Gemma, LLaMA-2, and Mistral - on the task of health misinformation detection across four datasets (Monkeypox-V1, Monkeypox-V2, COVID-19, and CoAID). The models were tested under five different settings: zero-shot classification, 5-shot random examples, 10-shot random examples, 5-shot sampled examples, and 10-shot sampled examples. Performance was evaluated using macro F1-score, and inter-model agreement was assessed using Cohen’s Kappa and Fleiss’ Kappa scores. By comprehensively benchmarking these LLMs, this study aims to determine which models excel in particular scenarios and provide insights into their potential for combating health misinformation in online environments.