HALIFacts: Evaluating Large Language Models for Domain-Specific Fact-Checking and Their Carbon Impact
摘要
Fact-checking has emerged as a crucial and in-demand field in recent years. Advances in machine learning and pretrained language models have enabled natural language processing techniques to assist in automating this task. Tailored neural architectures have previously been employed for fact-checking to address the dual objectives of veracity prediction and evidence selection. In this work, we demonstrate that comparable or even superior results can be obtained using large language models (LLMs). This is performed at a fraction of the training cost, particularly in terms of CO \(_2\) emissions. We introduce HALIFacts, a LLM training framework that leverages state-of-the-art LLM training techniques, special tokens and prompt engineering for this dual fact-checking task. We show that our approach either outperforms MultiVerS, the current state-of-the-art, or achieves comparable results across three specialised datasets. Furthermore, we demonstrate that it does so while using significantly less training data, resulting in a CO \(_2\) emission cost that is lower by an order of magnitude. The HALIFacts framework is very versatile, and it can thus be applied to any LLM and can scale easily.