ViFoodNLI: A Dataset for Vietnamese Natural Language Inference in Local Cuisine
摘要
This paper introduces the ViFoodNLI dataset, a natural language inference (NLI) dataset for Vietnamese. While recent efforts have been made to build high-quality NLI datasets for Vietnamese and some Cross-Lingual NLI Corpus (with support for Vietnamese) for multiple domains, our dataset specifically focuses on the field of local cuisine. The main reason for choosing this field is that cuisine is a significant component of Vietnamese culture, and thus the dataset encompasses many characteristics of the Vietnamese language. By collecting information on culinary topics from reliable news sources, we have developed various methods and logics such as knowledge graphs and Generative AI to create high-quality pairs of premise and hypothesis sentences. Through rigorous testing, the dataset has achieved significant results, creating momentum for future research and practical applications.