HCU: A Multi-lingual Health Claim Dataset and Its NLP Applications
摘要
We propose a multi-lingual health claim dataset for NLP applications and cross-discipline study. It has six languages including English, French, German, Poland, Romanian, and Hungarian. It has three sub datasets: EFSA authorized, manufacturer reworded, and consumer created Health Claims. We conducted basic data analysis of the dataset, described the data collection process, and discussed two NLP applications including text style transfer and machine translation. We applied recent machine learning models and conducted empirical evaluations of these models using this dataset. The experimental results show that this dataset can be used to investigate the design and evaluation of novel algorithms of NLP tasks. This dataset and its NLP applications can be used for food industry to generate multi-lingual health claims with different styles easily and automatically. This dataset bridges the gap of lacking good quality data for these applications in Health Claim area.