VSum-HB: A Vietnamese Text Summarization Dataset For Reinforcement Learning From Human Feedback
摘要
This paper presents a novel Vietnamese text summarization dataset named VSum-HB, designed specifically for Reinforcement Learning from Human Feedback (RLHF). A total of 5,000 samples from the Vietnews corpus were carefully selected, covering a wide range of topics. We introduce a hybrid method combining automated summarization using ChatGPT 4.o with human annotation to refine the summaries. This approach ensures the dataset is highly suitable for training RLHF models to generate summaries closely aligned with human preferences. Experimental results indicate that models trained on this dataset, enhanced by RLHF, significantly outperform traditional summarization models, with marked improvements in ROUGE scores and positive evaluations from human. Our work highlights the potential of this dataset to advance Vietnamese NLP research, with future studies aiming to expand its applications to other NLP tasks. Our datasets are also published in Hugging-Face framework ( https://huggingface.co/datasets/minhquy1624/VSUMHB ).