Improving Vietnamese Legal Question–Answering System Based on Automatic Data Enrichment
摘要
Question answering (QA) in law presents a significant challenge, as legal documents are often complex in terms of terminology, structure, and temporal and logical relationships. This is particularly difficult for low-resource languages like Vietnamese, where labelled data are scarce and pre-trained language models remain limited. In this paper, we address these limitations by developing a Vietnamese, article-level retrieval-based legal QA system and introducing a novel approach to enhancing language model performance through data quality improvement via weak labelling. We hypothesize that, in contexts where labelled data are limited, effective data enrichment can boost overall performance. Our experimental design evaluates multiple aspects and demonstrates the efficacy of the proposed technique.