Legar: A Legal Statute Identification System for Vietnamese Users on Land Law Matters
摘要
This paper introduces Legar, an innovative Legal Statute Identification system designed to cater to Vietnamese users seeking legal guidance on land law matters. Legar focuses on practical efficiency by harnessing the digital capabilities of the Vietnam Landlaw 2013. This is achieved through the utilization of a specialized legal-masked language model, based on the RoBERTa architecture, leading to the development of the law-oriented language model, LegaRBERT. Subsequently, LegaRBERT is incorporated into a multi-label classification model, XGBoots, to offer consultation services based on user-generated questions. As a result, Legar effectively addresses prevailing limitations found in comparable LSI research and prototypes. Empirical evaluations conducted using authentic datasets from Vietnamese legal consultations demonstrate that Legar surpasses existing baselines, particularly when considering our proposed K-Utility metric, which reflects the practical expectations of LSI users. The initial version of Legar is now publicly accessible and has garnered positive feedback from users.