错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Data Augmentation and Large Language Model for Legal Case Retrieval and Entailment

  • Minh-Quan Bui,
  • Dinh-Truong Do,
  • Nguyen-Khang Le,
  • Dieu-Hien Nguyen,
  • Khac-Vu-Hiep Nguyen,
  • Trang Pham Ngoc Anh,
  • Minh Le Nguyen

摘要

The Competition on Legal Information Extraction and Entailment (COLIEE) is a well-known international competition organized each year with the goal of applying machine learning algorithms and techniques in the analysis and understanding of legal documents. Two main applications of using machine learning in this domain are entailment and information retrieval. In the realm of legal text analysis, the scarcity of annotated data poses a significant challenge for training robust models. To address this limitation, we employ data augmentation methods to artificially expand the training dataset, enhancing the model’s ability to generalize across diverse legal contexts. Additionally, our approach harnesses the power of a state-of-the-art language model, enabling the extraction of nuanced legal information and improving entailment predictions. We evaluate the performance of our methodology on datasets from the competition, showcasing its effectiveness in achieving competitive results.