Improving Robustness in Language Models for Legal Textual Entailment Through Artifact-Aware Training
摘要
In this paper, we describe our participation in COLIEE 2024, focusing on legal textual entailment (Tasks 2 and 4). Our goal is to address language artifacts during language model training for improved robustness. Limited domain-specific datasets pose challenges, leading us to apply language artifact detection and mitigation methods tailored to legal textual entailment tasks. For Task 2, involving identifying relevant paragraphs in previous cases, we address annotation artifacts in premises. In Task 4, predicting if statute law articles entail or contradict legal bar exam questions, we identify and mitigate various artifact types in hypotheses. We caution against relying on specific language artifacts and advocate for data profiling and measures to balance these implicit biases, enhancing overall model robustness.