We developed a system designed to serve as the first consultant for individuals uncertain about whether they may have broken the law. The system predicts which law sections might have been violated based on provided case details. We focused on seven of the most commonly occurring law sections: three from Civil and Commercial Law (Sections 420, 537, 1382) and four from Criminal Law (Sections 59, 80, 83, 288). The system was trained using past verdicts, generating document embeddings as the sole feature for classification, with labels extracted from metadata specifying the violated law sections, thus eliminating the need for expert labeling. Treating each verdict as a long document, we employed clustering and deep learning methods for classification. Additionally, we compared the system’s performance with ChatGPT, a large language model. The best performance was found when the verdict was divided into multiple 200-token chunks and classified using the Bi-LSTM method, achieving an average F1 score of 0.53, with Section 537 reaching the highest F1 score of 0.69.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Predicting Violated Law Sections Using Document Classification Techniques

  • Taneeya Satyapanich,
  • Nantanat Wattanakul,
  • Tawan Lehiang

摘要

We developed a system designed to serve as the first consultant for individuals uncertain about whether they may have broken the law. The system predicts which law sections might have been violated based on provided case details. We focused on seven of the most commonly occurring law sections: three from Civil and Commercial Law (Sections 420, 537, 1382) and four from Criminal Law (Sections 59, 80, 83, 288). The system was trained using past verdicts, generating document embeddings as the sole feature for classification, with labels extracted from metadata specifying the violated law sections, thus eliminating the need for expert labeling. Treating each verdict as a long document, we employed clustering and deep learning methods for classification. Additionally, we compared the system’s performance with ChatGPT, a large language model. The best performance was found when the verdict was divided into multiple 200-token chunks and classified using the Bi-LSTM method, achieving an average F1 score of 0.53, with Section 537 reaching the highest F1 score of 0.69.