Legal natural language processing has recently received a surge in interest from experts due to its potential application in various domains. One of the most popular tasks in this area is legal case categorization and judgment prediction, which can be seen as a classification problem. Previous attempts at the classification task mostly employ traditional methods, such as the static embedding method in combination with machine learning or deep learning models. Even though such approaches can achieve a notable performance, they have yet to reach state-of-the-art performance. With the significant leap in text and natural language processing, namely the advent of transformer-based models, along with their consistent improvement, we see the opportunity to adopt this development to the legal classification task. Our proposed modeling pipeline, the WangchanBERTa-SFT-sc, can outperform the baseline model using the fastText embedding method and conventional machine learning models on the binary classification of whether or not the case is related to property offence, reaching exceptional test accuracy and F1-score of 94.5%. This finding highlights the capability and importance of contextual comprehension in dealing with the text classification task in a specific and highly sophisticated environment, particularly the legal domain.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Thai Legal Fact Classification of Property-Related Offences Using Finetuned BERT Modelling

  • Sirawit Chokphantavee,
  • Sorawit Chokphantavee,
  • Somrudee Deepaisarn

摘要

Legal natural language processing has recently received a surge in interest from experts due to its potential application in various domains. One of the most popular tasks in this area is legal case categorization and judgment prediction, which can be seen as a classification problem. Previous attempts at the classification task mostly employ traditional methods, such as the static embedding method in combination with machine learning or deep learning models. Even though such approaches can achieve a notable performance, they have yet to reach state-of-the-art performance. With the significant leap in text and natural language processing, namely the advent of transformer-based models, along with their consistent improvement, we see the opportunity to adopt this development to the legal classification task. Our proposed modeling pipeline, the WangchanBERTa-SFT-sc, can outperform the baseline model using the fastText embedding method and conventional machine learning models on the binary classification of whether or not the case is related to property offence, reaching exceptional test accuracy and F1-score of 94.5%. This finding highlights the capability and importance of contextual comprehension in dealing with the text classification task in a specific and highly sophisticated environment, particularly the legal domain.