Cost–benefit analysis of deploying shallow, deep learning and generative models for legal text classification
摘要
Recent advances in Generative Language Models (GLMs) have renewed focus on promising results in zero-shot text classification. However, their off-the-shelf performance on unfamiliar and domain specific tasks remains uncertain. In this legal clause classification task we evaluate a plug-and-play zero-shot prompting strategy for OpenAI’s GPT-4 GLM on a contract clause dataset. We introduce the new CUAD-SL dataset that has been refactored as a single label classification problem as a fairer and more robust legal classification benchmark. In a comparative study, we show that fine-tuning on legal domain data adapts smaller, less complex models to the task at hand, with significant classification accuracy improvement of up to 20.6%, with a best overall performance of 87.8% for the DeBERTa Transformer model compared to GPT-4's 67.2% performance. This study also takes the novel approach of assessing the business feasibility of deploying each of these machine learning models through a detailed cost–benefit analysis that measures the trade-off between performance metrics and low and high usage running costs.