<p>A non-disclosure agreement (NDA) is a legal contract between at least two business parties that restricts the disclosure of confidential and sensitive information. As part of the NDA document negotiation and review workflows, there is a need to identify, extract and track clauses to ensure they comply with the companies’ policies. The legal Natural Language Processing (NLP) landscape is rapidly evolving and offers a range of technologies and Machine Learning tools, such as text classification, that enable users to automate various key stages of the contract life cycle, thus easing the administrative burden. The proposed system was developed to classify individual clauses within an NDA contract. Automated legal text classification is challenging due to the high dimensionality of a word-based feature space. Additionally, corporate privacy concerns have limited the availability of publicly accessible legal corpora. To leverage the power of deep learning neural architectures, pre-trained word embeddings have attempted to resolve these issues by reusing an input feature space trained on a large, general-purpose dataset, and then fine-tuning it to adapt to a downstream classification task on a smaller annotated legal dataset. In this paper, we evaluate several shallow and deep supervised learning strategies for the classification of 26 individual mutually exclusive clause classes within an NDA contract. The impact of using pre-trained word embeddings on a small legal NDA contract dataset is also evaluated. The shallow SVM and XGBoost classification methods outperform the deep learning long short-term memory neural network (LSTM) approach in a small imbalanced dataset, even when supported by pre-trained embeddings. The potential of using novel transfer learning techniques that allow reuse of legal-domain specific NLP models and feature representation schemes is explored.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Evaluating Shallow and Deep Learning Strategies for Legal Text Classification of Clauses in Non-Disclosure Agreements

  • Niall McCarroll,
  • Philip McShane,
  • Eoin O’Connell,
  • Kevin Curran,
  • Muskaan Singh,
  • Eugene McNamee,
  • Angela Clist,
  • Andrew Brammer

摘要

A non-disclosure agreement (NDA) is a legal contract between at least two business parties that restricts the disclosure of confidential and sensitive information. As part of the NDA document negotiation and review workflows, there is a need to identify, extract and track clauses to ensure they comply with the companies’ policies. The legal Natural Language Processing (NLP) landscape is rapidly evolving and offers a range of technologies and Machine Learning tools, such as text classification, that enable users to automate various key stages of the contract life cycle, thus easing the administrative burden. The proposed system was developed to classify individual clauses within an NDA contract. Automated legal text classification is challenging due to the high dimensionality of a word-based feature space. Additionally, corporate privacy concerns have limited the availability of publicly accessible legal corpora. To leverage the power of deep learning neural architectures, pre-trained word embeddings have attempted to resolve these issues by reusing an input feature space trained on a large, general-purpose dataset, and then fine-tuning it to adapt to a downstream classification task on a smaller annotated legal dataset. In this paper, we evaluate several shallow and deep supervised learning strategies for the classification of 26 individual mutually exclusive clause classes within an NDA contract. The impact of using pre-trained word embeddings on a small legal NDA contract dataset is also evaluated. The shallow SVM and XGBoost classification methods outperform the deep learning long short-term memory neural network (LSTM) approach in a small imbalanced dataset, even when supported by pre-trained embeddings. The potential of using novel transfer learning techniques that allow reuse of legal-domain specific NLP models and feature representation schemes is explored.