<p>Legal document analysis presents significant challenges due to its complexity and domain-specific nature. This study introduces an innovative approach for classifying Indian court judgments into legal domains using diverse machine learning and deep learning techniques. The method incorporates feature engineering and deep learning algorithms for extracting meaningful features, supported by a wide range of classifiers, including voting classifiers, gradient boosting, and random forest. Embeddings are generated using models such as InLegalBERT, InCaseLawBERT, CustomInLawBERT, Mamba, T5, RoBERTa, CodeT5, SBERT, DistilBERT, XLM, XLM Large, LegalBERT, GPT2, ALBERT, Electra, DeBERTa, TFDeBERTa, FlanT5, FlanT-Large, BART, BigBird Pegasus, LongFormer, and LUKE. To address class imbalance, the SMOTE technique is employed, and dimensionality is reduced using PCA and forward feature selection. The T5+SMOTE+feature selection+voting classifier configuration achieves a notable accuracy of 98%, highlighting the effectiveness of the proposed approach. These advancements have significant implications for applications such as document retrieval, legal discovery, and case law analysis, enhancing the accuracy and efficiency of legal document classification.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Enhancing Indian legal judgment classification with embeddings, feature selection, and ensemble strategies

  • Priyanka Prabhakar,
  • Peeta Basa Pati

摘要

Legal document analysis presents significant challenges due to its complexity and domain-specific nature. This study introduces an innovative approach for classifying Indian court judgments into legal domains using diverse machine learning and deep learning techniques. The method incorporates feature engineering and deep learning algorithms for extracting meaningful features, supported by a wide range of classifiers, including voting classifiers, gradient boosting, and random forest. Embeddings are generated using models such as InLegalBERT, InCaseLawBERT, CustomInLawBERT, Mamba, T5, RoBERTa, CodeT5, SBERT, DistilBERT, XLM, XLM Large, LegalBERT, GPT2, ALBERT, Electra, DeBERTa, TFDeBERTa, FlanT5, FlanT-Large, BART, BigBird Pegasus, LongFormer, and LUKE. To address class imbalance, the SMOTE technique is employed, and dimensionality is reduced using PCA and forward feature selection. The T5+SMOTE+feature selection+voting classifier configuration achieves a notable accuracy of 98%, highlighting the effectiveness of the proposed approach. These advancements have significant implications for applications such as document retrieval, legal discovery, and case law analysis, enhancing the accuracy and efficiency of legal document classification.