Offense Feature Extraction and Comparative Analysis of Jaccard Similarity and Word Embedding Techniques for IPC Section Recommendations in First Information Report
摘要
This study emphasizes the pivotal role of advanced computational methods, particularly natural language processing (NLP), in offense analysis and IPC section classification within the realm of criminal justice. It underscores the importance of identifying IPC section applicability in First Information Reports (FIRs), detailing the process of achieving optimal results through comparisons between Jaccard similarity, word embedding techniques, and the offense feature extraction (OFE) method. A dataset comprising 100 FIRs was collected, with manual analysis conducted to select offense terms as feature selections and determine their applicable Indian Penal Code (IPC) sections. Subsequently, a model was trained using these offense terms to predict the applicable IPC sections, and tested on a set of 35 new FIRs to assess accuracy. The results highlight the effectiveness of leveraging advanced computational methods for offense analysis and classification, with implications for enhancing the efficiency and accuracy of legal proceedings. The average accuracy achieved by Jaccard similarity is impressively high. Moreover, when utilized in the mode Jaccard form, count vectorization consistently displays a tendency to attain superior accuracy across most FIRs. This indicates that Jaccard similarity, particularly when combined with count vectorization, showcases robust performance in effectively categorizing FIRs. Upon examining the offense feature extraction (OFE) method, the incidence of false negatives (FN) is found to be minimal. Future scope involves transitioning from individual word selection to group-based feature selection for improved accuracy and clarity, as well as automating and optimizing FIR interpretation and legal statute determination to enhance the efficiency, accuracy, and fairness of legal procedures.