Creation of a Unique Clustering Method Employing Novel Similarity Metrics for Legal Texts to Improve Information Management and Retrieval in the Legal Field
摘要
The goal of this study is to provide specialised clustering algorithms and similarity metrics for legal documents. Legal papers need specialised processing methods since they are highly structured and contain a lot of technical and legal jargons. We provide a clustering technique that uses domain-specific knowledge to find groups of related documents while taking into account the hierarchical structure of legal texts. The program uses a mix of graph-based and hierarchical clustering algorithms to account for the semantic connections between legal phrases and concepts. We also offer a measure of similarity that computes the similarity of legal writings by combining word embeddings with legal ontologies. On a sizable dataset of legal documents, the suggested algorithm and similarity measure are assessed and contrasted with contemporary clustering techniques and similarity measures. The findings demonstrate that our method performs better than existing methods in terms of clustering accuracy and offers a more understandable clustering solution that might be helpful for researchers and legal professionals. Other fields that demand specialised text data processing methods can use the proposed algorithm and similarity metric.