Analysis of Extractive Text Summarization Algorithms on Hindi Legal Document Corpus
摘要
Understanding legal documents is a tough job for someone who has no relationship with the judiciary of the country. These documents are long, cumbersome and contain legal jargon. We aim to understand these documents and make them less complex for the common masses. To make this task easier we choose the Hindi Legal Document Corpus. We aim to work on the HLDC dataset which contains around 9 lakh documentation on Hindi legal documents. We implement 4 different algorithms to perform extractive summarization on Hindi legal documents. TextRank, LexRank, LSA, and Luhn algorithms are implemented to get a concise, clear summary. These algorithms are then compared and evaluation is done to choose the best algorithm. We implement Rouge to get the best-performing algorithm. This is done to provide easy access to understand legal documents and their jargon. This provides a better understanding of complex and tedious legal documents in our vernacular language.