In contemporary Indian parliamentary debates, understanding argumentation patterns is crucial for practical policy analysis and decision-making. Exploring Indian Parliamentary debates helps grasp their linguistic characteristics, rhetorical strategies, and political implications. Analysing speeches aims to understand how language is used to persuade, argue, and negotiate in this context, shedding light on the dynamics of Indian democracy and parliamentary discourse. This research paper highlights the use of natural language processing to comprehend the hidden patterns in Lok Sabha’s speeches, aka the House of People, which were collected from its official website. 5000 speeches were extracted, and a subset was manually annotated into five categories—Appreciate, Neutral, Call For Action, Blame, and Issue. The data was preprocessed to determine the sentiment, key phrases, similarity, gender distribution, and category classification. Multiple baseline machine learning models were implemented, such as logistic regression, Naive Bayes, SGD classifiers, and deep neural networks. The logistic regression model gave the best performance with an accuracy of 55%.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Exploring Argumentation in Indian Parliamentary Debates Using NLP

  • Samya Jain,
  • Kashika Akhouri,
  • Punerva Singh,
  • Tanya Chhikara,
  • Rahul Sachdeva

摘要

In contemporary Indian parliamentary debates, understanding argumentation patterns is crucial for practical policy analysis and decision-making. Exploring Indian Parliamentary debates helps grasp their linguistic characteristics, rhetorical strategies, and political implications. Analysing speeches aims to understand how language is used to persuade, argue, and negotiate in this context, shedding light on the dynamics of Indian democracy and parliamentary discourse. This research paper highlights the use of natural language processing to comprehend the hidden patterns in Lok Sabha’s speeches, aka the House of People, which were collected from its official website. 5000 speeches were extracted, and a subset was manually annotated into five categories—Appreciate, Neutral, Call For Action, Blame, and Issue. The data was preprocessed to determine the sentiment, key phrases, similarity, gender distribution, and category classification. Multiple baseline machine learning models were implemented, such as logistic regression, Naive Bayes, SGD classifiers, and deep neural networks. The logistic regression model gave the best performance with an accuracy of 55%.