Multiple questions with the same intent can cause a lot of wastage of time to the active readers because they will have to spend a lot of time trying to find information to their questions and the best possible answer to that question, this also creates confusion to the reader since there are multiple versions of the same answer. This paper provides a solution to this problem by identifying the duplicate questions on social media. The paper is an empirical comparison of a random model, Logistic Regression, Linear Support Vector Machine (SVM) and XGBoost Machine Learning (ML) algorithms. The experiment shows improved accuracy using Matthew's Correlation Score (MCC) metric with the XGBoost algorithm.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Empirical Comparison of Machine Learning Algorithms to Identify Duplicate Questions on Various Social Media Forums

  • Karthik Sridhar,
  • Pranita Mahajan

摘要

Multiple questions with the same intent can cause a lot of wastage of time to the active readers because they will have to spend a lot of time trying to find information to their questions and the best possible answer to that question, this also creates confusion to the reader since there are multiple versions of the same answer. This paper provides a solution to this problem by identifying the duplicate questions on social media. The paper is an empirical comparison of a random model, Logistic Regression, Linear Support Vector Machine (SVM) and XGBoost Machine Learning (ML) algorithms. The experiment shows improved accuracy using Matthew's Correlation Score (MCC) metric with the XGBoost algorithm.