In the e-commerce sector, identifying the weights of factors influencing cancellation rates and forecasting these rates in advance can enable businesses to make more informed and data-driven decisions. In this study, real sales data from an e-commerce company operating in Turkey and engaged in the sale of fashion jewelry and silver accessories on major online marketplaces was analyzed. The weights of the factors affecting product cancellation rates were determined using XGBoost, CatBoost Feature Importance, Permutation Importance, and SHAP (Shapley Additive Explanations) methods. Additionally, the performances of various machine learning algorithms—such as boosting models, Decision Tree, and Support Vector Regressor—were compared in terms of their ability to predict cancellation rates. According to the results, the CatBoost model achieved the highest performance across all metrics, providing the most accurate predictions with an R2 score of 0.9986. Based on feature importance analyses, the variable Customer_Cancelled_Order_Quantity was identified as the most influential feature by SHAP, CatBoost, and Permutation Importance methods, whereas the Net_Sales_Quantity variable was found to be the most significant according to the XGBoost model. The findings suggest that in order to reduce order cancellations, pricing strategies should be optimized, inventory management should be strengthened, and customer-focused processes should be improved.

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Analysis of Order Cancellation Rates and Feature Weighting with Machine Learning: An E-Commerce Case Study

  • Tuba Irmak

摘要

In the e-commerce sector, identifying the weights of factors influencing cancellation rates and forecasting these rates in advance can enable businesses to make more informed and data-driven decisions. In this study, real sales data from an e-commerce company operating in Turkey and engaged in the sale of fashion jewelry and silver accessories on major online marketplaces was analyzed. The weights of the factors affecting product cancellation rates were determined using XGBoost, CatBoost Feature Importance, Permutation Importance, and SHAP (Shapley Additive Explanations) methods. Additionally, the performances of various machine learning algorithms—such as boosting models, Decision Tree, and Support Vector Regressor—were compared in terms of their ability to predict cancellation rates. According to the results, the CatBoost model achieved the highest performance across all metrics, providing the most accurate predictions with an R2 score of 0.9986. Based on feature importance analyses, the variable Customer_Cancelled_Order_Quantity was identified as the most influential feature by SHAP, CatBoost, and Permutation Importance methods, whereas the Net_Sales_Quantity variable was found to be the most significant according to the XGBoost model. The findings suggest that in order to reduce order cancellations, pricing strategies should be optimized, inventory management should be strengthened, and customer-focused processes should be improved.