<p>Customer churn prediction is an essential research topic in the telecommunications industry, as the market is becoming saturated, fierce competition is prevalent, and the financial cost of losing customers is high. It is much cheaper to retain existing customers than to acquire new ones, and therefore, predictive analytics is a strategic requirement for telecom operators. Churn prediction methodologies have developed significantly in the last 20&#xa0;years, moving beyond traditional statistical modeling techniques to state-of-the-art artificial intelligence-based systems with the ability to scale to large volumes and high dimensions of telecom data and time-warped telecom data. This is a review paper that gives a comprehensive, systematic analysis of the state-of-the-art techniques that are used in predicting telecom churn. The systematic search was carried out in three major academic databases (Scopus, ACM Digital Library, IEEE Xplore) from January 2010 to December 2024. The Boolean search string used was (telecom* OR telecommunication* OR mobile) AND (churn OR attrition OR "customer retention") AND (predict* OR classif* OR "machine learning" OR "deep learning"). Studies were selected if they (i) introduced or tested a model that could predict customer churn in a telecom firm, (ii) provided quantitative performance indicators, and (iii) were published in peer-reviewed sources. Non-English publications, as well as books and patents, were not included. After full-text review, 199 of 299 candidate studies were included. The research summarizes the results of benchmark datasets, including Orange, Cell2cell, IBM/Kaggle Telecom datasets, and actual operator billing and usage data. It shows a critical overview of classical machine learning models such as the Logistic Regression, Decision Trees, Support Vector Machines, Random Forests, and Gradient Boosting before proceeding to deep learning architectures such as the Convolutional Neural Network (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM) networks, Autoencoders, and hybrid deep machine learning models. Besides, the review points out to new evolutionary computation methods, such as genetic programming-based models, which focus on automated feature evolution and cost-sensitive optimization. Among the main problems discussed in the literature, there are the imbalance in classes, high-dimensional information about customer behavior, a model of temporal dependency, asymmetry of misclassification cost, constraints in interpretability, and the problem of scalability in the real-time telecom setting. Performance metrics employed in the study to predict churn, including Accuracy, Precision, Recall, F1-score, AUC-ROC, G-mean, cost-sensitive metrics, and profit-based ones, are also extensively evaluated, with limitations of the accuracy-focused evaluation of imbalanced and business-critical problems highlighted. The results demonstrate a distinct change in methodology to hybrid, cost-conscious, and explainable models that optimize the goal of predictive modelling with telecom revenue optimization policies. Deep learning models are better at extracting nonlinear features, whereas the ensemble and evolutionary methods are more robust and financially aligned. The review concludes with future research directions and with a discussion on explainable AI, federated learning, multi-objective optimization, graph-based churn models, and research protocols based on profit. The piece suggests a systematic approach to designing smart, scalable, and economically viable churn prediction systems for today's telecom environment.</p>

错误:搜索内容不能为空,请输入英文关键词
错误:关键词超出字数限制,请精简
高级检索

Telecom customer churn prediction: a detailed systematic review of data, methodology, and evaluation methods

  • Hemlata Jain,
  • Ajay Khunteta,
  • Nikita Sharma

摘要

Customer churn prediction is an essential research topic in the telecommunications industry, as the market is becoming saturated, fierce competition is prevalent, and the financial cost of losing customers is high. It is much cheaper to retain existing customers than to acquire new ones, and therefore, predictive analytics is a strategic requirement for telecom operators. Churn prediction methodologies have developed significantly in the last 20 years, moving beyond traditional statistical modeling techniques to state-of-the-art artificial intelligence-based systems with the ability to scale to large volumes and high dimensions of telecom data and time-warped telecom data. This is a review paper that gives a comprehensive, systematic analysis of the state-of-the-art techniques that are used in predicting telecom churn. The systematic search was carried out in three major academic databases (Scopus, ACM Digital Library, IEEE Xplore) from January 2010 to December 2024. The Boolean search string used was (telecom* OR telecommunication* OR mobile) AND (churn OR attrition OR "customer retention") AND (predict* OR classif* OR "machine learning" OR "deep learning"). Studies were selected if they (i) introduced or tested a model that could predict customer churn in a telecom firm, (ii) provided quantitative performance indicators, and (iii) were published in peer-reviewed sources. Non-English publications, as well as books and patents, were not included. After full-text review, 199 of 299 candidate studies were included. The research summarizes the results of benchmark datasets, including Orange, Cell2cell, IBM/Kaggle Telecom datasets, and actual operator billing and usage data. It shows a critical overview of classical machine learning models such as the Logistic Regression, Decision Trees, Support Vector Machines, Random Forests, and Gradient Boosting before proceeding to deep learning architectures such as the Convolutional Neural Network (CNN), Recurrent Neural Networks (RNN), Long Short-Term Memory (LSTM) networks, Autoencoders, and hybrid deep machine learning models. Besides, the review points out to new evolutionary computation methods, such as genetic programming-based models, which focus on automated feature evolution and cost-sensitive optimization. Among the main problems discussed in the literature, there are the imbalance in classes, high-dimensional information about customer behavior, a model of temporal dependency, asymmetry of misclassification cost, constraints in interpretability, and the problem of scalability in the real-time telecom setting. Performance metrics employed in the study to predict churn, including Accuracy, Precision, Recall, F1-score, AUC-ROC, G-mean, cost-sensitive metrics, and profit-based ones, are also extensively evaluated, with limitations of the accuracy-focused evaluation of imbalanced and business-critical problems highlighted. The results demonstrate a distinct change in methodology to hybrid, cost-conscious, and explainable models that optimize the goal of predictive modelling with telecom revenue optimization policies. Deep learning models are better at extracting nonlinear features, whereas the ensemble and evolutionary methods are more robust and financially aligned. The review concludes with future research directions and with a discussion on explainable AI, federated learning, multi-objective optimization, graph-based churn models, and research protocols based on profit. The piece suggests a systematic approach to designing smart, scalable, and economically viable churn prediction systems for today's telecom environment.