Predicting Customer Churn in the Telecommunications Industry using Machine Learning Techniques
Adeline Makokha, Kevin Obote, Henry Muchiri, Kennedy Senagi
Strathmore University
阅读操作
确认中在文库中上传 PDF 后可生成中文音频讲解。
摘要与影响
Voluntary customer churn constitutes a persistent financial risk for telecommunications operators, particularly within enterprise customer segments where high-value accounts administer complex, multi-subscription portfolios. Industry data indicate that acquiring a new account costs between five and seven times more than retaining an existing one. Despite heightened industry awareness, the majority of operational retention platforms remain reactive, detecting departure only after the event has occurred. This investigation constructs and evaluates a machine learning pipeline engineered to identify enterprise customer churn risk proactively, drawing on authentic operational records extracted from a business-tobusiness telecommunications environment. The study follows the Cross-Industry Standard Process for Data Mining (CRISP-DM) lifecycle. A dataset of 8,454 unique business accounts, characterised by 14 raw attributes and enriched to a final 22-variable feature set, underpins the empirical work. Pronounced class imbalance, churned accounts representing approximately 6.5minority ratio of 14.3:1, necessitated specialised resampling prior to classifier training. Five oversampling strategies were benchmarked; SVMSMOTE produced the largest gain in minority-class sensitivity and was adopted for all subsequent training cycles. Ten classifier families were trained and assessed, including EasyEnsembleClassifier, RUSBoostClassifier, XGBoost, LightGBM, CatBoost, Histogram Gradient Boosting, Balanced Bagging, a multilayer perceptron, a soft-voting ensemble, and a stacking ensemble. EasyEnsembleClassifier emerged as the leading model, attaining an F1-score of 0.129 and a recall of 38.242 of 110 churned accounts. Post-hoc explainability analysis through SHAP and LIME identified active subscriber rate, geographic billing zone, and engineered interaction terms as the dominant predictive signals. The framework was operationalised within a FastAPI-based application supporting realtime individual scoring, batch CSV prediction, and retention campaign monitoring. The projected annual revenue protection under conservative assumptions exceeds 74,000 currency units. The study illustrates that interpretable, explainability-augmented machine learning frameworks can bridge the gap between quantitative model output and managerial action, offering a replicable blueprint for data-driven churn governance in both emerging and mature telecommunications markets.
逐年被引趋势
暂无年度引用数据
关键指标
同类平均 = 1
同领域 · 同年份 · 同类型
Google Scholar 与 OpenAlex 的被引统计范围不同,数值存在差异属正常。
AI 辅助阅读
依据:摘要
可就本文提问;依据不足时会说明。
学术脉络
学科主题
经济 / 管理Customer churn and segmentation
AI and HR Technologies · Imbalanced Data Classification Techniques
参考文献 10
此处列出前 3 条