Please use this identifier to cite or link to this item:
http://dspace.dtu.ac.in:8080/jspui/handle/repository/23172| Title: | PREDICTING CUSTOMER CHURN USING MACHINE LEARNING TECHNIQUES: A CASE STUDY ON TELECOM CUSTOMERS |
| Authors: | KUMAR, AASHISH Sharma, Yogesh (SUPERVISOR) |
| Keywords: | PREDICTING CUSTOMER CHURN MACHINE LEARNING TECHNIQUES TELECOM CUSTOMERS SVMSMOTE |
| Issue Date: | Sep-2026 |
| Series/Report no.: | TD-9261; |
| Abstract: | The telecommunications industry faces one of the most persistent and financially consequential challenges in customer management: the prediction and prevention of customer churn. Churn, defined as the voluntary discontinuation of a subscription service, directly impacts revenue stability, inflates customer acquisition costs, and undermines long-run profitability. Industry research consistently estimates that acquiring a new customer costs five to seven times more than retaining an existing one, making every prevented defection a measurable financial contribution to the firm's bottom line. As network infrastructure becomes increasingly commoditized and service offerings difficult to differentiate on technical grounds alone, the strategic focus of leading telecom operators has shifted decisively from acquisition to retention, placing churn prediction at the center of competitive customer management strategy. This dissertation investigates the application of machine learning classification techniques to telecom customer churn prediction using the IBM Telco Customer Churn dataset sourced from Kaggle, comprising 7,043 customer records and 21 features spanning demographic attributes, account characteristics, service configuration variables, and financial metrics. Three classifiers — Logistic Regression, Decision Tree, and XGBoost — are evaluated under both imbalanced baseline conditions and three SMOTE-based resampling configurations: Standard SMOTE, Borderline-SMOTE, and SVMSMOTE. Systematic hyperparameter tuning using RandomizedSearchCV with five-fold stratified cross-validation is applied to optimize each configuration for minority-class detection performance. Model performance is assessed using five metrics — Accuracy, Precision, Recall, F1-score, and ROC-AUC — with Recall and F1-score designated as primary evaluation criteria, consistent with the operational objective of maximizing the identification of genuine at-risk customers. Exploratory data analysis identifies contract type, customer tenure, internet service type, and value-added service subscriptions as the primary structural drivers of churn propensity. v Month-to-month customers churn at 42.7 percent compared to only 2.8 percent for two-year contract customers — a ratio exceeding fifteen-to-one that underscores the enormous protective effect of long-term contractual commitment. Newly acquired customers in their first twelve months of service exhibit the highest attrition risk, while subscribers to value added services such as online security, technical support, online backup, and device protection consistently show lower churn rates across all customer segments, consistent with the switching cost hypothesis that each additional service increases the complexity and cost of switching to a competitor. Senior citizen customers exhibit disproportionately elevated churn rates of approximately 41.7 percent compared to 23.6 percent for non-senior customers, identifying them as a structurally distinct high-risk segment warranting dedicated retention attention. Baseline model results confirm the severity of the class imbalance problem. All three classifiers trained on the raw imbalanced dataset achieved accuracy values near 0.80 — figures that are almost entirely explained by the correct classification of the dominant non churn majority rather than genuine predictive capability. Recall values for the minority churn class ranged from only 0.497 to 0.550, meaning that between 45 and 50 percent of actual churners were misclassified as non-churners. A deployed model performing at this level would fail to flag nearly half of all customers who subsequently leave, rendering any retention program built on its outputs severely incomplete and operationally ineffective. The application of SMOTE-based resampling produced substantial and consistent improvements in minority-class performance across all three classifiers, with Recall gains ranging from 18 to 29 percentage points depending on the classifier-resampling combination. Among all twelve model-resampling configurations evaluated, XGBoost combined with SVMSMOTE achieved the strongest overall performance — an F1-score of 0.636, Recall of 0.758, and ROC-AUC of 0.845. Compared to the baseline XGBoost model, the SVMSMOTE-balanced configuration reduced missed churners by 52 percent, from 235 to 113 False Negatives, while increasing correctly identified churners from 232 to 354 True Positives. The increase in False Positives from 115 to 293 represents an operationally acceptable trade-off, as the cost of incorrectly targeting a non-churner with a retention offer is substantially lower than the cost of failing to intervene with a genuine at-risk customer. The study concludes with five principal findings: class imbalance is a genuine and consequential methodological problem that accuracy-focused evaluation cannot adequately capture; SMOTE-based resampling — SVMSMOTE in particular — is a mandatory rather than optional component of any production churn prediction pipeline; XGBoost with vi SVMSMOTE is the recommended configuration for deployment, delivering the strongest F1 score, most stable ROC-AUC, and greatest reduction in missed churners; Logistic Regression with SVMSMOTE represents a competitive interpretability-friendly alternative for organizations with regulatory or stakeholder transparency requirements; and contract type combined with early tenure represents the highest-value target for proactive retention intervention. Together, these findings provide both a validated analytical framework and actionable strategic guidance for telecom operators seeking to implement effective, data driven customer retention programs. |
| URI: | http://dspace.dtu.ac.in:8080/jspui/handle/repository/23172 |
| Appears in Collections: | MBA |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| Aashish Kumar mba.pdf | 1.09 MB | Adobe PDF | View/Open | |
| Aashish Kumar plag.pdf | 21.95 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



