Please use this identifier to cite or link to this item:
http://dspace.dtu.ac.in:8080/jspui/handle/repository/23151| Title: | CUSTOMER LIFETIME VALUE PREDICTION FOR RETAIL ANALYTICS: A COMPARATIVE STUDY OF MACHINE LEARNING MODELS |
| Authors: | AGGARWAL, ARNAV Sharma, Yogesh (SUPERVISOR) |
| Keywords: | RETAIL ANALYTICS CUSTOMER LIFETIME VALUE PREDICTION MACHINE LEARNING MODELS |
| Issue Date: | Sep-2026 |
| Series/Report no.: | TD-9232; |
| Abstract: | Modern retail organizations are confronted with an increasingly complex question at the heart of their marketing strategy: which customers deserve the greatest investment of time, money, and attention? The answer, in its most rigorous form, is found in Customer Lifetime Value a forward- looking estimate of the total financial contribution a customer is expected to make over the entire duration of their relationship with a firm. When computed accurately and acted upon intelligently, CLV transforms marketing from a cost center into a precision instrument for long-run revenue maximization. This dissertation investigates the prediction of Customer Lifetime Value in a retail e-commerce setting, employing four machine learning regression techniques Linear Regression, Random Forest, Gradient Boosting, and Multivariate Adaptive Regression Splines (MARS) and subjecting them to a rigorously designed empirical comparison. The dataset used is the publicly available Online Retail II record of transactions from a UK-based online gift-ware retailer spanning December 2009 through December 2011. Rather than computing features and targets from the same time window a common methodological error in the published literature this study enforces a hard temporal boundary at December 1, 2010, ensuring that all predictive features are derived exclusively from historical transactions while the prediction target is measured entirely from future ones. This design eliminates data leakage and produces performance estimates that are genuinely indicative of what each model would achieve in a live deployment. Customer behavior is represented through the Recency-Frequency-Monetary (RFM) framework. Recency measures how many days have elapsed since the customer's last purchase; Frequency counts the number of distinct transactions made in the historical window; Monetary Value records the total spending accumulated over that same period. These three variables, standardized via z- score normalization prior to modeling, form the entire feature set fed to all four models, ensuring that differences in performance are attributable to algorithmic properties rather than to differences in data preparation. The empirical results establish a clear ordering of predictive capability. Random Forest achieves the strongest out-of-sample performance, with a test-set coefficient of determination (R²) of 0.7244 and a Root Mean Squared Error of £3,727. MARS performs closely behind, reaching a test R² of 0.6910 and an RMSE of £3,947. Both models demonstrate meaningful and practically significant predictive power. Linear Regression, by contrast, collapses dramatically from training to testing achieving a training R² of 0.9766 but only 0.3022 on the test set revealing that the relationships between RFM v variables and future customer value are non-linear in ways that a linear model cannot represent. Gradient Boosting fails entirely under default hyperparameter settings, producing a test R² of −1.4598, a result worse than simply predicting the mean value for every customer, and serving as a sharp empirical warning against the uncritical application of complex algorithms without adequate configuration and tuning. Beyond the accuracy comparison, this study makes a substantive contribution through the interpretation of the MARS model. Because MARS represents its predictions as a set of explicit piecewise linear equations built from hinge functions at identified threshold values, the model's logic is fully transparent and directly actionable. The analysis extracts six customer segments whose boundaries are defined by data-derived threshold rules: Champions (7% of customers, average predicted CLV £10,855), Loyal Customers (9%, £5,164), At-Risk High-Value Customers (2%, £4,877), New Customers (9%, £776), Lost Customers (6%, £616), and a broad Others category (67%, £1,244). Each segment is accompanied by targeted marketing recommendations grounded in the behavioral characteristics that define it. A central argument of this dissertation is that the modest accuracy advantage of Random Forest over MARS roughly three percentage points of test R² does not straightforwardly justify choosing the less interpretable model in business and regulatory contexts. As algorithmic accountability frameworks such as the European Union's General Data Protection Regulation place growing legal obligations on firms to explain automated decisions, the ability to articulate a CLV model's logic in plain language becomes not merely a preference but a compliance requirement. MARS satisfies this requirement by construction; Random Forest requires approximation methods that are inherently less reliable. The dissertation concludes that CLV prediction methodology in retail analytics must simultaneously pursue three objectives predictive validity, methodological rigor, and practical interpretability and that MARS, properly applied within a temporally sound validation framework, provides a compelling solution to all three. |
| URI: | http://dspace.dtu.ac.in:8080/jspui/handle/repository/23151 |
| Appears in Collections: | MBA |
Files in This Item:
| File | Description | Size | Format | |
|---|---|---|---|---|
| Arnav Aggarwal BMBA.pdf | 739.15 kB | Adobe PDF | View/Open | |
| Arnav Aggarwal PLAG.pdf | 7.94 MB | Adobe PDF | View/Open |
Items in DSpace are protected by copyright, with all rights reserved, unless otherwise indicated.



