Case Study · Manufacturing
Kale Seramik: a demand forecasting PoC — and why we did not continue
An MVP that started on the same roadmap as Kalekim: two product functions, 34 methods, three encoding approaches. The result fell short of the target and the second phase was not started. This page explains what did not work, and why.
- 141
- DFUs
- 200+
- Product attributes
- 34
- Methods tried
- 1 / 3
- Phases completed
29 of them carry 80% of sales
Most sparsely populated
Encoding × imputation × feature combinations
Phase 2 not started
Problem
Like Kalekim, Kale Seramik wanted a more accurate demand forecast as the input to production planning. The roadmap drawn up with Kale Group in early 2023 was the same for both companies — MVP, automated forecasting process, forecasts managed through screens — but because their data and structures differ, they were run as two separate projects.
Kale Seramik chose two functions: wall tiles (SKM – main) and bathroom – fittings and sanitaryware, with all sub-groups. The forecast was requested at sales-office level on the customer side and DFU level on the product side.
What we did
Data came from Excel: monthly sales from 2019 to mid-2023, a material characteristics table with more than 200 columns, the material-to-DFU mapping and external factors (exchange rate, natural gas, special days). Part of the sales data arrived with months as columns; it was reshaped into rows. Out-of-scope and deleted-material flags were filtered. Missing attributes were filled as “unknown” or with the most frequent value within the DFU.
Modelling used the same framework as Kalekim: three methods for carrying material attributes up to DFU level (M1 most frequent value, M2 counts and ratios, ME target encoding), XGBoost and LightGBM under FLAML AutoML. Combinations of imputation method, encoding method, hierarchy encoding and feature selection produced 34 different methods; results were generated for each of eight months and measured with sales-weighted MAPE.
What we found
The results were not acceptable at the level the forecast was made. Some months looked good, but there was no consistency across months; three methods stood slightly apart from the rest, and that was all. M2 was the best method; M1 and ME did not contribute.
We documented the reasons:
- The long tail. 29 of 141 DFUs carry 80% of sales. The remaining 112 sell sparsely and irregularly; at DFU × sales-office level that tail leaves no repeating pattern.
- Attributes exist but are not populated. Of more than 200 attribute columns, most are sparsely filled. The attributes that separate fast sellers from slow sellers were not sufficiently present in the data; this is where the limit of predicting a numeric target from categorical attributes showed.
- External factors were not enough. Emphasis on external drivers had been agreed at the outset, but the number of factors that made it into the model could not explain the outcome.
- Structural breaks. The economic swings of the last two years and the 2023 earthquake produced movements that historical data had not taught.
- The hierarchy dilemma. Going down, data grows but repeatability falls; going up, the reverse — but carrying row-level external factors up a level requires an assumption, and that assumption can mislead the model.
Why we stopped
The roadmap was designed to allow it: whether to proceed to Phase 2 was Kale Seramik’s decision, based on Phase 1’s output. That was exactly the purpose of the MVP — to show the feasibility and likelihood of success of a forecasting system before a large investment. The results did not justify a second phase with this setup; the project was left there.
In the document we delivered we wrote what continuing would require: defining the attribute set that could drive the forecast and populating it at material level to a large extent, and trying forecasts in different customer hierarchies for different product hierarchies.
The lesson
Same roadmap, same methods, same team: at Kalekim it grew into a three-year system, at Kale Seramik it stayed at one phase. The difference was set by the data, not the method — how many products the sales concentrate in, how well the product attributes are populated, whether the forecast level leaves a repeating pattern.
That is what the MVP is for. The cheapest failure of a forecasting project is the one measured and stopped in its first phase.