Case Study · Retail
Defacto Norm Staffing: calculating the staff each store should have, from data
A model that takes the store staffing decision away from actual payroll and subjective requests and grounds it in measured workload and an efficient peer in the same segment. It produces a 2025 norm, a 2026 target and an increase / reduce / keep decision for every store × role — end to end on BigQuery, fully parametric.
- Store × role
- Decision unit
- Frontier
- Reference
- Gradual
- Approach
- Increase · reduce · keep
- Output
Sales · Stockroom · Management
The segment's efficient peer that also hits its target
No jump to the reference in one step
Per-store decision for 2026
Problem
How many people does a store need? In most retailers the answer is a continuation of the past: however many there were last year, plus or minus a little. The staffing request comes from the store manager, the budget comes down from above, and the two meet somewhere in the middle. Nobody can show with a number which store is genuinely short-staffed and which is running with too many.
What Defacto wanted was to make the staffing decision from data, not from human judgement: the staff each store should have for each role, calculated from the volume of work that store does and from the most efficient peer store in the same segment; and from that result, an increase / reduce / keep decision.
The difficulty lies in defining “same segment” and “efficient” correctly. Comparing a high-revenue mall store with a small high-street store is meaningless; taking the store with the fewest people as the benchmark would reward understaffing.
The effort layer — how many hours is the workload?
The model starts from the store operations effort standard: a unit time for every task item. Sales and returns, goods receipt, unpacking, tagging, transfers, stock counts, till closing, label changes, membership sign-ups, omnichannel tasks (ship-from-store, click and collect), visual merchandising, training, management… Each task’s annual hours = volume × unit time.
Those hours are distributed to stores by driver share: transaction count, inbound waybills, labels printed, till closings, membership sign-ups — each task is tied to the driver that best represents its volume. Tasks not tied to volume are distributed either equally per store (management, daily meeting, visual merchandising) or by operational share. Every task is split into Sales, Stockroom and Management by role weights; for mixed tasks (unpacking, stock counts, transfers) the weights were set together with store operations.
The effort layer has its own validation: how closely does the model’s effort track actual headcount? This “quality gate” stops the model from moving to the next step until it is sound.
Segmentation — compare apples with apples
Stores are clustered by hierarchical dimensions: revenue › velocity segment › m² › store type › location › floors. The narrowest cluster consists of stores sharing all six dimensions. If a cluster has too few stores, the weakest dimension is dropped and the search moves to a wider cluster — revenue is always kept. Every store thus gets a statistically sound peer group.
Frontier — a reference with two gates
Each cluster’s reference store passes two filters:
- Performance gate: budget / target attainment must be within the segment’s band. A structurally troubled store, or one that looks “lean” because it is understaffed, cannot be the reference.
- Efficiency gate: among those candidates, the store working with the fewest FTE per unit of revenue.
The reference is not a headcount but an efficiency ratio: the frontier’s FTE/revenue is multiplied by each store’s own revenue. Without this scaling, low-revenue stores would be forced to carry a large store’s staffing; scaling only kicks in when the revenue gap widens.
Convergence — the heart of the model
Equating every store directly to the frontier is too aggressive. Instead, each store is moved gradually toward its scaled reference: part of the gap between current staffing and the reference is closed, not all of it. The result is a realistic, actionable 2025 norm — per store × role.
Edge cases are rule-based: no norm is proposed for a role the store does not have at all (assumed to be handled centrally); room-sized stores stay out of the frontier pool; stores without a full year are annualised by monthly run-rate and flagged for manual review.
2026 — the multiplier
There is no re-convergence for 2026. The ratio between “raw effort” and “converged norm” in 2025 is stored as an efficiency multiplier for each store × role; the 2026 workload forecast is converted to effort and multiplied by it. The norm carries the efficiency level; effort carries the year-on-year change in the work. The result is compared store by store with the company’s 2026 hours budget: where the budget is below the reference, “aggressive / understaffing risk”; where above, “room to spare”.
Data decisions
The model’s reliability rests on a few honest data decisions. Because time-and-attendance data for 2025 had incomplete coverage and excluded support staff, payroll was taken as authoritative for actual hours. Stores showing zero stockroom staff were found to have stockroom work booked under sales in payroll and were corrected using the fleet ratio. 2026 drivers that had undergone process or coverage changes (footfall, put-away, outbound waybills) were left out of the model, with reliable drivers used as proxies instead. For label data, two sources were compared and one chosen.
Parametric structure and version discipline
The model was built end to end on BigQuery as step-by-step SQL, driven by parameters from a single config: monthly net hours (the FTE divisor), the convergence coefficient, the performance-band scenario, frontier scaling on/off, the cluster threshold, the segment hierarchy. Every improvement request that followed a presentation — a change of divisor, role mappings, removal of calibration, a budget gate on the frontier — was applied as a parameter change, and every version was frozen as a checkpoint. Alternative requests such as a “direct frontier” were answered with comparison rather than persuasion: two outputs from the same data, side by side.
The output package: one master table per store with segment, cluster, frontier, budget attainment, 2025 (actual / model / frontier) and 2026 (norm / budget / proposal) FTE, plus combined, proposal and segment-detail Excel files for management, and methodology and data-dictionary documents.
Status
Phase 1 was presented to management; eight improvements were applied afterwards and Phase 2 is under way. The HR analytics roadmap defines two continuations of the model: an adaptation for international stores, calibrated per country, and a cold-start staffing estimate for stores not yet opened. The norm gives what is “required”; the surplus/shortfall analysis against actual staffing is a separate piece of work.
This project is the backbone of the HR analytics engagement with Defacto.
How the model was built
| Objective | Direction | Tension |
|---|---|---|
| Required staff (FTE) per store × role | ↓ | While moving toward the efficient peer, a store that misses its target cannot be the reference — leanness caused by understaffing does not count |
| Reference efficiency (FTE / revenue) | ↓ | The narrower the cluster, the fewer the peers; without enough peers the search moves to a wider segment |
Constraints
- Performance gate — a store whose budget / target attainment is outside the segment's band cannot be the frontier
- If a cluster is too small, the hierarchy moves up one level (revenue is the most protected dimension)
- The frontier is an efficiency ratio, not a headcount; it is scaled by each store's own revenue
- Gradual convergence — part of the gap is closed, giving an achievable target
- No norm is proposed for a role the store does not have (the work is assumed to be handled centrally)
- Very small "room" stores stay out of the frontier pool and follow a separate rule
- Stores without a full year are annualised by run-rate and flagged for manual review
- No re-convergence for 2026; 2025's efficiency level is carried forward with a multiplier