CASE STUDY / RETAIL & E-COMMERCE
Forecasting fresh demand across 400 grocery stores
SKU/store-level probabilistic forecasting driving automated replenishment — 28% fewer stockouts and 23% less fresh-category waste across 400 stores.
28%
reduction in stockouts on top-selling SKUs
23%
less fresh-category shrink, worth $17M annually
31%
forecast-accuracy improvement over the legacy system
92%
of order proposals accepted without adjustment by month six
THE CLIENT
Context
The client operates 400 grocery stores with a strong fresh and prepared-foods identity — the categories that drive loyalty and the ones hardest to forecast. Ordering combined a legacy replenishment system for center-store with manager gut-feel for fresh, produce, and bakery, where a wrong order becomes either an empty shelf or a shrink write-off within 48 hours.
The numbers told the story: fresh shrink ran 9.4% of category sales against a 6% industry benchmark, while availability audits found top-selling fresh SKUs out of stock in 1 of 8 store visits. Two floors of the same building were losing money in opposite directions.
THE CHALLENGE
What was at stake
Fresh forecasting is unforgiving: demand swings with weather, local events, promotions, and day-of-week patterns that differ store by store; short shelf lives mean forecast errors can't be buffered with safety stock; and new products launch weekly with no sales history. The legacy system forecast at category/week level — useless for deciding how many rotisserie chickens store #217 needs on a rainy Tuesday.
Equally hard: 400 store managers who had been ordering fresh by experience for years would receive system-generated order proposals. If early proposals were visibly wrong, the program would be dead regardless of aggregate accuracy.
Engagement at a glance
- Client
- A top-10 regional grocery chain
- Region
- United States
- Duration
- 8 months
- Team
- 8-person team: ML engineers, data platform, supply-chain analyst, change lead
Services applied
THE SOLUTION
What we built
We built a probabilistic forecasting platform generating SKU/store/day demand distributions — not point estimates — for 38,000 SKUs across 400 stores nightly. Gradient-boosted models handle established products using weather forecasts, promotion calendars, holidays, local events, and store-specific seasonality; similarity-based cold-start models cover new products by matching them to demand-alike predecessors. Order proposals are then optimized against each SKU's economics: shelf life, case pack sizes, margin, and the asymmetric cost of over- versus under-stocking fresh items.
Store managers kept control by design. Proposals arrive in a purpose-built ordering app that shows the reasoning — expected demand curve, weather effect, promo lift — and managers approve or adjust by exception. Adjustments feed back into the models, and a weekly accuracy scorecard per store made the system's improvement visible to the people using it.
// ARCHITECTURE
Nightly pipelines in Databricks process POS, inventory, promotion, and third-party weather/events data into a feature store; LightGBM quantile models generate demand distributions, and an optimization layer converts them into order proposals under supplier lead-time and truck-schedule constraints. Forecast serving and the store ordering app run on GCP with offline-tolerant sync for stores with weak connectivity.
MLflow manages weekly retraining with automated backtesting gates — no model ships unless it beats the incumbent on the trailing eight weeks.
Core stack
- Databricks
- LightGBM
- BigQuery
- dbt
- MLflow
- GCP
- Kafka
- React Native (store app)
- Airflow
HOW IT WAS DELIVERED
Implementation approach
Value delivered in phases with go/no-go evidence at each gate — never a big-bang bet.
- 01
Backtest diagnostic (weeks 1–6)
Trained candidate models on two years of history and proved a 31% forecast-accuracy improvement over the legacy system before build commitment.
- 02
Data platform (months 2–4)
Unified POS, inventory, promotions, and external signals; repaired perpetual-inventory accuracy at store level, the silent killer of replenishment systems.
- 03
20-store pilot (months 4–6)
Live ordering proposals in 20 stores across three regions, with manager feedback sessions every week and the accuracy scorecard published openly.
- 04
Chain-wide rollout (months 6–8)
Cohort rollout of 60–80 stores per fortnight with train-the-trainer enablement through regional managers.
“We'd been told for years that fresh couldn't be forecast — it was 'an art.' Turns out it was a data problem. Our store managers are the system's biggest advocates now, because it made them better at the part of the job they're judged on.”
Tom Kowalczyk
EVP, Merchandising & Supply Chain — Top-10 regional grocery chain
KEEP READING
More case studies
Want results like these?
Bring us the metric you need to move. Our architects will map how a comparable engagement would work in your environment — systems, timeline, and expected impact.
