CASE STUDY / MANUFACTURING
Predicting failures before they stop the line
Edge-deployed predictive maintenance across 9 plants and 1,100 assets — 32% less unplanned downtime and a 9-point OEE gain inside a year.
32%
reduction in unplanned downtime on covered assets
9pt
OEE improvement across the nine plants
86%
prediction precision on top failure modes
11 mo
program payback, verified against maintenance and production baselines
THE CLIENT
Context
The client manufactures heavy industrial equipment across nine plants on three continents. Machining centers, presses, and paint-line assets ran on calendar-based preventive maintenance, yet unplanned downtime still consumed 11% of scheduled production time — with a single constrained machining line in Monterrey costing an estimated $140K per lost shift.
Plant data existed but was stranded: 20 years of historian tags, maintenance logs in three CMMS systems and two languages, and vibration routes collected monthly on clipboards. Two prior 'Industry 4.0' pilots had produced dashboards but no changed maintenance decisions.
THE CHALLENGE
What was at stake
The failure modes that hurt most — spindle bearing degradation, hydraulic press seal failures, paint-line conveyor faults — develop over weeks but were being caught in hours. Detecting them required continuous high-frequency vibration and process data the monthly clipboard routes could not provide, plus models tuned per asset class rather than generic anomaly scores that maintenance planners had learned to ignore.
Operationally, the constraint was trust and workflow: predictions had to arrive as prioritized, evidence-backed work orders inside the CMMS that planners already used — with enough diagnostic context that a technician knew what to inspect — or they would join the previous pilots as ignored dashboards.
Engagement at a glance
- Client
- A global industrial-equipment manufacturer
- Region
- US, Mexico & Germany
- Duration
- 12 months
- Team
- 9-person team: OT integration, ML engineers, edge platform, reliability engineer
Services applied
THE SOLUTION
What we built
We instrumented 1,100 critical assets with continuous vibration and current sensing where needed, streaming alongside existing PLC and historian tags through a unified namespace into plant-level edge clusters. Per-asset-class models — spectral anomaly detection for rotating equipment, sequence models for hydraulic cycles — run at the edge, so detection never depends on WAN connectivity.
Detections become CMMS work orders automatically, ranked by predicted time-to-failure and production impact, each carrying the evidence: trend charts, frequency signatures, and the closest historical matches from digitized maintenance logs. A reliability-engineering feedback loop closes every work order with a confirmed/refuted disposition, which retrains the models and — just as importantly — published a running precision score that earned planner trust plant by plant.
The same data foundation delivered accurate OEE and loss attribution across all nine plants for the first time, which the COO now reviews weekly — turning the maintenance program into the beachhead for a broader production-intelligence platform.
// ARCHITECTURE
Edge layer: OPC UA and MQTT/Sparkplug B connectivity into K3s clusters per plant running ONNX-exported models with local buffering. Cloud layer: AWS-based lakehouse consolidating all plants for cross-fleet model training, OEE analytics, and fleet-level monitoring. CMMS integration via APIs into Maximo and SAP PM, with bidirectional work-order status sync.
Model lifecycle runs through MLflow with per-asset-class registries; updates deploy to plants over the air during maintenance windows with automatic rollback.
Core stack
- OPC UA
- MQTT / Sparkplug B
- K3s edge clusters
- PyTorch
- ONNX Runtime
- TimescaleDB
- AWS IoT SiteWise
- Databricks
- MLflow
- Grafana
HOW IT WAS DELIVERED
Implementation approach
Value delivered in phases with go/no-go evidence at each gate — never a big-bang bet.
- 01
Criticality and loss mapping (months 1–2)
Ranked 4,000+ assets by downtime cost and failure history; scoped the program to the 1,100 that drive 85% of loss.
- 02
Monterrey lighthouse (months 3–6)
Full deployment at the highest-cost plant: sensing, edge platform, models, and CMMS integration, validated against real failure events.
- 03
Model hardening (months 5–8)
Closed-loop disposition feedback pushed precision above 85% on the top failure modes before any fleet expansion.
- 04
Fleet rollout (months 7–12)
Templated deployment kits brought the remaining eight plants live, with cross-fleet transfer learning accelerating each new site.
“We'd bought dashboards twice before. Ilmora gave us work orders — with evidence attached — inside the system our planners already lived in. When the Monterrey spindle prediction saved us a full weekend of lost production, the plant managers stopped asking why and started asking when.”
Ingrid Baumann
VP, Global Manufacturing Operations — Global industrial-equipment manufacturer
KEEP READING
More case studies
Want results like these?
Bring us the metric you need to move. Our architects will map how a comparable engagement would work in your environment — systems, timeline, and expected impact.
