AI LABS
Where frontier research becomes deployable
AI Labs is Ilmora's applied research group — 30 researchers and engineers, led by VP of AI Research Lena Vogel, who pressure-test emerging techniques on realistic enterprise workloads. What survives becomes a pattern, an accelerator, or a product. What doesn't saves our clients from finding out the expensive way.
THE LAB
An engineering lab, not an ivory tower
Every Labs experiment starts with a question from a live engagement: why do agent workflows fail unrecoverably, why do extraction pipelines rot when a form changes, what does an auditor need to trust an autonomous decision. Researchers pair with delivery engineers, run the experiment against realistic (or synthetic-but-honest) enterprise data, and publish the result internally whether it worked or not.
Experiments that clear the bar graduate into practice patterns — reference implementations with evaluation suites attached — which is how a technique gets from arXiv to a regulated production environment in months instead of years. Roughly one in three graduates; the rest are documented failures we never bill a client to rediscover.
FOCUS AREAS
Three problems worth a decade
We keep the lab narrow on purpose — three areas, chosen because they decide whether enterprise AI can be trusted at all.
Agentic systems
Reliability patterns for autonomous and semi-autonomous agents: orchestration topologies, permission models, failure containment, and the audit trails regulators actually accept. Everything our agentic practice ships started as a Labs pattern.
- Multi-agent orchestration
- Tool-use permissioning
- Recovery & rollback
- Action audit trails
Multimodal intelligence
Document, image, and voice understanding for enterprise workflows — extraction that survives messy scans, layout-aware reasoning, and speech pipelines for regulated call environments. The engine room behind Ilmora Lens.
- Layout-aware extraction
- Vision-language grounding
- Speech & call analytics
- Synthetic training data
Evals & safety
The discipline that makes the other two shippable: regression-gated evaluation harnesses, domain-specific benchmarks, red-teaming playbooks, and drift monitoring that catches behavior change before users do.
- Regression-gated evals
- Domain benchmarks
- Red-team playbooks
- Behavioral drift monitoring
EXPERIMENT LOG
On the bench right now
A sample of the current log — codenames real, details lightly redacted. Clients get the full log, including the failures.
Supervisor trees for agent fleets
Borrowed Erlang's supervision model for multi-agent systems: a hierarchy of watchdog agents that restart, quarantine, or escalate failing workers. Cut unrecoverable workflow failures 83% in pilot; now standard on agentic engagements.
Eval-set distillation from production traces
Automatically mining anonymized production traces for the hardest 2% of cases and compiling them into regression suites. Early results: eval sets that predict incident rates 4x better than hand-curated ones.
Layout-blind document extraction
Testing whether structured extraction can stay accurate when trained with aggressive layout randomization — the goal is models that survive the form redesigns that silently break production pipelines.
Cost-aware model routing
A router that picks the cheapest model predicted to clear the task's quality bar, learning from eval outcomes rather than static rules. Currently holding 96% of frontier-model quality at 38% of the cost across three pilot workloads.
Verifiable action logs for agents
Tamper-evident, hash-chained logs of every agent action and its authorization context — designed so an auditor can replay any decision. Exploring what regulated clients would need to certify autonomous workflows.
Voice agents under adversarial callers
Stress-testing speech agents against prompt injection, social engineering, and accent/noise edge cases using a synthetic adversarial caller framework built in-house.
Get Labs thinking on your roadmap.
Enterprise partners get research previews, experiment briefings, and first access to graduated patterns. Ask what the lab is learning about your problem.
