Edition 2026 Talk Evals & Observability

Why Ship an AI You Can't Measure? Building an LLM Judge You Can Actually Trust, at the Scale of Millions of Products

Language EN · FR

Speakers

Adrien Morvan

Adrien Morvan

AI Engineering Manager

Oussama Taki Amrani

Oussama Taki Amrani

AI Engineer

Description

An AI in production without a performance measure is a bet, not a product. At Mirakl, our Catalog Transformer converts raw seller catalogs into marketplace-ready product listings every day: over 60 million products processed. But how do you know whether the result is actually good, when there is often no single right answer, only defensible ones?

This talk tells the story of how we built a panel of LLM judges that continuously evaluates our system on real production traffic, at a controlled and predictable cost. And above all, why the judge was never the hard part: in our first annotation campaign, we, the experts, were unanimous on only 51% of cases.

On the agenda: adaptive GTIN-hash sampling that makes the score reproducible and the budget bounded; a labeling system that separates genuine errors from defensible ambiguities; building a golden set and calibrating AI against humans; and a prompt iteration loop that guards against regressions. Reusable recipes for evaluating any AI system facing the real world, where "correct" is never binary.