Edition 2026 Talk Data Platform Evolution

Towards self-serve analytics: a year of data transformation at leboncoin

Language FR

Speaker

Patrice Chaperon

Patrice Chaperon

Product Director Data

Description

We migrated to Databricks. Here is what we really learned.

Leboncoin: sixty data people, twenty-seven million users, more than three hundred Tableau dashboards. In 2024 we took a decision: migrate from Redshift and Athena to Databricks, deploy Unity Catalog, document our domains in Coalesce, and make business teams autonomous. The kind of programme that makes beautiful slides.

A year later we are midstream. The stack is moving forward. What we really learned is that the technical problem was the easy one.

This talk covers what we concretely did, incremental migration, semantic layer, data products, an end-to-end real estate pilot, and above all what it revealed: tables with no owner for five years, four different definitions of the same KPI, data analysts who want autonomy but not accountability. And why solving all of that is also the only foundation on which AI agents can be placed without going off the rails.

Agenda. The diagnosis: sixty data people, a bottleneck, business analysts outside the data teams; the legacy stack of Redshift, Athena and Glue as the visible symptom; the real problem, four definitions of a lead, pipelines with no producer, non-existent governance; the decision not to patch but to transform, and to create a dedicated taskforce so the programme is not drowned in business as usual.

What we built: the stack migration to Databricks and Unity Catalog, why that choice and what it changes for governance and access; Coalesce to document business domains and align definitions, not sexy but indispensable; incremental migration rather than big bang, what it costs in duration and what it avoids in risk. The semantic layer in practice, with data stewards per domain: who, how, and why it is harder than choosing a tool. The real estate pilot as the first end-to-end domain, with tables migrated, Coalesce documented, analysts autonomous on Databricks, and what we delivered, did not deliver and got wrong.

What is genuinely hard: the ownership debt, where migrating a table takes two days and finding who owns it takes two weeks, with inherited tables, orphan pipelines and producers who left the company, and business teams who want autonomy but not accountability for the data product. The skills gap, where analysts trained on Tableau do not become autonomous on Databricks in two sprints, what we put in place with data literacy, onboarding per vertical and explicit roles, and what we still cannot measure. Governance against pace, building solid foundations while everyone wants to ship fast, and how the real estate pilot forced the trade-offs we had been postponing for eighteen months.

Why this also matters for agents: what we target next quarter, business users generating insights through AI agents; what we learned, that an agent relies on the same foundations as an analyst, clear ownership, stable definitions, governed access, except that it tolerates no ambiguity; the direct link, Unity Catalog as the access layer, unified KPIs, data products with an owner, the same work rather than an extra AI project; and what is still missing, honesty about where we stand.

Takeaways: the technical part is the easiest; a well chosen pilot forces the trade-offs you were postponing, so picking the right domain matters; making data usable by agents means finishing the work we should have done for humans.