Malt field report: from experimentation to production, adapting your observability infrastructure for generative AI
Speaker
Description
Internal use and large-scale production deployment of generative AI solutions bring their share of challenges: exploding costs, confidentiality concerns, service level agreements and non-determinism at execution. In that situation, how do you stay in control?
A first answer we brought at Malt was to update our observability stack to meet these new needs.
In this technical field report we explore the backstage of our architecture:
- OpenTelemetry standardisation: how we use OTEL as the standard for LLM traces and how we adapted it to our different technical stacks (Python and the JVM and Spring ecosystem).
- Architecture and data pipeline: the configuration of our collection pipeline (OTEL Collector, selective routing to Datadog) and the tricks to debug it effectively.
- Steering ROI: how we monitor AI adoption and costs internally, in particular the coding agents used by our developers (Claude Code).
- Democratisation with Langfuse: why and how we gave trace access to a much wider audience through this LLM engineering platform.
In short: the good practices, the standards of tomorrow and the mistakes to avoid in order to operate your LLMs with peace of mind.