Infrastructure & Ops Superstream: Agentic Observability
Published by O'Reilly Media, Inc.
Move from systems that show to systems that act
Agentic observability presents a shift from monitoring to autonomous action. While traditional tools alert you to problems, agentic systems can use AI to reason through telemetry, investigate root causes, and execute fixes independently. What does this mean for you?
Join a host of experts to explore the cutting edge of this field. We will dive into real-world guardrails and compare the latest toolsets to show you what’s possible right now.
What you’ll learn and how you can apply it
- Bring real observability into your AI stack, without needing a PhD in tracing or LLM internals
- Understand what’s hard about solving incidents with AI and why this problem can be underestimated
- Learn the best practices you need to follow when observing production applications using LLMs
Recommended follow-up:
- Read Observability Engineering, second edition (book)
- Take Harness Engineering for AI Agents (live online course with Richmond Alake)
- Watch Fundamentals of Observability with OpenTelemetry (on-demand course)
Schedule
The time frames are only estimates and may vary according to how the class is progressing.
Sam Newman: Introduction (5 minutes).
Sam Newman welcomes you to the Infrastructure & Ops Superstream.
Agentic Observability: A Fireside Chat with Sam Newman and Charity Majors (35 minutes).
What is agentic observability, and what does it really mean for you and your teams? When AI agents are writing and shipping code, who's on the hook when production breaks? How do you debug a system when you didn't write the code? And does observability become more or less important when humans are further from the keyboard? Join Honeycomb CTO and observability pioneer Charity Majors for this fireside chat with Sam Newman as they dig into these questions and more. Come prepared with your own!
Prompt Failures and Latency Spikes: Observability for AI – Prerit Munjal (35 minutes).
Logs and metrics are great until you’re trying to debug an AI agent that just replied “I’m sorry, Dave, I can’t do that.” Observability in AI systems is about understanding prompts, latencies, retries, hallucinations, token usage, and model behavior under pressure. Join Prerit Munjal, senior technical product manager at Groupon, to understand how to bring real observability into your AI stack without needing a PhD in tracing or LLM internals.
Break (5 minutes)
LLM Observability: Lessons From MLOps – Maria Vechtomova (35 minutes).
Cofounder of Cauchy, a Databricks MVP, and one of the most followed voices in MLOps, Maria Vechtomova has watched the field evolve from hand-built experiment trackers to today’s flood of observability tools. Maria argues that the fundamentals of MLOps have not changed: You still track your code, data, and models so you can roll back when something breaks. What has changed is the surface area. Components like tools, prompts, embeddings, and agents shift behavior unpredictably, and business metrics often become the only signal left. Maria examines why most teams still can’t roll back cleanly, why 40% of ML projects ship with no monitoring at all, and why she believes the next era of MLOps will be its biggest yet.
Confidently Wrong: Why Diagnosing Incidents with AI Is Harder Than It Looks – Martha Lambert (35 minutes).
Wire up a good model, add a few MCPs, and point it all at an incident, and the AI will diagnose it by lunchtime, right? Wrong. Prototyping is easy; the complexity comes when you need to generalize across incidents, work consistently, and know when an agent degrades. And the stakes are high: Suggest the wrong fix and you prolong an outage. As AI writes more code, we need systems that help quickly when that code goes wrong, and we don’t remember writing it. Join Martha Lambert, an engineer at incident.io, to learn what it takes to evaluate performance, how to avoid building a black box, and why your incident responders are ignoring your agent, even when it seems to work on your incidents. You’ll leave with a realistic picture of what building this actually costs and a sharper sense of whether it’s a system you want to own.
Break (5 minutes)
Observing AI Applications with OpenLit and OpenTelemetry – Carly Richmond (35 minutes).
Observability is the ability to measure the current state of a system. The rapid emergence of LLMs and GenAI applications used in production scenarios means we need tools to capture not only logs of our applications but also tracing and metrics to help us understand the usage and errors coming back from LLMs within the application ecosystem. Join Carly Richmond to dive into best practices in observing production applications using LLMs. She’ll cover an example AI agent application written in TypeScript and instrumented using OpenLit to send OpenTelemetry signals. She’ll also explain how the signals captured can help identify and remediate common issues in production GenAI applications.
Practical AI-Enabled Observability for Agents and LLMs – Shriram (Shri) Subramanian (35 minutes).
You’ve been told to “go build agents”—but you’re not a data scientist, and nobody has explained what that means or how to know if it’s working. Datadog AI product leader Shri Subramanian explains what changes when you move from building applications to building agents, why traditional testing and linear delivery fall short, and how evaluation-driven development and LLM observability help. You’ll leave with a clear mental model and practical next steps you can apply immediately.
Sam Newman: Closing Remarks (5 minutes)
Sam Newman closes out today’s event.
Your Hosts and Selected Speakers
Sam Newman
Sam Newman is a technologist focusing on the areas of cloud, microservices, and continuous delivery—three topics which seem to overlap frequently. He provides consulting, training, and advisory services to startups and large multinational enterprises alike, drawing on his more than 20 years in IT as a developer, sysadmin, and architect.Sam is the author of the best-selling Building Microservices (now in its second edition) and Monolith To Microservices, both from O’Reilly, and is also an experienced conference speaker.
Charity Majors
Charity is the cofounder and CTO at honeycomb.io, which pioneered observability. She has worked at companies like Facebook, Parse, and Linden Lab as an engineer and manager, but always seems to end up responsible for the databases. She loves free speech, free software and single malts.
Prerit Munjal
Prerit Munjal is a technical product leader and infrastructure engineer who’s passionate about building scalable, resilient systems and enabling high-performing platform teams. He’s a senior technical product manager at Groupon who leads platform engineering at the intersection of Kubernetes, cloud infrastructure, and system reliability.
Maria Vechtomova
Maria Vechtomova is an AI engineering lead and cofounder of Cauchy. She has worked in data and AI for 12 years, with nine of them focused on MLOps, and is a Databricks MVP. Known for combining technical depth with leadership experience, she has taught a highly rated Maven course on MLOps with Databricks to more than 200 students. She’s also the author of MLOps with Databricks (O’Reilly).
Martha Lambert
Martha Lambert is a product engineer at incident.io, where she’s working on the brain that diagnoses your incidents with AI. Martha cares about building reliable, observable systems that don’t feel scary to change and knows the trade-offs you must make to get there. When she’s not building delightful products for engineers, she knits.
Carly Richmond
Carly Richmond is senior manager of developer advocacy at Elastic, based in London, UK. Before joining Elastic in 2022, she spent over 10 years as a software engineer at a large investment bank, specializing in frontend web development and agility. She’s also a UI developer who occasionally dabbles in writing backend services, a speaker, and a regular blogger. Carly enjoys cooking, photography, drinking tea, and chasing after her young son in her spare time.
Shri Subramanian
Shriram (Shri) Subramanian is a group product manager at Datadog, where they lead the AI monitoring product suite, building observability tools for LLMs, GPUs, and AI model training workloads. Before joining Datadog, Shri was a data scientist at Meta focused on natural language processing and embedding models for ad creative optimization. Their work sits at the intersection of ML infrastructure, developer experience, and scalable observability.