AI Product Lab: The AI PM’s Guide to Understanding Evals with Laurie Voss and Aman Khan
Published by O'Reilly Media, Inc.
What evals are and why they matter for PMs shipping AI features
Demos of AI features are often impressive. But what separates the ones that hold up in production from the ones that quietly degrade is whether the team has a repeatable practice for evaluating what the model is actually doing at scale. This practice, commonly referred to as evals, is one of the most important and misunderstood skills in AI product management. In this session of AI Product Lab, Laurie Voss and Aman Khan will walk product managers and builders through what evals are, why they matter, and how leading AI teams actually run the workflow. You’ll come away with a working mental model of the discipline and a clearer sense of the kinds of failures evals are designed to catch.
AI Product Lab is a live online event where AI product leader Aman Khan demonstrates how AI tools can enhance product management through real-time experimentation and guided discovery. You’ll watch live tool comparisons, from-scratch prototyping (including dead ends), and honest discussion of what works and what doesn’t when applying new AI tools and techniques to product management. Embracing the value of experimentation over polished demos, each event will help you understand how AI can accelerate core product workflows like research synthesis, documentation, and feature prototyping, while helping you build the confidence to practice on your own and evolve into an AI-empowered product manager.
What you’ll learn and how you can apply it
- Understand what evals are, why they’re critical for PMs shipping AI features, and where they fit in the AI product lifecycle
- Get a clear picture of how AI teams identify failure modes in production, including the kinds of problems that a working demo and manual spot-checking will never catch
- Follow an eval workflow from error analysis through eval writing and validation, and come away with a realistic first conversation to have with your own team
This live event is for you because...
- You’re a product manager whose team has shipped or is about to ship an AI feature, and you want to know how to evaluate whether it’s actually working.
- You’re an AI product leader who wants to build a rigorous evaluation practice into your team’s workflow and needs the vocabulary and mental model to lead that conversation.
- You’re a forward-thinking professional in any role who wants to understand what separates AI features that hold up in production from ones that quietly break.
Your Hosts and Guests
Aman Khan
Aman Khan is a lead AI product manager at Google, working on agentic governance, evals, and observability on the agent platform team. Aman is also writing AI Product Management, an O’Reilly book in early release, that covers how to build, evaluate, and ship successful AI products from prototype to production. He also teaches courses on AI product management and writes about what he learns on his Substack (AI Product Playbook). Previously, Aman led products at Arize, Spotify, Cruise, and Apple.
Laurie Voss
Laurie Voss is head of developer relations at Arize AI, the leading company for AI observability and evaluations. He’s been a developer for over 30 years and was cofounder of npm, Inc. He believes passionately in making the web bigger, better, and more accessible for everyone.