Skip to Content
AI Evals in Practice
book

AI Evals in Practice

by Caio Incau
September 2026
Intermediate
216 pages
5h 33m
English
Packt Publishing

Overview

Build evaluation systems that reveal whether your LLM applications are improving or regressing, using calibrated judges, RAG and agent metrics, CI regression tests, online evals, and cost-aware pipelines

Key Features

  • Build a complete Python eval harness from datasets and scorers to CI and dashboards
  • Evaluate RAG, agents, and prompts with calibrated metrics and human feedback
  • Move from offline testing to production evals, guardrails, and reliability workflows

Book Description

LLM applications can look healthy in dashboards while their answers quietly become less accurate, less useful, or less reliable. AI Evals in Practice gives developers and AI engineers a systematic way to measure quality, catch regressions, and make evidence-based improvements before and after deployment.

You will build evalkit, a complete Python evaluation harness, while learning the core components of an eval: datasets, scorers, runners, and golden test sets. You will create deterministic scorers and LLM-as-judge evaluations, then calibrate judges and mitigate common biases. The book applies these foundations to prompt regression testing and CI, RAG retrieval and generation metrics, agent trajectories and tool calls, and human annotation workflows. You will then extend evaluation into production with online sampling, guardrails, and cost- and latency-aware pipelines, while comparing tools such as DeepEval, promptfoo, Langfuse, and Braintrust. A case study and final project bring the pieces together into an eval-driven development workflow with CI and a dashboard.

By the end, you will be able to design evaluation pipelines that help you ship LLM systems with measurable, repeatable quality.

What you will learn

  • Design eval datasets, scorers, runners, and golden test sets
  • Build deterministic metrics for repeatable quality checks
  • Calibrate LLM judges and reduce common evaluation bias
  • Add prompt regression tests to continuous integration
  • Measure retrieval and generation quality in RAG systems
  • Evaluate agent trajectories, tool calls, and multi-turn behavior
  • Run human annotation workflows and production online evals
  • Control evaluation cost and latency without losing signal

Who this book is for

This book is for AI engineers, LLM application developers, machine learning engineers, platform engineers, and technical leads who build or operate production systems using large language models. It is especially useful for teams working with prompts, RAG pipelines, agents, or AI features that need measurable quality and regression protection. Readers should be comfortable with Python and familiar with building or integrating LLM applications.

Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.

Read now

Unlock full access

More than 5,000 organizations count on O’Reilly

AirBnbBlueOriginElectronic ArtsHomeDepotNasdaqRakutenTata Consultancy Services

QuotationMarkO’Reilly covers everything we've got, with content to help us build a world-class technology community, upgrade the capabilities and competencies of our teams, and improve overall team performance as well as their engagement.
Julian F.
Head of Cybersecurity
QuotationMarkI wanted to learn C and C++, but it didn't click for me until I picked up an O'Reilly book. When I went on the O’Reilly platform, I was astonished to find all the books there, plus live events and sandboxes so you could play around with the technology.
Addison B.
Field Engineer
QuotationMarkI’ve been on the O’Reilly platform for more than eight years. I use a couple of learning platforms, but I'm on O'Reilly more than anybody else. When you're there, you start learning. I'm never disappointed.
Amir M.
Data Platform Tech Lead
QuotationMarkI'm always learning. So when I got on to O'Reilly, I was like a kid in a candy store. There are playlists. There are answers. There's on-demand training. It's worth its weight in gold, in terms of what it allows me to do.
Mark W.
Embedded Software Engineer

You might also like

AI Agents: The Definitive Guide

AI Agents: The Definitive Guide

Nicole Koenigstein
Evals for AI Engineers

Evals for AI Engineers

Shreya Shankar, Hamel Husain

Publisher Resources

ISBN: 9781808817250