October 2026
Intermediate to advanced
225 pages
4h 39m
English
The process of developing robust evaluations for LLM applications is inherently iterative. It involves creating test cases, assessing performance, and refining the system based on those observations. High-level guides, such as Anthropic’s documentation on creating empirical evaluations for Claude Anthropic 2024, often depict the evaluation process as a cycle of developing test cases, engineering prompts, testing, and refining (Figure 3-1).1 This chapter covers the Analyze step of our Analyze-Measure-Improve lifecycle (Figure 1-2). We walk through how to bootstrap an initial dataset, read traces, and build a structured understanding of where ...
Read now
Unlock full access