Chapter 4. Collaborative Evaluation Practices
Error analysis (Chapter 3) gives us a list of failure modes. The next question is whether different people on the team agree on what qualifies as a failure. This chapter covers how to align human judgment.
In the previous chapter (Chapter 3), we walked through error analysis: reading traces, identifying failure modes, and building a taxonomy of how your system goes wrong. That process depends on human judgment at every step. You decide what counts as a failure. You decide how to categorize errors. You decide whether a trace is acceptable or not.
Error analysis (Chapter 3) depends on human judgment at every step. You decide what qualifies as a failure, and how to categorize errors. But evaluation criteria are often subjective (what one person considers “helpful,” another might find verbose), and individual judgment is inconsistent (the same person might rate the same trace differently on different days). If the labels are unreliable, everything built on top of them (failure mode counts, automated evaluators, ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access