Chapter 4. Collaborative Evaluation Practices
Error analysis (Chapter 3) gives us a list of failure modes. The next question is whether different people on the team agree on what qualifies as a failure. This chapter covers how to align human judgment.
In this chapter, you will learn:
-
When a single “benevolent dictator” expert is enough for evaluation, and when you need multiple annotators
-
A nine-step collaborative annotation workflow for producing consistent rubrics
-
How to measure inter-annotator agreement with Cohen’s Kappa (and when to use alternatives like Fleiss’ Kappa or Krippendorff’s Alpha)
-
How to run alignment sessions that resolve disagreements and improve rubrics
-
Common pitfalls in collaborative evaluation and how to avoid them
Error analysis (Chapter 3) depends on human judgment at every step. You decide what qualifies as a failure, and how to categorize errors. But evaluation criteria are often subjective (what one person considers “helpful,” another might find verbose), and individual judgment is inconsistent (the same person might ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access