September 2026
Intermediate
216 pages
5h 33m
English
Everything we have built so far runs offline: load a dataset, run the eval, check the results. This is essential for development and CI, but it misses a critical class of problems — the ones that only appear with real user inputs, real-world distribution shifts, and the full production environment.
Online evals run on live production traffic. They sample requests, score them in real time or near-real time, and alert you when quality drops. They are the last line of defense between your system and your users.
But online evals introduce new constraints. You cannot afford to run an LLM judge on every request — the latency and cost would be prohibitive. You cannot block user requests ...
Read now
Unlock full access