Chapter 11. From Compute to Cost: Designing Efficient Agentic Systems
About 40% of agentic AI projects may be canceled over the next few years. Not because the models aren’t good enough, but because the systems become too expensive to run or never translate into real business value. Estimates from firms like Gartner already point in that direction.
You learned throughout the book that agents operate through iterative reasoning, tool calls, and multi-step coordination. While this is their great strength, it’s also their biggest weakness. Every additional step adds latency and cost. What looks trivial in a single interaction becomes a serious problem once the system runs at scale.
To understand how this compounds, consider your weakest agent relying on retries to get things right. If that agent has a success rate of 60%, but you want to push it to near certainty, you need multiple retries. The following calculation assumes that each retry is an independent event. As you’ve seen in “Instructor: Thinking at the Failure Level, Not the Tooling Level”, frameworks like Instructor or LangGraph can condition each attempt on prior failures and improve the odds. However, that improvement is the result of explicit feedback loops and additional context, and still costs more tokens per step:
This is where many systems break. Not in accuracy, but in economics. Because for a 99% success rate, it requires a 5x increase in your token ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access