Chapter 4. Context Engineering
The user types under 5% of what the model reads. This chapter is about the other 95%.
For a couple of years, the industry called this work “prompt engineering,” and the phrase implied a magic incantation waiting to be found. Job listings chased it. Conference talks chased it. In hindsight, the framing was a photograph of where we were.
Here’s what changed. Working with a coding agent, you aren’t shaping the initial input to a prompt. You’re operating on the working memory of a system across many turns. The user prompt, the system prompt, the AGENTS.md, the accumulated tool outputs from the last 40 steps: different memory slices, competing for a limited sliver of attention inside the model. In practice the user prompt is often less than 5% of the tokens in the window, so 95% of what you’re trying to influence is the output of your harness rather than the sentence a person typed (see Figure 4-1).
Figure 4-1. A fixed-width context ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access