Chapter 3. The Anatomy of an Agent: Model, Harness, Loop
When I was 16, I built a web browser called XWebs for a national science competition, and I gave it an assistant named Phoebe. She sat next to the browser, slightly animated, and she could talk: press a button, say a thing, get a response. She read pages out loud. I’d wired in the ideas I most wanted from software at the time: speech recognition, text to speech, theming lifted from my Winamp obsession. What I actually wanted was a browser you could pair with to get stuff done.
Phoebe could talk. She couldn’t act. Two decades on, the models hold the conversation fine, and the conversation turns out to be the smallest part of the problem. Two agents built on the same model can behave nothing alike: two of them drew from the same weights and ran the same benchmark, Terminal-Bench, and one landed near thirtieth while the other landed in the top five. Everything I wished I could have given Phoebe, the ability to touch things safely, remember things durably, and know when she was done, is the part that decides ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access