Cracking LLMs Open: Emmanuel Ameisen Live with Tim O’Reilly
Published by O'Reilly Media, Inc.
What Claude's internals reveal about how it actually thinks
Anthropic's interpretability team studies Claude's internals the way neuroscientists study a brain, but the science is younger, messier, and far less mapped. Emmanuel Ameisen has been part of that effort for the last two years, and at O'Reilly's recent Foo Camp, he shared some of what he’s learned about what happens inside of Claude when it processes information.
Token prediction is often framed as simply “pattern matching,” Emmanuel noted in his talk, but it turns out that in order to be superhuman at token prediction, models build a complex understanding of the world. Ask Claude to finish a sentence about a short bike ride across the “GG bridge” and it correctly infers that the bridge must be the Golden Gate, placing the user in San Francisco. From there, the model can build on this understanding to answer questions, for instance how long it will take to get to “world-class skiing” at Lake Tahoe. His team can see this world model emerge inside the model as different kinds of activation patterns light up when the model completes complex tasks.
Tim asked Emmanuel to join him for this episode to extend his Foo Camp presentation. They’ll start by establishing some foundations about what we mean when we talk about world models, and whether what we find inside language models matches that description. They’ll discuss examples of these world models at work when doing math, or writing poetry, where the models choose how a line will end before writing it. They’ll also discuss the tools Anthropic's interpretability team has built to develop this understanding, steering models by "pushing or pulling" on their activations to get them to behave accordingly. Along the way, Emmanuel and Tim will think big-picture about interpretability: What should people take on faith about model behavior right now, and what should they not?
Join in to find out what Emmanuel’s learned from two years of looking inside Claude. And bring your questions and comments. They’ll help guide the conversation.
Recommended prep or follow-up:
- Listen to Generative AI in the Real World: Emmanuel Ameisen on LLM Interpretability (podcast)
Schedule
The time frames are only estimates and may vary according to how the class is progressing.
Wednesday, September 9, 2026, at 9:00am PT / 12:00pm ET
- Interactive discussion and Q&A (60 minutes)
Your Hosts and Guests
Tim O'Reilly
Tim O’Reilly is the founder and CEO of O'Reilly Media, Inc. His original business plan was simply "interesting work for interesting people," and that's worked out pretty well. He publishes books, runs online conferences, invests in early-stage startups, urges companies to create more value than they capture, and tries to change the world by spreading and amplifying the knowledge of innovators. He’s perhaps best known for his role in shaping big ideas like open source software, unconferences (Foo Camp), Web 2.0, and government as a platform. His 2017 book WTF? What’s the Future and Why It’s Up to Us explored the role of human agency in shaping the future in the face of the coming AI wave. These days, he’s focused on mechanism design for the human-AI economy. Mechanism design is sometimes described as “reverse game theory”, whereby you start with the outcome you want and then figure out what rules of the game will produce that outcome. He explores these ideas at the non-profit AI Disclosures Project, which he co-founded with Ilan Strauss. He writes frequently on Substack at the O’Reilly Radar, the AI Disclosures Project’s Asimov’s Addendum, and his own Conversations with AI.
Emmanuel Ameisen
Emmanuel Ameisen works on mechanistic interpretability at Anthropic: building tools that trace the computations inside large language models, and using them to study how those systems do what they do. His work includes finding evidence that LLMs plan ahead, an account of the general circuits they learn for representing and manipulating numbers, and explanations for some of the surprising behavior we see in frontier systems. Before Anthropic, he was a Staff ML Engineer at Stripe, and led Insight Data Science's AI program, directing more than a hundred applied ML projects. He is the author of the O'Reilly book Building Machine Learning Powered Applications.