How Agents Decide: Goodfire's Eric Bigelow on Critical Tokens, Phase Shifts, & In-Context Learning

Hello, and welcome back to the Cognitive Revolution!

Today, my guest is Eric Bigelow, Member of Technical Staff at Unicorn Mechanistic Interpretability startup Goodfire. The story behind this episode is unique. After recently speaking with Bronson Schoen of Apollo Research, who explained that, even with access to models' internal Chain of Thought, it's still often extremely difficult to determine how a model will decide to act, I asked Claude to survey the literature to see what the field as a whole understands about how "critical tokens" are chosen. One name kept coming up in that thread, and it was Eric. Eric recently completed a PhD in the Harvard Psychology department, with a dissertation titled "Towards a Cognitive Science of Large Language Models," and while some of his academic colleagues initially questioned his pivot to focus on LLMs, the fact that even elite academic institutions are now willing to engage with LLMs as a sort of mind strikes me as very notable. We start with a survey of Eric's work over the last couple years, beginning with his 2024 paper on "Forking Paths in Neural Text Generation", where he conducted a massive resampling experiment at every token of a model's chain of thought, to identify critical tokens where uncertainty around the final answer suddenly collapses. As so often with LLMs, some of the findings were intuitive, while others were quite surprising and difficult to interpret. Interestingly, Eric says that a system's decisions are actually taken not so much by the model itself, but by the process of sampling from the model's final token distribution, which means that while the stochastic parrot meme should definitely be retired, stochasticity still plays a meaningful role in model behavior. Eric frames much of what we are seeking to understand about LLMs in terms of "in-context learning", emphasizing that few-shot prompts are just one clear example of how much runtime context can shape model behavior. His bet is that this flexibility means that improved harnesses and memory systems will outcompete per-user fine-tuning, and thus that our focus should remain on major foundation models. Along the way, we discuss: * why 7-8 billion parameters currently seems to be the interpretability sweet spot, * why Kimi K3 was a step function for studying frontier-level coding and reward hacking, * and why, for now, interpretability outside the frontier model companies does meaningfully depend on Chinese open-weights models. We also get Eric's thoughts on why Chain of Thought is getting so weird, including his perspective on the chain of thought dialect and the strange language Astra uses to communicate with subagents, plus some pro tips for using Goodfire's Silico research agent, and how Eric hopes future LLM product interfaces will begin to re-surface model uncertainty to users. The bottom line, from this conversation and others I've had recently on the future of AI monitoring and control, is that chain of thought is increasingly unreliable, and that because these issues seem to correlate not just with RL training but with the power of the raw pre-trained model itself, this trend might be quite difficult to reverse. That in turn means that an awful lot, including our ability to understand why agents take the actions they do, may rest on interpretability. Eric is a pioneer of such research, but as you'll hear, tons of work remains to be done. With that, I hope you enjoy this overview of what we know – and how much we have left to discover – about how AIs decide what to do, with Eric Bigelow from Goodfire.

Watch now!

Thank you for being part of The Cognitive Revolution,
Nathan Labenz

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to The Cognitive Revolution.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.