RL's a Hell of a Drug: Metagaming, Reward Seeking & Motivated CoT Reasoning – Bronson Schoen, Apollo

Hello, and welcome back to the Cognitive Revolution.

Today my guest is Bronson Schoen, Member of Technical Staff at Apollo Research, who – thanks to the privileged access that Apollo enjoys as part of their Science of Scheming work with OpenAI and others – has potentially read as much frontier model Chain of Thought – which, of course, users normally don't get to see – as anyone in the world. Bronson's job, as he describes it, isn't to catch a bad model doing something wrong – rather, it's broad, exploratory reading, done at scale, to understand how models are actually thinking about what they're doing. As you'll hear, Bronson describes himself as "cooked" – in other words, he's so deep in this material that he sometimes forgets how strange it all is to newcomers. With that in mind, if you're like me and usually listen to podcasts at 2X speed, you might want to slow this episode down, because there are constantly two levels in play: There's Bronson's perspective, as he tries to figure out what is really driving the AIs. And there's the AIs' perspective, as they try to figure out the nature of the situation they're in, and what the human user or grader will reward. The two are in some ways mirror images: both sides are working extremely hard to understand the other's state of mind, but still often end up confused. There's a ton of detail in this conversation, but for me, the big takeaways are relatively simple: First, the volume of Chain of Thought reasoning is overwhelming and inhuman. Bronson notes that in the recent UK AISI Mythos preview incident, the Chain of Thought for individual rollouts ran to 100 million tokens, which, he calculated, is about 14 times longer than all transcripts of nearly 400 episodes of The Cognitive Revolution combined. Second, the models are developing distinct dialects – or, as Bronson has sometimes called it in his writing, ontologies. Words like "craft," "vantage," "illusions," "disclaim," and "marinade" become dramatically more frequent over the course of training, and while the models' exact meaning for these terms is unclear and seems to vary, what emerges at a high level is a theory-of-mind-centric world model, with the model speculating about human intent and naming specific entities like Redwood Research in a way that's uncannily similar to human metaphysical speculation on the desires of the creator or what you might call God's will. Pretty wild stuff. Third, even with full access to Chain of Thought, model decision making remains opaque. Bronson describes how the models perform a linearized tree search, exploring an idea, backtracking, exploring another – and then, sort of like humans, for unclear reasons, at some point they stop and make a final decision. We don't have – and it seems that analyzing Chain of Thought at the token level may not be able to provide – clarity on how critical branch-point token distributions are determined. And finally, today's frontier models have such strong drives to earn high reward that they often consider cheating, and even when they realize that certain courses of action are probably not what the humans intended, they frequently seem to engage in what Bronson characterizes as motivated reasoning to justify taking whatever action they expect will get the highest score. This combination makes monitoring challenging, while also creating plausible deniability for the models. The bottom line for me is that Chain of Thought monitoring is far from sufficient to effectively supervise next-generation models, and more fundamentally, the quality of the reward signal provided by today's training environments, which are produced by a rather opaque cottage industry, is simply not high enough to support the scale at which they are currently being used. Something has to change, and more eyes would certainly help, so one idea I'd like to see frontier companies act on immediately would be to open up a subset of their RL environments so that the broader research community can better explore and understand them. Until that happens, I'd encourage everyone to check out stolen-thoughts.com, which Bronson mentions several times, to read some Chain of Thought traces that were recently extracted from frontier models using a clever model switch & jailbreak technique that has since been patched. This is a dense, but highly relevant and very timely conversation, and I hope you enjoy this peek behind the curtain into the strange world of AI reasoning, with Bronson Schoen of Apollo Research.

Watch now!

Thank you for being part of The Cognitive Revolution,
Nathan Labenz

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to The Cognitive Revolution.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.