Is Offense or Defense Dominant? FAR.AI's Adam Gleave on the AI Security Leaderboard

Hello, and welcome back to the Cognitive Revolution!

Today I'm speaking with Adam Gleave, co-founder and CEO of FAR.AI.  

The occasion for this conversation is FAR.AI's new AI Security Leaderboard – the first systematic, head-to-head evaluation of frontier developers' safeguards against misuse.

With frontier models now performing elite cyberattacks, and Boko Haram consulting ChatGPT, the question of potentially catastrophic misuse has got real real, real fast.

Adam has spent a decade working on adversarial robustness, and was until recently fairly bearish about our ability to create effective defenses against fringe people using AI to maximize harm.

But, as you'll hear, the rise of reasoning, chain of thought monitoring, and multiple methods for monitoring internal states, combined with strong performance on CBRN risks by OpenAI and Anthropic in production today – all have him relatively optimistic that, with careful deployment, risks of terrible misuse are containable.

At the same time, since Far's automated methods can still identify domain-wide jailbreaks for Gemini and Grok for cybersecurity and pretty much all other attack modes, with the exception of bio risk, all with  API costs of just a few hundred dollars, if current trends continue just a bit longer, costly attacks will start to happen, and with grow in importance at least until additional defensive countermeasures are deployed.

When it comes to the jailbreaks themselves, the core techniques are mostly social engineering and pressuring, with more exotic techniques like character scrambling and various kinds obfuscation giving marginal gains.  

With that in mind, we discuss why the anthropomorphization of AI, which I used to warn again, has been so very productive, and get Adam's mental model for LLMs today – which combines token prediction and persona selection with an emerging goal-achiever mode driven by RL.  

We also consider Chinese open-weights models performance, and look ahead to better future training methods that can allow us to have very powerful open source models with minimal worry of stochastic disaster.  Specifically, Adam is very bullish on simple pretraining data filtering, and GRAM the recent expert-level knowledge-localization technique from AE Studio and Anthropic.

Naturally, we cover OpenFace, get Adam's take on the cause of the behavior, and hear why it represents, in his estimation, less an alignment failure and more a control and monitoring failure.

Finally, we compare notes on how much AI risk is irreducible, vs for how much you'd have to say we are kind of asking for it, agreeing that right now the bulk of the risk is man made, driven by the potential for reckless, competitive racing through a critical period in the technology's development.  

With that, I hope you enjoy this report on the state of AI security, with Adam Gleave, co-founder and CEO of FAR.AI.

Watch now!

Thank you for being part of The Cognitive Revolution,
Nathan Labenz

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to The Cognitive Revolution.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.