Thinking in Silico: Goodfire CTO Dan Balsam on Concept Manifolds & a $1000/Month ML Research Agent

Hello, and welcome back to the Cognitive Revolution!

Today I'm speaking with Dan Balsam, CTO of Mechanistic Interpretability startup Goodfire.

The occasion for this conversation is the launch of Silico, a long-horizon, agentic ML research platform, which began as an internal tool, that's meant to democratize access to Goodfire's hard-won expertise, including GPU cluster management, research taste, and all sorts of experimental, visualization, and validation techniques. 

At $1000 per month per seat for enterprise customers, it's not cheap, but compared to Goodfire's high-touch research engagements, which are staffed by forward-deployed research engineers and can easily reach 7 figures, it is 2 orders of magnitude more affordable.  And consistent with the company's public benefit charter and safety-focused mission, they will soon be introducing special pricing and grants for safety & alignment researchers, which I definitely intend to apply for myself.

Of course, we cover a lot more than the platform, beginning with an update on Goodfire research.

We discuss their work on Predictive Data Debugging, which uses interpretability techniques to identify the concepts that network updates are likely to affect, thus making it possible to identify and address anomalies before they become unpleasant surprises.  

We go deep on their series of papers on the intricate and often quite beautiful geometries that LLMs use to represent advanced concepts, including how we should understand this as an evolution of the linear representation hypothesis, how they're using this deeper understanding to improve model steering, and how they've identified spatial representations of such advanced concepts as the Periodic Table and the Tree of Life.  

Along the way, we get Dan's take on key issues, including:

  • the importance of open source models for avoiding dangerous concentration of power,
  • how worried he is about AIs contributing to future pandemics or other biodisasters,
  • the steps Goodfire is taking to prevent misuse of their platform,
  • his reasons for signing on to the recent "pacing the frontier" letter and what he thinks the AI research community would ideally do going forward,
  • why he's reasonably optimistic about monitoring techniques but nevertheless believes that we ultimately have no choice but to intentionally design techniques to control what models learn in the training process,
  • and which training techniques he believes are sufficiently likely to prove problematic that they should be avoided.  

He also shares some of his favorite use cases of Silico so far, dispels internet rumors about hidden research agendas at Goodfire – stating plainly that there are none – and even offers a glimpse into what the researchers at Goodfire are actively discussing around the lunch table these days, which, unsurprisingly given recent research, emphasizes the mysteries around AI welfare and consciousness.  

With that, I hope you enjoy this educational and thought-provoking conversation about the shape of AI thought, and the new highly-autonomous ML research platform, Silico, with Dan Balsam, co-founder and CTO of Goodfire.

Watch now!

Thank you for being part of The Cognitive Revolution,
Nathan Labenz

Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to The Cognitive Revolution.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.