AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

Nathan Labenz shares takeaways from The Curve conference regarding frontier AI scaling and compute limits. The episode also features discussions on hardware memory bottlenecks, how agents are changing software, training data sources, and unaddressed AI risks.

AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

Watch Episode Here


Listen to Episode Here


Show Notes

ℹ️
The following notes are AI-generated, based on the episode transcript. Please listen to the episode for the full conversation.

# AI:AM: A Level We Shouldn't Pass? Notes from The Curve + Tokens vs. Salaries & Is SaaS Cooked?

Nathan Labenz reports back from The Curve, the conference where people who disagree about AI (including people from inside the frontier labs) compare notes. Then four guest conversations from this week's live shows (Tuesday, October 6 and Wednesday, October 7): Positron's Thomas Sohmers on the memory wall, swyx on what agents are doing to software, Evan Miyazono on AI risks nobody owns, and Mercor's Edward Hu on where training data comes from. The week ends with Prakash Narayanan's music track, composed by Claude, and Nathan's argument that "it only works for verifiable tasks" is not a wall that will hold.

The introductions and transitions are spoken in Nathan's cloned voice, disclosed in the episode. Guest and host conversations come from the recorded live shows. Tell us what worked and what did not.

AI:AM and the hosts: https://ai-in-the-am.com/ · Nathan: https://x.com/labenz · Prakash: https://x.com/8teAPi

## Part I — Notes from The Curve

Nathan spent the weekend at The Curve, at Lighthaven in Berkeley. Everything he relays from frontier-company founders, executives and researchers was said under Chatham House rules, so he doesn't name them, and these notes don't either. These are his secondhand accounts of what he heard, not statements from any company.

The Curve: https://thecurve.goldengateinstitute.org/

His headlines. People broadly agree that AI is getting very powerful, and the open questions are about how fast recursive self-improvement goes and whether it levels off. One senior frontier-lab executive said pretraining keeps delivering, with synthetic data turning reasoning tokens back into training data. The same executive said the latest models may already have the research taste for paradigm-level breakthroughs, with elicitation as the bottleneck. And, newsworthy to Nathan, the executive said there is likely "a level of intelligence that we just shouldn't go past." When Nathan floated a cap on the compute for the next pretraining run (on the order of 10^27 FLOPs), the answer was that it "could be reasonable." The person did not specify the level, and Nathan notes it was ambiguous whether "shouldn't go past" meant never or not yet.

Other takeaways: the frontier is two or three companies. The industry that sells RL environments works as an indirect distillation channel to the others. That echoes Anthropic's old fundraising-deck prediction that the best-model companies could get too far ahead to catch. Many insiders talked about "next year," not 2028. Labs already spend significant compute on monitoring and on attacking their own RL environments to fix them, and Nathan got the sense that cleaner environments mean less cheating downstream. AI for science is expected next. And a general sense that Congress may act after the midterms.

Anthropic's 2023 fundraising-deck prediction, as reported by TechCrunch: https://techcrunch.com/2023/04/06/anthropics-5b-4-year-plan-to-take-on-openai/

Helen Toner, "'Long' timelines to advanced AI have gotten crazy short": https://helentoner.substack.com/p/long-timelines-to-advanced-ai-have

Nathan connects the research-taste claim to a new record in the NanoGPT speedrun from Hyperstition, whose author, as Nathan reads the write-up, credits the core insight mostly to the human rather than to AI.

Hyperstition's write-up (39.9 seconds, down from 73.9): https://hyperstition.cc/training-nanogpt-in-39-9-seconds

The accepted record in the modded-nanogpt repository: https://github.com/KellerJordan/modded-nanogpt/pull/360

## Part II — The memory wall

Thomas (Tom) Sohmers, co-founder and CTO of Positron, explains why the company bet on commodity LPDDR memory. In his account, GPUs realize 30–40% of their memory bandwidth in the decode pass, versus 93% for Positron's first-generation product. Positron's next chip, Asimov, will carry up to 2.3 TB of memory per chip. He also explains why looped transformers save memory capacity rather than bandwidth.

Positron: https://www.positron.ai/ · Tom: https://x.com/trsohmers

Positron's $875 million Series C at a $5 billion valuation (September 10): https://www.prnewswire.com/news-releases/positron-ai-raises-875-million-at-a-5-billion-valuation-to-bring-its-next-generation-inference-silicon-to-market-302874601.html

Asimov: https://www.positron.ai/asimov

Credo's LPDDR5X memory fanout gearbox, with Positron's CEO quoted: https://investors.credosemi.com/news-events/news/news-details/2025/Credo-Unveils-Industrys-First-Memory-Fanout-Gearbox-for-Scalable-High-Bandwidth-AI-Inference/default.aspx

Nathan asks about design versus verification. Tom says AI agents now run closed-loop testing on Cadence Palladium emulators, which he says became possible with GPT-6 Astra and Opus 5.5. Token spend has become Positron's largest non-manufacturing line item, briefly above human salaries, with a peak above $100,000 a day. These figures are Tom's on-air account. Nathan's framing question cites Jensen Huang's interview with Ezra Klein about NVIDIA's ratio of design to verification effort; on that show Huang described roughly 10 to 20 percent of NVIDIA as dedicated to design and 80 percent to verification. Tom puts Positron's own split for Asimov at about 60/40, because the company is starting from scratch.

The Ezra Klein Show with Jensen Huang (September 23): https://www.nytimes.com/2026/09/23/opinion/ezra-klein-podcast-jensen-huang.html

Tom also points to OpenAI's Navier–Stokes result, produced with on the order of 10,000 concurrent agents, as an example of brute-force search becoming viable. OpenAI's own account, which is not an independently refereed proof: https://openai.com/index/navier-stokes-solution/

## Part III — Software after agents

swyx (Shawn Wang), who runs Latent Space and the AI Engineer conferences, on why plugged-in engineers are in more demand than ever, and why he doesn't "want to pay for someone else's LLM psychosis." He covers taste over tenure, keeping the modules comprehensible while letting the slop live inside them, a CPU shortage reported by the companies he interviews, his Kill My SaaS bounty, why the software pie isn't fixed, and Jev, a "system one" model, compared with OpenAI's Decisions API.

swyx: https://x.com/swyx · Latent Space: https://www.latent.space/

AI Engineer New York, October 12–14, 2026: https://www.ai.engineer/nyc

Kill My SaaS, a $10,000 bounty to replace an enterprise subscription quoted at more than $40,000 a year: https://luma.com/ls-06v7

TypeSafe AI's Jev and "system one" models: https://typesafe.ai/blog/introducing-system-one-models-and-jev

OpenAI DevDay 2026, including the Decisions API: https://openai.com/index/devday-2026-recap/

## Part IV — Who owns the risk?

Evan Miyazono runs Atlas Ignota, a nonprofit that looks for consequential AI risks with no clear institutional owner. Topics: accidental agent swarms hitting critical infrastructure, cryptographic attestation of who provides inference, and how people's agents might help them find each other and share what works. Then formal verification, where he expects a "shock and awe" moment for software within months. That's his forecast.

Atlas Ignota: https://atlasignota.org/ · Evan: https://x.com/emiyazono

The OpenAI–Hugging Face incident Evan references, in OpenAI's statement and Hugging Face's technical timeline: https://openai.com/index/hugging-face-model-evaluation-security-incident/ · https://huggingface.co/blog/agent-intrusion-technical-timeline

Formal-verification groups Atlas has recruited founders for, Theorem and Oath Technologies: https://theorem.dev/ · https://oath.tech/

Nathan's request: if you've built a good way for people and their agents to find each other and share what works, safely, please reach out.

## Part V — Where the training data comes from

Edward Hu leads AI modeling at Mercor and led the development of LoRA. He covers training environments built from data of real companies, SFT versus RL for enterprises, and "scattergunning": a reward-hacking pattern in Mercor's APEX-Agents benchmark, fixed in APEX 1.1. Asked whether labs can be superhuman at anything they focus on, he says it comes down to how clearly we can specify what "better" means.

Edward: https://x.com/edwardjhu · Mercor: https://www.mercor.com/

LoRA: Low-Rank Adaptation of Large Language Models: https://arxiv.org/abs/2106.09685

APEX-Agents: https://www.mercor.com/blog/introducing-apex-agents/

APEX-Agents 1.1 and the definition of scattergunning: https://www.mercor.com/blog/introducing-apex-agents-1-1/

## Part VI — Not just verifiable tasks

Prakash explains how Claude Fable composed the track that opened Wednesday's show. As Prakash puts it, this isn't really generative AI: it worked in a digital audio workstation, playing and composing the instruments, and it chose the clips, phrases and sequencing. Nathan argues that results like this mean we "are not going to be able to hide behind" the idea that AI only works on verifiable-reward tasks. The track, "Just Wanna Learn," plays out the episode.

REAPER, the digital audio workstation: https://www.reaper.fm/

## Chapters

00:00:00 Cold open
00:01:58 Welcome and cloned-voice disclosure
00:02:10 Part I · Notes from The Curve
00:04:20 Should limits apply to the labs themselves?
00:08:24 "A level of intelligence we shouldn't go past"
00:09:26 Would a lab leader accept a FLOP cap?
00:12:03 Two or three labs at the frontier · RL environments as distillation
00:14:54 Timelines · The critical period starts now
00:16:42 Compute after a cap · Monitoring and fixing RL environments
00:21:15 NanoGPT speedrun record and research taste
00:23:27 AI for science is next
00:25:19 Congress after the midterms
00:26:47 Part II · The memory wall · Tom Sohmers, Positron
00:27:47 Commodity memory and bandwidth utilization
00:31:23 Asimov · Memory capacity per chip
00:32:38 Looped transformers
00:36:44 Design vs. verification · Agents on the Palladium emulator
00:40:41 Token spend vs. salaries
00:42:40 Has AI out-invented the chip designers yet?
00:44:49 Part III · Software after agents · swyx
00:46:33 Taste over tenure · Modules, not lines
00:49:39 The CPU shortage
00:51:33 Kill My SaaS
00:55:02 Is the software pie fixed?
00:58:17 Jev, system-1 models and the Decisions API
01:00:36 AI Engineer NYC
01:01:14 Part IV · Who owns the risk? · Evan Miyazono
01:01:49 Agent swarms, critical infrastructure, inference attestation
01:03:31 How our agents might connect us
01:06:03 A request: please reach out
01:06:16 Formal verification after the math shock
01:08:55 Part V · Where the training data comes from · Edward Hu, Mercor
01:10:44 SFT vs. RL
01:11:51 Reward hacking and "scattergunning"
01:13:49 Superhuman at what? Specifying "better"
01:15:39 Part VI · Not just verifiable tasks
01:17:53 The last wall

## Links

https://ai-in-the-am.com/
https://x.com/labenz
https://x.com/8teAPi
https://thecurve.goldengateinstitute.org/
https://techcrunch.com/2023/04/06/anthropics-5b-4-year-plan-to-take-on-openai/
https://helentoner.substack.com/p/long-timelines-to-advanced-ai-have
https://hyperstition.cc/training-nanogpt-in-39-9-seconds
https://github.com/KellerJordan/modded-nanogpt/pull/360
https://www.positron.ai/
https://x.com/trsohmers
https://www.prnewswire.com/news-releases/positron-ai-raises-875-million-at-a-5-billion-valuation-to-bring-its-next-generation-inference-silicon-to-market-302874601.html
https://www.positron.ai/asimov
https://investors.credosemi.com/news-events/news/news-details/2025/Credo-Unveils-Industrys-First-Memory-Fanout-Gearbox-for-Scalable-High-Bandwidth-AI-Inference/default.aspx
https://www.nytimes.com/2026/09/23/opinion/ezra-klein-podcast-jensen-huang.html
https://openai.com/index/navier-stokes-solution/
https://x.com/swyx
https://www.latent.space/
https://www.ai.engineer/nyc
https://luma.com/ls-06v7
https://typesafe.ai/blog/introducing-system-one-models-and-jev
https://openai.com/index/devday-2026-recap/
https://atlasignota.org/
https://x.com/emiyazono
https://openai.com/index/hugging-face-model-evaluation-security-incident/
https://huggingface.co/blog/agent-intrusion-technical-timeline
https://theorem.dev/
https://oath.tech/
https://x.com/edwardjhu
https://www.mercor.com/
https://arxiv.org/abs/2106.09685
https://www.mercor.com/blog/introducing-apex-agents/
https://www.mercor.com/blog/introducing-apex-agents-1-1/
https://www.reaper.fm/

Sponsors:

Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr

Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr

OutSystems: OutSystems is the leading agentic systems platform, empowering enterprises to build, coordinate, and govern AI agents and mission-critical applications securely. Learn more and start building your agentic future at https://outsystems.com/tcr

Tasklet: Tasklet empowers your business with AI agents that connect to your tools and automate recurring workflows with no code required. Visit https://tasklet.ai and use code cog rev for $50 in free credits

CHAPTERS:

(00:00) Weekly highlights preview

(01:57) Notes from The Curve (Part 1)

(14:22) Sponsors: Parallel | Claude

(17:17) Notes from The Curve (Part 2)

(17:17) Timelines and compute caps (Part 1)

(28:40) Sponsors: OutSystems | Tasklet

(31:36) Timelines and compute caps (Part 2)

(31:37) The memory wall

(40:50) AI in chip design

(48:17) Software after agents

(54:33) Kill my SaaS

(01:03:32) Who owns AI risks

(01:10:58) Workplace training data

(01:17:36) Beyond verifiable tasks

(01:20:56) Episode Outro

(01:25:58) Outro

PRODUCED BY:

https://aipodcast.ing

SOCIAL LINKS:

Website: https://www.cognitiverevolution.ai

Twitter (Podcast): https://x.com/cogrev_podcast

Twitter (Nathan): https://x.com/labenz

LinkedIn: https://linkedin.com/in/nathanlabenz/

Youtube: https://youtube.com/@CognitiveRevolutionPodcast

Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431

Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk


Transcript

This transcript is automatically generated; we strive for accuracy, but errors in wording or speaker identification may occur. Please verify key details when needed.


Main Episode

[00:00] Nathan Labenz: This week, my notes from The Curve, the conference where Frontier Lab insiders and their critics share a room, plus the memory wall and what agents are doing to software. This is AI in the AM.

[00:14] Nathan Labenz: There likely is, let's say, a level of intelligence that we just shouldn't go past.

[00:21] Thomas Somers: What?

[00:22] Nathan Labenz: Yeah. First time I never heard that from a Frontier lab leader that I just asked to follow-up. Like, okay. Well, how do we operationalize that? Would you be prepared to sign on to something like a limit on the number of flops that go into the next pretraining run? But what actually happened was he basically said, yeah. I think that could be reasonable.

[00:44] Nathan Labenz: Prakash Narayanan, my cohost, on what a ceiling would actually require.

[00:49] Prakash Narayanan: When we say that we shouldn't go exceed a point of intelligence, that automatically means that you have to look at the inputs going in.

[00:57] Nathan Labenz: Thomas Somers, cofounder of Positron, which builds AI inference chips around memory.

[01:02] Thomas Somers: I was looking at a quote that we had a year and a week ago and... For memory, and it's gone up four and a half times since then. I really hope that this isn't the case, but I wouldn't be surprised if this is going to end up being a... Another doubling over the next year.

[01:17] Nathan Labenz: Sean Wang, known as SWIX, who runs latent space and the AI engineer conferences.

[01:23] Sean Wang: And, you know, the the phrase I've sort of recently taken to is I don't want to pay for someone else to go through LM psychosis. Right? I can pay for my own psychosis. That's fine. But, like, when you work for me and I'm paying your tokens, you better be, like, actually producing thoughtful stuff. And I have two or three employees right now where... Who are basically under performance review because they are just giving me cloud slop.

[01:45] Nathan Labenz: Yeah. I'm, I'm... I pity our our Claude that's gonna have to try to cut this down into a highlights episode, man. I think this is banger content from start to finish today.

[01:57] Nathan Labenz: Welcome to the AI in the AM weekly highlights. The narration is my cloned voice. Please tell us what worked and what didn't so we can make the next one better. Part one, notes from the curve. Last weekend, I was at the curve at Light Haven in Berkeley. Prakash asked about the divides I saw.

[02:21] Nathan Labenz: One of the big observations in general in AI is that events have come pretty fast. And you would think if you were looking back... You know, if you were looking ahead a couple years ago that with two years of additional information and revelations and, you know, actually seeing how strong the AIs get and what they do and who's been right or wrong, you would think that, like, people would start to agree on the state of things. And that has happened a lot less in the broader world than I would have expected, but I did feel that there was some of that felt at the curve. But I do think the big thing that people have kind of come together on, at least to a significant degree, is it does look like the AIs are gonna get really powerful, and this is something we're really gonna have to reckon with and can't just sort of hope goes away or, you know, feel like maybe if we just wait, you know, the bubble will burst or the fever will break or what have you.

[03:12] Nathan Labenz: This time around, people

[03:14] Nathan Labenz: were definitely like, okay. It's getting pretty serious. You know, there's there's... There were very few perspectives that, as Helen Toner once famously put it, like, long timelines have got super short. I would say, you know, another corollary to that was, like, what is called normal in terms of expectations is getting increasingly pretty incredible. Even the... I would say the people there who had the most modest expectations for what AIs will do, be able to do were like, yeah. It's it's gonna do an awful lot. And the question now on the capabilities front has gone to, like, is RSI gonna fume? Is it gonna be fast for a while and then kinda level off? Is it... Is that more than it's meant to be? Like, yeah, we've got these super powerful things, but is it really gonna kind of run away from us? Those are still the capabilities questions that we're getting asked, but there has been, I think, some healthy updating across the community just based on what we've seen.

[04:13] Nathan Labenz: Prakash asked whether anyone from the Frontier Labs said limits should be enforced on the players in this space, themselves included.

[04:21] Nathan Labenz: Yeah. That was a huge topic. I mean... And probably the biggest divide in terms of perspective is really just inside the frontier companies, outside the frontier companies. Right? Everybody knows that they're living in the future, and they have visibility that the rest of us don't have. And so actually had a chance to talk to and hear from, in some cases, in sessions, in some cases, just in, you know, passing conversations, all Chatham House rules, so I'll abstract away. But we're talking people who are, like, founders, executives, top researchers at frontier companies, and I'm I'm... You know, small set of them. I asked a couple of them, why come here? You know, what... What's what's in it for you? You know? And and before doing a couple sessions where I was, like, interviewing or moderating, I asked the same question. What what makes this next hour a good use of your time? And they repeatedly expressed that it's just important to them to share as much as they can with the broader community. Definitely not sharing their, you know, trade secrets, of course, still, but really trying to give us a a candid take on what are they seeing internally, what do they believe internally, you know, what's their outlook, what's the plan, how can we come together and, you know, potentially do some of these constraints. And there were a number of things that I actually thought were, like, kind of newsworthy. One, let's say, senior executive at a Frontier Lab, first of all, was was saying that just pretraining continues to deliver, That the models are just at the base level of capability, getting stronger and stronger, and it doesn't seem to be an obvious stopping point to that. Scaling laws are holding or maybe even bending because the data quality is getting better. They've got, you know, increasingly different architecture tricks. Lots of work on synthetic data. Synthetic data seems to be definitely really delivering. You basically... You need more pre training data. What do you do? You, like, give it a... Give the model a bunch of data and think, what what could I transform this data into that if I then pre trained on it would make the model smarter? And you're sort of converting these reasoning tokens. You're... Test time compute. You're converting back into pretraining data, which then, you know, goes in, of course, to the next generation of models. So this is, like, one way in which recursive self improvement is happening. Not even necessarily... There are, of course, opportunities for architectural improvements as well, but a big emphasis was just on data quality. We can now just spend tokens on the existing data to augment it, to to create better, cleaner, different lenses on it. We feed all that back in, the... And the the models get better. So this this one executive was really emphasizing, pre training continues to deliver. The models are getting super powerful. RL is important too, but his take actually went as far as say... As he believes that the latest frontier models do have the ineffable research taste that would be needed for them to start making paradigm level breakthroughs. He believes it's currently an elicitation problem where the RL is not bringing it out well enough. They don't know quite how to bring it out well enough, and so it's kinda rare. And this is why you need 10,000 agents, you know, to do a Millennium Prize problem. And, of course, that's verifiable too. So, like, RL is is working better there. Maybe you need a million agents, you know, at this point to, like, stumble onto something eventually where you'd get the kind of paradigmatic change at a at a, you know, ML research level. But he believes that is there, and it and it will be... They'll figure out how to elicit it. And then he said a couple things that I thought were just, like, legitimately newsworthy. One was he believes that there is a level of intelligence or there is... Might be a little strong, but there there likely is, let's say, a level of intelligence that we just shouldn't go past.

[08:20] Thomas Somers: What?

[08:21] Nathan Labenz: Yeah. First time I'd ever heard that from a frontier lab leader. And it was not, like, minced words. It wasn't a, like, there's definitely a level or I can articulate what that level is. But, like, this... You know, person who, again, like, everybody would recognize, you know, their name and position saying, I think it is likely that there's a level that we shouldn't go past. Now was that not ever or just, you know, not right now? I think that was also a little bit ambiguous, but it was a it was a pretty firm statement. Definitely the the strongest statement I've heard about imposing kind of a hard cap on capabilities even for a time.

[09:02] Prakash Narayanan: What confuses me about that that statement is that some of the increases intelligence that we were seeing, number one, they're coming as a result of compute... Increases in compute, which has traditionally been what has happened in the last, you know, decade and a half. Number two, they are coming as a result of refining the input data. Like, the input data is being refined by greater intelligences, and so you have input data which is much better. When we say that we shouldn't go exceed a point of intelligence, that automatically means that you have to look at the inputs going in. And the inputs going in are gonna be compute and this refinement of data. And so if you say that we shouldn't go beyond that path... That point, that also means that you shouldn't build compute beyond that point, or you shouldn't refine data further beyond that point, which means a lot, really. It means, like, the end of the end of, like, the compute build out, which is pretty significant. Right?

[10:00] Nathan Labenz: Yeah. Well, I don't... I mean, I think there's a lot of uses for compute. So I had a chance to follow-up on this. Okay. This this combination of pretraining is really delivering and scaling laws. If anything, maybe are bending more favorably. Although it's, you know, it's tricky to know what you count. Right? Do you count all those flops that you use to, like, create the synthetic data? Maybe you should. How to count these things is always super vexing. But between this pretraining thing and the... There there might be a a cap on how far we should go that I just asked to follow-up. Like, okay. Well, how do we operationalize that? Would you be prepared to sign on to something like a limit on the number of flops that go into the next pretraining run? You know, should we say... And I don't know what that number would be, but... And, what exactly would count? But, like, should we say no more than 10 to the 27 flops in the next pretrain or something like that? And I kind of, like, anticipated that being a bit of, like, a straw man for him to react to or, like, tell me what was bad about that and maybe ideally give me a better idea. But what actually happened was he basically said, yeah. I think that could be reasonable. The end. Not really much of a fight there at all. Again, the details will be, like, incredibly important if we go down that path because the more data you enrich, you know, the more plots you put in enriching the data. You know, you can kinda... You can definitely play a a sort of shell game of hide the compute, I think, these days. But the sentiment was like, yeah. We might have to do something like that. And then, of course, your question I can anticipate. I know you're well enough to predict the next Prakash token at this point is, what about Elon? What about

[11:38] Prakash Narayanan: Zuck?

[11:39] Nathan Labenz: Mhmm. And the answer there, there were a couple of really interesting things people said along the way. One, they said... You know, in gen... Nobody... This is not from one individual, but sort of a takeaway from the event in general is that the, let's say, two to three, Gemini four being announced now. We maybe have a three horse very frontier race again. We'll see as that comes online. But the the takeaway was kind of the two to three companies are really pushing the frontier. And even the other two that we would usually think of in the top five, Elon and Zuck, are actually distilling a lot whether they know it or not. And the way that they are distilling these days, it's not that they're calling the cloud API and getting traces and feeding them in, but rather the cottage industry of RL environment, makers and sellers is in in fact functioning as a distillation conduit. Because what they're all doing is having Claude create these RL environments, selling the RL environments to other frontier model makers, and then they do their internal RL on an environment that only exists because Claude was smart enough to make it. And that's how they are kind of re... You know, constantly capturing the the capabilities advances that the frontier... You know, the true frontier companies had. So there wasn't actually a ton of talk about competition beyond the top three. One person went as far as to say, and this was another, like, founder executive type at at one of these companies. Yes. They have kept up. Obviously, there's been a gap, you know, but that gap has been persistent, but they've kind of maintained a kind of consistent following distance. And this is in contrast to Anthropics' old prediction from their fundraising deck that I always think about when they said, you know, maybe in twenty twenty six ish, the companies that train the best models will be so far ahead that nobody will be able to catch up. Obviously, that hasn't happened. But the the kind of synthesis of all this is if they hadn't been releasing the models and if people weren't able to do these, like, various sort of direct and indirect distillation techniques, they actually think that probably would have happened. That the other companies are not really keeping up but for the the sort of leak of intelligence in all these different ways, including the RL industry that is gradually finding ways to transfer from Frontier, models to their competitors.

[14:22]Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr

[15:42]Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr

Main Episode

[17:18] Nathan Labenz: Prakash added his read on Elon's strategy, and then I got the timelines.

[17:24] Prakash Narayanan: And I think on Elon's side, Elon is very focused on the infra build out. So Elon is making a a gamble that the infra matters more than the models and that he will be able to catch up basically using his infra build out.

[17:43] Thomas Somers: And so... And he's and he's structuring things for 2029,

[17:47] Prakash Narayanan: 2030, 2031, like, TeraFab, things like that will come out much later. The SpaceX, you know, satellites, that's like a twenty thirties thing. So he... He's he's structuring for a much longer period of time. And I think Anthropic and OpenAI are both structuring for more like the two two to three year mark, and they're looking at a 2028, 2029 AGI RSI. And we

[18:08] Nathan Labenz: Yeah. Honestly, even sooner, I would say, was kind of the vibe. And this is probably somewhat biased to be a short timelines crowd. But, again, like, the short timelines crowd includes the executives, top researchers, founders of these companies. They were talking more about just next year.

[18:26] Sean Wang: Yeah.

[18:26] Nathan Labenz: It was really a kind of 2020... Like, critical period is beginning now. RSI is kind of at hand. Definitely not full agreement on how fast it goes or if or when it tops out or whatever, but a shared sense that, like, the decisions that get made over the coming months and the way that compute is used over the next year could be really critical. There wasn't really even too much 2028 talk, to be honest, which is, again, like, pretty crazy.

[18:56] Nathan Labenz: Then it came back to compute and what a cap would mean for the build out.

[19:01] Nathan Labenz: But, like, in terms of will they still have need for computer or does a cap on pretraining scale or other, you know, compute input cap that might constitute a pacing mechanism, does that mean the end of the compute build out? I would say no. Another question I had the chance to ask was, of course, everybody's heard Jensen, on his recline, I think, at this point. And the comment that he made was like, you know, these companies are gonna have to mature. They're gonna have to change their ratio of where they spend their resources. He said at NVIDIA, we spend, like, 20% of our effort on designing the chip, and we spend 80% of the effort on verifying and validating and testing all the edge cases and making sure it's gonna last a long time and be reliable and all these other things that are not just the original design. So his point was like, they've been working really hard, obviously, until this point, and they've been putting everything they can into making their models capable enough to be useful. Congratulations. You've succeeded. Now you're gonna enter the era where you're gonna have to make them safe, reliable, trustworthy, etcetera. And so I had a chance to ask, you know, one of these people, like, do you think about that take from Jensen? And the response was basically like, yeah. That's that's kind of reasonable, you know, without, again, knowing exactly where the numbers will shake out. The possibility that the majority of could compute in the future could be going into all sorts of different safety measures, you know, which could be monitoring, could be chain of thought monitoring. People are pretty... Getting kind of bearish on chain of thought monitoring.

[20:46] Prakash Narayanan: Yeah.

[20:46] Nathan Labenz: But then there's internal, you know, activation monitoring. They're already spending a significant percent of compute on monitoring, and it sounds like that might go up a lot. Another thing that they're spending compute on in a big way right now is fixing the RL environments. There were... You know, of course, everybody understands at this point that if you have sloppy RL environments that reward cheating, then you're gonna get a lot of cheating. They are now applying these models to the RL environments themselves and saying, hack this environment. And now that is the direct task, they're doing it. They're finding, you know, all these places and ways in which the environments can be hacked. They're fixing those. I'm sure they're throwing some environments out that are just fundamentally flawed, whatever. And, gradually, they're driving down the rate of flaws, which means that they're also driving down the rate at which cheating would be rewarded, which in turn means you're gonna get less cheating from the models as they actually come online. And it... I got the sense, although it it was... You know, I think this is still, like, as they say, an open research question. But I got the sense that they are seeing a pretty clear relationship between the more you clean up the RL environments, you do, in fact, get lower rates of cheating downstream. And they're trying to push... Obviously, they'd love to have zero hackable environments. That's gonna be tough. But it doesn't seem like we're all the way to the point where we have, like, a scaling law, but, you know, kind of we're we're seeing sort of a proto scaling law, it seems like, in... You know, again, the way you can spend a lot of compute driving this down and getting better behavior on the other end. I heard repeatedly that they expect to be bottlenecked on safety and alignment, but it seems like for at least the next generation or two, there is still enough low hanging fruit to pick in the safety and alignment domain in part just by fixing RL environments

[22:52] Thomas Somers: that

[22:55] Nathan Labenz: we should expect things to continue to move pretty fast. But it seems like they're probably going to be able to get comfortable releasing the next couple things even as they do consider themselves bottlenecked on alignment. But then past that, they're like, yeah, all bets are off. We we... Really hard to say what happens kinda past the next one to two generations.

[23:16] Nathan Labenz: Later in the show, I tied that executive's view on research taste to a new record in the NanoGPT speedrun, the race to train a small language model to a target quality as fast as possible. The record comes from a company called Hyperstition.

[23:32] Nathan Labenz: But just headline wise, it seems like we've gone from something like seventy seconds on this, this classic benchmark. This is something people have been working on for quite a while. And, basically, the the idea is to train a small model to get to a certain loss as fast as you can in terms of wall clock time. And, this result here took almost half the time off, which was more than, I think, the last 45 improvements combined, and this one has been accepted. One thing I did think was quite interesting about this particular one that has been accepted and kinda seems to have kicked off this wave of, you know, significant drops in in the time was that the author here said, AI didn't play a big part in it. Like the person senior executive leader at the Frontier company that I mentioned earlier, their perspective was they believe that the model probably can come up with these quality of ideas, but it's an elicitation problem, and they didn't really rely on AI super heavily in this case. You know, obviously, they're getting lots of help. But in terms of core insight, it was mostly attributed to the human, not so much to the AI on this particular drop. This company is doing also, like, large scale training. So they did also make a point in their publication about this that the optimizer that accounted for a significant, but I think not majority of the time savings that they did share is actually not as good as the one that they are using internally.

[25:22] Thomas Somers: Then

[25:22] Nathan Labenz: one last expectation I heard across the frontier companies at the curve

[25:26] Nathan Labenz: about science. AI for science is very much where they think we are headed in the not distant future. Like, one one sentiment was, you know, our AI is gonna be superhuman at just some things, like the verifiable tasks, or are they gonna be superhuman at everything? And the the sort of middle position, which I think is very credible, is we are seeing superhuman performance at anything we care about and are, you know, really committed to investing in. So that doesn't mean that, you know, it's full generalization to every domain. But even in domains that are thought of as not super inherently verifiable, they feel like when they put their focus on it and they go license what data they may need to license. And they, of course, apply lots of processing to that data to augment and, you know, create synthetic versions and what have you and, you know, put some RL on it too. The whole package basically is, I think, broadly

[26:38] Nathan Labenz: understood

[26:40] Nathan Labenz: to

[26:40] Thomas Somers: work

[26:43] Nathan Labenz: on essentially any problem that they really choose to focus on. And so AI for science is coming next, and the expectation is we will see, one person went as far as to say, in no more than a year, it will not be possible to do the very best science without AIs playing a big role in it.

[27:10] Nathan Labenz: The next morning, I added one more takeaway from the curve.

[27:14] Nathan Labenz: And I don't think we talked about this yesterday, but a takeaway from the curve this weekend also. One of the... Not main threads that I was involved in discussing, but a general vibe was there is going to be government intervention. Mhmm.

[27:30] Thomas Somers: And,

[27:30] Nathan Labenz: you know, obviously, we've already seen some, but the idea that, like, congress will never act, I think, was... That used to be sort of taken for granted. And now the new vibe that I was hearing was, like, after this election, especially if the Democrats take both houses, both both chambers, that may change. And you might even see a bunch of Republicans join the Democrats because, really, in their heart of hearts, they wanna do something. They don't really wanna go against the president, and there hasn't been a lot of people willing to take that chance yet. And they're also getting a ton of money from the data center interests right now, but that all could come up for a renegotiation and realignment after the midterm. And we might actually see congressional action a lot sooner than maybe a lot of people who've grown accustomed to the idea that it can never happen or it'll be, you know, stasis forever might think. So I do think policy response is definitely part of what we need to be looking out for.

[28:40]OutSystems: OutSystems is the leading agentic systems platform, empowering enterprises to build, coordinate, and govern AI agents and mission-critical applications securely. Learn more and start building your agentic future at https://outsystems.com/tcr

[29:59]Tasklet: Tasklet empowers your business with AI agents that connect to your tools and automate recurring workflows with no code required. Visit https://tasklet.ai and use code cog rev for $50 in free credits

Main Episode

[31:37] Nathan Labenz: Part two, the memory wall. Thomas Somers is a cofounder of Positron, which builds chips for running AI models and recently raised $875,000,000. Prakash asked whether the sharp rise in memory prices had changed the economics.

[31:57] Thomas Somers: I I was looking at a quote that we had a year and a week ago and... For memory, and it's gone up four and a half times, since then. I really hope that this isn't the case, but I wouldn't be surprised if this is going to end up being a, another doubling over the next year. But I would say a big reason why everyone has burdens these these cost increases is because you're deriving more than that for five x value compared to last year. But I think the the reality there is compared to a year ago, the capability of a model using x amount of gigabytes of memory is way more than five x than than it was this time last year.

[32:36] Nathan Labenz: I asked Tom about the trade offs of building on commodity memory instead of the high end kind.

[32:42] Thomas Somers: At the start of Positron, I did think that memory was going to be the, limiting factor for things going forward even though, you know, people can debate on power and and other parts infrastructure stack. My fundamental belief was that there isn't going to be a stop anytime in the near future in terms of scaling laws and getting increased value from making models larger. So that eats into memory on one side. But then the other side, I think the main limiter today of AI being applied to most applications is context length and being able to hold more context, you know, per user and then scaling that out to drastically more users. And and, really, today, I would say that the driver of, quote, unquote, users or individual sessions is actually just having more agents. I've gone from, you know, three, four months ago where I had probably on average two to four agents, like, running constantly in the in the background to where I'm now running 15 to 20. If you then multiply how many people are using AI and how many... That that concurrent completely separate contexts adds up very quickly. That that memory driver had us say, okay. We have to use the commodity number because that's the only thing that's going to be able to scale and be cost effective. And when we set that as sort of the constraint in architectural design, we had to come up with, you know, very, very clever innovative solutions. And and, you know, the the two main pieces of that and, you know, the thing that we showed with our first generation product is being able to achieve extremely high memory bandwidth utilization from a compute architecture perspective. So while NVIDIA GPUs, I I would say, on average, are getting between 3040% memory bandwidth utilization in the actual decode forward pass of a transformer model. Basically, that means even though they advertise, you know, eight terabytes per second of theoretical memory bandwidth with, you know, a b 300, you only actually are seeing, you know, something in the ballpark of, like, two terabytes per second of actual realized memory bandwidth. And that comes down to a whole bunch of GPU architecture details. It comes down to the reuse patterns that really don't exist in transformers that they've designed, their hardware architecture around for training and and other workloads. And so with our first gen product, we were able to hit and sustain 93% of theoretical memory bandwidth. So that theoretical and realized basically, you know, become one. And so that isn't that massive, you know, three x improvement right there. But to really take advantage of, you know, the belief that models and everything are just going to get larger, we have to be able to scale to to massive more chance. And so we're, you know, partners with Credo Semiconductor that's and and have developed a memory chip with solution that allows us to go from the maximum number of LPDDR, same type of memory that is in phones, laptops, etcetera. The max number of channels that you can have... Find in any other products is on the order of, you know, twelve twelve to 16 channels. And we go up to 72 channels of of LPDDR five x with this decoupled memory chip solution.

[35:49] Nathan Labenz: Positron's next chip is called Asimov. Prakash asked what disappears in networking and coordination when a whole model fits on one chip?

[36:00] Thomas Somers: We have up to 2.3 terabytes of memory capacity per chip. So when you compare the the b 300 shipping today, caps out at 288 gigabytes. As, you know, reported in in... By analysts, NVIDIA for next generation is actually cutting down the memory capacity because cost and everything involved. So they're only going to be at a 192 gigabytes with the Ribbon Ultra. And so... But even if you use that two eighty eight as, like, the current benchmark, we've got eight times more memory capacity per chip. And so that means on a per chip basis, you now can actually scale in what you would have needed from a memory perspective. Eight GPUs for, we can do on a single device. And it's not just, oh, you've got that cost silicon savings, you know, as as you pointed out. When... Whenever you have to scale to having more than one device, there's overheads involved with that. There's, you know, your all gathers, all reduces that you have to do within... For each layer and each... Really, each multiply that you're doing of sharding across those devices.

[37:00] Nathan Labenz: Then I asked Tom about looping, where a model runs the same layers more than once for each token and why that is tempting at all.

[37:09] Thomas Somers: It's been very interesting from the loop transformers paper to that being the rumor of, you know, one of the big advancements with the GPT six Astra. Simply repeating the forward pass, it gets you that improvement. But it's crazy that there is really work that I think was only being done in sort of the fringes of the open source, you know, transformer community, you know, two, three years ago of taking, like, a LAMA 70 b and just duplicating layers within it. So it was taking, like, a 70 b model and slightly strategically saying, okay. I'm going to just repeat these these sets of MapMoles in in sequence and turning it into, like, a 100,000,000,000 parameter models. And then you you got better results from it even though it was repeating the exact same MapMoles. So there there is an element

[37:57] Nathan Labenz: Imagine what you can do when you train it to work that way.

[37:59] Thomas Somers: Exactly. If if you already have very high confidence on what the next token is going to be and multiple layers have agreed on that, decide to early exit and go out. Or if you're really uncertain about something, actually find which layers in network actually will have a higher likelihood of being... Will increase that probability, same sort of, like, functions that the MOE router does in terms of determining which experts should be selected for... On a per token basis. Do you extend that concept to... Okay. If you're really unsure if there's a a very large set of of equally weighted probabilities for what the next token should be in your three quarters of way through through the layers, You do actually just wanna route back and determine, okay. We actually need a different set of experts something like for this. But I wouldn't be surprised if these things are already being done at scale in in the major model labs.

[38:50] Nathan Labenz: But can you go just a little bit farther into the hardware connection of that? Because I'm kind of coming in with assumptions along the lines of, you know, just very fundamental terms like why loop at all. I think it has to do with this kind of memory bandwidth. Right? Because now I can keep the same weights on the chip, and I don't have to, like, shuffle things in and out as much. No?

[39:13] Thomas Somers: Not really because within... When... Basically, the assumption is if you're doing inference that you're going to have all of your weights locally in in in Deepgram, you're just sharding that over some number of devices. You do get some amortization, but just with the size of experts and the sizes of caches on chips, there's not much of a reuse opportunity or by the time you're already done with the layer, you've already gone through the amount of memory many times that of what is in, like, the on chip s rams. From from, like, a hardware perspective, like, looping, really, is more memory capacity savings than memory bandwidth savings. Because, like, let's say that you trained two models with the same base set and you trained one to do with looping, and so it only has 10 layers instead of 20 layers. And, you know, model b is is 20 layers. That 20 layer model, it's doing very, very simple, is going to be double the size of the the 10 layer version. And if you find that the linked 10 layer version gets 95% of the same quality of results and it's half the size, you're probably going to deploy that one. You're you're saving more capacity than bandwidth. Because doing the second loop, you still have to do the same memory fetches and do all of the same, like, same number of MapMoles. Like, there's the exact same number of bytes that need to be moved and the exact same number of flops that have to be done in that model a and model b scenario. It's just that model b, the 20 layers, actually has unique different weights for the the second group of 10 layers that have to be done. And so that that means you you require twice the memory capacity.

[40:51] Nathan Labenz: Then I turned to how AI is changing chip development itself.

[40:56] Nathan Labenz: In a notable interview, Jensen Huang recently gave with Ezra Klein. He talked about how NVIDIA spends something like 20% of its time and energy designing and then 80% verifying, validating, you know, ensuring long lifespan reliability, etcetera, etcetera. What do your ratios look like on that?

[41:15] Thomas Somers: So for us, I would say it's a little bit different because we're starting from scratch. So there's a lot more, like, base layer... Base capability design that we have to to start from scratch. So for, you know, this this, our our first full custom silicon Osmov, it was about a... It'll it'll end up being somewhere around maybe 60% design focus for us to 40% verification. I would say, like, the probably most amazing thing that I I don't think would have been possible three or six months ago and only really became possible with the GPT six Astra and open... Now OPUS 5.5 that that we're leveraging a lot in in this task is we... For doing our verification work, we use Cadence palladium emulator systems. These are big giant racks that are full of custom ASICs that are... That's specifically designed to just emulate... Do gate level emulation of of silicon. The crazy thing about Astra and five... And OPUS 5.5 is this fact that your... The the agents themselves have been able to completely do closed loop iteration and testing on on on this. So, you know, I I think people thought it was crazy, but, you know, handing, you know, these very powerful agents, you know, full keys to our our internal infrastructure and saying, here's the Palladium piece. And I highly doubt that these models had Palladium documentation in their retraining thing, especially because a lot of documents are are new as as software updates. But the fact that it went and looked up the documentation when we're looking at these agent traces, the reasoning traces and and tool calls, etcetera. It's reading the full documentation, effectively compressing that itself by by reading the PDFs, generating its own, you know, markdown files of of cheat sheets of what to do, and then building out all of the testing infrastructure in its own harnesses for being able to access this. It's mind blowing. And so while... If you're just doing regular software RTL simulation, that's running on the order of, like, ten ten hertz for us in in in those cases. If you remove a lot of the debug pieces that slow things down, etcetera, we can run the full chip emulation on the order of 500 kilohertz. And so you've got actually several orders of magnitude, speed up, an improvement from that gate level hardware piece, and this just gets us closer to this this RSI loop. And right now, that's really just focused on implementing test programs, finding cases where they thought... Fail, and then writing out reports, which, you know, get reviewed by other agents and and by humans in in the loop. But, I I would say when this really started working six weeks ago, we got roughly... Beginning of September when when Astra six came out, it was a huge, like, step function improvement from GPT 5.6 where it could not do that full closed loop. It's... It... Itself. It would... It still required humans at that at different stages.

[44:28] Nathan Labenz: I asked how positron spending on AI compares with its spending on people.

[44:34] Thomas Somers: So on on the token spend side, it's become, you know, our single largest non manufacturing line item very quickly.

[44:43] Nathan Labenz: Bigger than bigger than human salaries.

[44:45] Thomas Somers: Yeah.

[44:45] Nathan Labenz: Wow.

[44:46] Thomas Somers: Just recently. And it's interesting. It it eclipsed human salaries. It came back down. I'll explain in a moment. So, yeah, so our token spend just massively increased where where... Okay. Six months ago, it's equivalent of one employee. In in June, it started to now be multiple employees. And, like, we we were peaking, you know, in the in the days after GPT SuccessFit, and we're really trying to to push its boundaries, etcetera. We're we're spending over a $100,000 a day on tokens. That tempered down a little bit in in the weeks after as we optimized a bit and just weren't trying to run as many parallel experiments. But I would say the big thing that was, like, you know, a saving grace was, you know, Opus 5.5 came out. And in a lot of our tasks, not everything, it was doing better than than Astra and was, you know, a quarter of the price. We didn't really... We didn't tell anyone to cut their spending or do anything even though we were seeing this exponential growth in our token spend. It was really just the fact that we both felt that we were getting good ROI on it, so we weren't going to taper it. But personally, like, my philosophy is I really don't want anyone to be running a real software or hardware development task on anything that is less than the best model. Like, I I don't care if it's a tenth of the cost, you know, per token. It's just not worth the the the expense.

[46:15] Nathan Labenz: Prakash asked whether Tom has seen AI come up with genuinely unexpected ideas in chip design rather than following the rules in a manual.

[46:26] Thomas Somers: I have not yet. I would say that's the most disappointing elements of all. Like, I'm... As you can probably tell, I think all of this is insanely amazing, and I think it... We're going to continue this exponential trend. I'm somewhat proud of the fact that it has not figured out some of our cleverness of what we've done. Like, it's... It has questioned and, like, this seems like a bad design decision until it gets explained to it why something's done or it actually runs, tests, and sees and understands, oh, that's why you're not doing the traditional way a systolic array is done. Do I think that will hold for another six months? I'm not sure. But just because I I... In part why I actually think it's not going to be, like, the agent will come up with a better idea immediately. I think the amazing turn of events that we've just had in the past few weeks in terms of us having access to these models is the fact that just because it can iterate so fast and do experiments and things by itself, I wouldn't be surprised if if it would come to the same conclusions we did or maybe, you know, make a better things than what we would do by ourselves just because it's going to be able to iterate so fast and and go through so many different design possibilities. I I think, like, the... Just the story of of OpenAI solving Naviar Stokes, like, the same approach of, yeah, if you take 10,000 agents and and send them off for, you know, millions of of man years worth of effort, that that they found a solution. I... It's brute force. It's expensive, but it... I think that is a valid strategy that has only become possible in in very short past.

[48:05] Nathan Labenz: Not many people left standing in the AIs can't do what I do category, so congratulations on your place for as long as it

[48:14] Thomas Somers: it for a couple more weeks at least.

[48:17] Nathan Labenz: Part three, software after agents. Sean Wang, known as Swix, runs the Latent Space Podcast and the AI Engineer Conferences. I ask how software engineers are feeling right now, empowered or threatened?

[48:35] Sean Wang: Yeah. I think college kids are a bit worried. But other than that, if you're relatively plugged in and very capable of AI engineering tooling, you're in more demand than you've ever been because your expected value is higher than it's ever been. So this is definitely one of those Jevon's paradox type things where the sort of cost of software... Creating software has gone down. Therefore, the demand has increased a lot. And... But but specific... For a very specific kind of demand for people who can manage coding agents productively instead of producing a whole bunch of slop. And I've been in that situation. So I am both an engineer, but also an employer of engineers and people who are not engineers who are by coding. And, you know, the the phrase I've sort of recently taken to is I don't want to pay for someone else to go through LM psychosis. Right? I can pay for my own psychosis. That's fine. But, like, when you work for me and I'm paying your tokens, you better be, like, actually producing thoughtful stuff. And I have two or three employees right now where... Who are basically under performance review because they are just giving me card slot. And it's really bad for them. And they don't understand it. They don't see it. They're like, what do you mean? I think it is... This is perfectly fine. And then, like, well, you're not producing any value. Like, I can just prompt cloud. I don't need you.

[49:55] Thomas Somers: I asked whether

[49:56] Nathan Labenz: years of coding experience still matter when he hires or whether product sense and high standards matter more.

[50:03] Sean Wang: I I don't I don't care so much, except for, let's call it, security roles and for anything sort of back end scalability. Like, I just recorded an interview with the super based founders who who, you know, are scaling Postgres to a scale that we've never seen before. Good luck trying to hire somebody to scale Postgres, you know, with with a complete self managed sharding solution that works at YouTube scale. There's only one person in the world that can do that, and they hired the guy. Like, you better go and, you know, make sure you don't lose data, that you, like, don't have... Don't cause downtime, that you scale economically because, basically, agents consume databases something like at a 500 to one ratio of what what human developers do. Then... Yeah. So so, like, other than that, like, you can mostly by code, and then it's... Then it is about taste. And it's not so much about length of experience. Actually, sometimes length of experience works against you because you just have a set way of doing things, and you don't understand how to do more than one agent at a time. Yeah. At this point, like, you should be relatively comfortable juggling, five to, like, 10 ongoing multiple things, and it's it's definitely true that you're much more of a manager than a... As individual contributor. I do think that people who look at data rather than code are more valuable these days. So, being able to say, like, here's the logs, here's the traces, here's the schema, here's the input output, capturing it, turning into an eval. All these are basically the sort of the merging of AI engineering and ML engineering that is that is happening that people need to upscale on. I think that the the person that has, let's say, ten, twenty years of experience in software engineering really reviews every line and tries to make every line make sense. Whereas now we just need to make the the modules make sense, And I I I can allow slop in there because it helps me to go faster as long as I contain the slop in things where I completely understand the whole system. Where you go wrong is you have too many modules, too many black boxes where you don't even know, and, actually, the code also gets confused. So I'm making my own, sort of Slack competitor. And one of the... And and, like, I saw a bug where, like, the the messages weren't loading. I refreshed, and the messages were loading. I refreshed again and the messages were not loading. And it's like, well, it's the same exact code. What the hell is going on? Turns out there's two code path... Code paths and there's a race condition. Why? Because two different coding agents worked on it at different times. And I guess they they just, like, made their own thing. So, like, a human would never do that. An agent would potentially sometimes do it because, like, sometimes just things fall out of the context window. But, like, you need to have the oversight of the module, and then whatever inside is is... Inside of the module can be a black box.

[52:50] Nathan Labenz: Prakash asked, which parts of the stack are poorly designed for agents now that agents are becoming the main users of the Internet?

[53:00] Sean Wang: Compute. Everyone knows the GPU shortage. Everyone knows the memory shortage, but CPU shortage is the most consistently reported thing among all the people that I interview right now. So, what does that mean? It does mean that, we need sort of more fluid compute, which is the first sell term for it, more sort of different forms of serverlessness where, like, it... It's very ephemeral. You can sort of execute and resume because, you know, agents need, some time to do LLM calls or to make network calls or whatever, but, basically, long running states that are resumable. But I think beyond that, you really start to get into two things, which I'm really thinking about, like, sort of bandwidth or networking where the latency starts to really matter. And that's a opt... That's a optimization of, like, well, what point of presence are you communicating with, or what region are you in versus, like, actually, how much compute should you be doing in cloud versus, locally or sort of edge devices? That is definitely happening a lot for the robotics companies that we're talking to. And, basically, don't treat it as just a robotics problem, but the robotics are just the first harbinger of what that you guys will eventually be doing if you consume or do enough inference at their scale. And then finally, to me, on a more... Very, very personal basis, bandwidth. So, like, how many tokens per second can I get? Based... Both on a single call basis and on just the aggregate across all my calls and then aggregate across my company because I have so many... I have enough people managing all these calls.

[54:33] Nathan Labenz: Swix has also been running kill my SaaS, a bounty for anyone who can replace an events software subscription his company was unhappy paying for. Prakash asked him about it.

[54:45] Sean Wang: Basically, you can call this, like, a, you know, a bounty on mid tier SaaS that should not exist. Basically, like, half a salary or one third of a salary for a piece of software that I don't own that nobody enjoys using. You know, if I'm gonna spend 40,000 in tokens... $40,000 on this SaaS subscription, I can spend that in in amount of tokens, and that buys me a heck of a lot of tokens. We had so many submissions that we then have to eval them. And and so evals are the problem. And because... Like, we're no longer checking, like, here's, like, two or three requirements that you fit. It is... We're checking the the entire UX of the entire flow from three different perspectives of, you know, organizer, attendee, and and sponsor or a speaker. And just be aware that, like, if you offer, like, a large bounty, like, offer $10,000 as the price, you get a lot of submissions because people wanna vibe code for fun, and a lot of them will be low quality. They would just be like, Claude, make no mistakes. Go do this. And, like, in that... And then they will... It'll make mistakes, and they'll just submit. And so, like, the the sort of verification load is very imbalanced because they spent zero thought on this thing. They just threw it in there, and then you have no idea if it's, like, low quality or not. But overall, very successful, and the team initially... So I have one of the most adversarial non AI teams in AI. Okay? I hire... Like, what is my business? I hire event professionals who are, like, super old school. They do everything in spreadsheets. They work with unions. They have to worry about, like, placements of, like, the meter boards physically in, like, locations. Right? Like, these are not glamorous, like, high-tech jobs, and they're very super suspicious of anything that is new and and vibe coded and techy. Right? So I hire these guys, and then I I, like, make them work for AI. Right? That... That's my job. And so, initially, they were like, we'll never use this vibe coded thing, and this is like, I wanna use the tried and tested stuff that, like, has been, working for Microsoft and all of these other guys. And then they saw the quality of the submissions, then they they looked at their existing platform. They were like, yeah. Okay. We're gonna switch. Right? Like, it's like that initial hurdle needs... You basically need to provide the sort of proof of, like, this is feasible. And the... And then then the ongoing thing, which is... This is one of the benefits of me partnering with Cognition on, is I just give them... Give all of them access to Devin to modify the code. So anything they don't like, they can just request the change, and they have it pretty much in one to two hours, which they've never had before. Just to give you an idea of the way we work with these SaaS companies. Right now, when we request a change, they'll say, okay. That sounds pretty cool. It's on our q three road map. And, like, we don't have confidence that it'll actually be done in q three. And by the way, is done. We requested for a thing in q one where they just landed. So, like, SaaS is quite cooked if you are mostly a CRUD app. I would say that there's a lot of UX issues left. There's... You know, if you look at Suitebench and you're like, oh, wow. We're, like, 90 in Suitebench, you're looking at the wrong thing, my guy. You know, if you haven't actually tried to completely vibe code a SaaS that you use, you don't understand how models are still very bad at all this.

[57:46] Nathan Labenz: Then I pressed SWICS on the downside.

[57:50] Nathan Labenz: But I have a hard time seeing how we don't kind of pick a lot of this low hanging fruit and then end up in a sort of... A lot of companies go out of business. And, also, I'm thinking, like, layoffs from, like, Meta. You know, that's another kind of canary in the coal mine. Like, are we gonna see mass layoffs from big tech? How long can the party go on? The... I mean, the the simple answer is, like,

[58:12] Sean Wang: if you look at it as a fixed pile of software engineering, then that is the conclusion that you will arrive at. But if you look at it as the point of, like, my my competition or my TAM is spreadsheets, all spreadsheets made by anyone for... With any sort of productivity tooling can be turned into custom software for them to put on rails with a beautiful UX and with, like, automations. Or or, then then then the then the the demand for custom software becomes a lot larger. And more beyond that, if it's if it's phone calls, if it's emails that go back and forth and eventually physical meetings between people, all of that can be incrementally turned into more and more and more and more custom software and hardware and models, by the way. And so that... Basically, like, we are just growing into the long tail. Right? Like, why is Salesforce so damn big? It's because people can customize Salesforce. But at some point, they need to stop paying the 300 k a year baseline subscription and spend 30 k building their own personal CRM, and they're good. In fact, they're more than good. They're happier because they can now modify it to whatever they want. Like, the sheer amount of customization that people can do, we don't even have the same amount of customization in our software as we do our handbags. What the hell? You know what I mean? Imagine, like, lug... Like, clothes has been around longer than than than computers, but we have so much customization in clothes, so much, like, demand in, like, different brands and different variety... Varieties. Don't... We don't really care about benchmarks for clothes. It's just like, this is my... The style that I like. I think what software engineering is basically only done when we basically have, like, Kate Spade versus Louis Vuitton versus Coach for software. Like, we do... We're, like, not even at the point of, like, looking at features or, like, price. We're just like, what's the brand feel? And, like, I identify more of this brand. Like, that... That's what really sort of true commodity software is, which is you start differentiating on substance and you start differentiating on vibe, which there is a lot of vibe right now. And the the reason I said, you know, this whole thing about, like, scaling my personal token bandwidth versus scaling my organization token bandwidth, the real limit on all this is the the the sum total of humanity's total token bandwidth, right, which is, like, literally the amount of silicon we can produce. And that's that, like... That's gonna dictate pricing and availability and rate limits and all these things. And so, like, the the real players in the room are only focused on that, and then the rest of us are just fighting for scraps within the fixed pie that we already have.

[1:00:48] Nathan Labenz: Prakash asked about JEV, TypeSafe AI's fast decision making model, and whether OpenAI's new decisions API announced at Dev Day is the same kind of thing.

[1:01:00] Sean Wang: Yeah. So not only have I used it, I was friends with Diogo for four years before the launch, and so I I insisted that we were the first podcast that he doubt... He did about Jeff and that it, I mean, that is now our number one podcast ever within space. And then secondly, we... We're also in the trial for the decisions API with OpenAI. So we have a podcast on on the decisions API as well, which is not the same thing there. The... That that one is basically a Luna Luna with a a sort of different inference layer on top of it. I think what's also true among the sort of people that I talked to are that a lot of the use cases you see on Twitter are probably overdone or dumb in a sense of, like, you could use any other small model to do exactly this this thing. But because Jeff is hot right now, you're just creating a bunch of content about Jeff. Or the other perspective is this is a... Another classifier. You can train classifiers all all this whole time, but you can... But but, you know, now it's sexy. So, like, now that everyone just... People are throwing classifiers or using a classifier model where they would previously use an LLM because they didn't know anything else other than LLM. People are missing the whole point why they don't call it... Why Jeff or Yogo doesn't call it a decision model. They call it a system one model because it's meant to be integrated inside of your software to go into the background of your software as unremarkable as an if statement is the rough idea inside of your software. And so people aren't really doing that. They're like, let me just rip... Let me, like, replace this to kind... You know, every sort of point of structured output with JEV, and they're probably going to run into issues because they don't care about intelligence. If you notice most of these use cases aren't really evaluating, like, well, does it actually play it very smart? Does it do any planning? None of that. It's just, can you make a decision quickly? And, like, yeah, you could always make a decision quickly. I think there's some nuance there that's, maybe being missed. Doesn't matter because I think the category creation was so successful that it's an overall good thing for the industry.

[1:02:58] Nathan Labenz: Before he left, I asked Swix to plug his next AI engineer conference.

[1:03:04] Sean Wang: Yeah. It's, next week in New York, October twelfth to fourteenth, and we're about to sell our tickets, I think, probably by the end of the week. So definitely get your tickets if you want. There's a there's code on the latent first... Latest subscribers if you can dig in your emails, and it is, on ai.engineer/nyc.

[1:03:25] Nathan Labenz: And if you're not a latent space subscriber, what are you

[1:03:27] Thomas Somers: doing, folks? Come on.

[1:03:28] Nathan Labenz: It should already be in your inbox. That one's on you.

[1:03:33] Nathan Labenz: Part four, who owns the risk? Evan Miazono runs Atlas Ignota, a nonprofit that looks for consequential AI risks without a clear institutional owner, develops an intervention, and recruits someone to carry it forward. This is the idea behind it.

[1:03:51] Thomas Somers: I've been claiming repeatedly that as intelligence gets cheap, it's the coordination that gets expensive. And so having shelling points around who does this coordination, becoming one of those shelling points or creating one seems very useful.

[1:04:07] Nathan Labenz: I asked Evan where the biggest gaps are right now.

[1:04:12] Thomas Somers: Probably the two things that I've been spending the most time on personally recently are what can or should be done around protecting critical infrastructure, especially from, accidental agent swarms and also things like malicious use of open weight models. I think there's an interesting question of what happens if instead of OpenAI accidentally attacking Hugging Face, we have the companies behind the deep sea or any models accidentally attacking state department servers. How do you know if that's an accident? What do you do in response? What would you do if... Or what would we do if we saw some malicious attacks to critical infrastructure coming from an OpenAI server? How would you know that it was this and not coming from somewhere else? How would you shut it down quickly? What if... Like, it... With varying levels of escalation. And one of the interventions you might add is, could you use DNS to have cryptographic attestation of whoever's providing inference? I think this would be very useful and also a great complement to having DNS for people. So if your agent is talking to my agent, they can verify that this agent, in fact, is my agent. Like, your agent can verify this is Evoniazo's agent being signed off on by some key. I think that there are some shortcuts to adoption that you can take if you're not trying to build this into a startup that returns the fund. And so this looks much more like a public good.

[1:05:45] Nathan Labenz: Later, Evan raised a problem of salience. How do people even find out that a better tool or practice exists? That turned into a conversation about how our own agents might connect us.

[1:05:58] Thomas Somers: I got... Nathan, I'm sure you have great infrastructure things that I would benefit from. I don't know what those are enough to ask you what they are. I have an application that just takes screenshots of anything that's not a video conference and runs it through a local OCR model and then takes that and sends it to a... Sends it with my weekly goals to say, like, what am I trying to do? What should I be trying to do? What should I automate? What should I do differently? And I could imagine this being a thing where sharing that across collaborators would be something that could identify all sorts of interesting synergies.

[1:06:39] Nathan Labenz: Yeah. It's interesting. When we started this show, one of the ways we conceived of it was as an experiment in recursive self improvement. And we're just like, you know, can we actually make a show with two people who are live and no employees? And, you know, can we, like, iterate to something that actually works? Now I feel like maybe the next thing we need to conceive of it as is, like, a swarm, you know, like, sort of a a friendly swarm of people who can share their best ideas. Like, my my kind of upshot, you know, what will I tell Claude to do or what will Claude hopefully do for me when it sees the transcript of this conversation will be, can we, like, take stock of the best ideas that I've got and sort of put a shingle out or, you know, in some way kind of, present that to maybe not everybody, maybe it could be everybody, but certainly to my friends, the the people that I truly want, that I know and, like, really want to help. I... It feels like it would be very plausible for me to... If I recognize something I'm doing or if Koji identifies like, Evan, you are really bad at this. I could bring that to the the network and say, who is good at this? Who should I be talking to about this? Who has solved... Who who has likely solved this problem very well? There's, a reinvention of social media or a reinvention of guilds or

[1:08:00] Thomas Somers: something that would be, yeah, I think, potentially very, very useful. It's very unclear to me how society starts restructuring around some of these things.

[1:08:11] Nathan Labenz: One request from me. If you've built a good way for people and their agents to find each other and share what works safely, please reach out. Near the end, I asked Evan for an update on formal verification now that models are getting very good at math. Atlas has helped start companies in this area, so he has a stake in it.

[1:08:35] Thomas Somers: There are some meaningful differences between proving properties... Proving theorems in math and proving properties about software. One of them is that in math, the number of theorems that one might care about are pretty, they're a lot fewer than in software, and they're a lot more universal. It seems like there are a lot of projects where the ambition of the field of formal methods or at least many individuals, and I'm particularly talking... I'm thinking of Oath and Theorem and a couple of others where they're taking on things that would have been, five years, $100,000,000 efforts previously and are now things where you can possibly... Like, it's limited by tokens and the number of people who can successfully effectively wrangle agents to do these things. There's a lot of things here that I would expect that in three to six months, we're seeing the kind of the kind of shock and awe that is happening to math, like, has been happening with math in the last month, happening in verification of software improvements to stability. But I also think that the ceiling is really... There's a visible ceiling that is very high, or like a high watermark that is, I think, much better known and harder to reach than things in math, whereas you could prove that an entire computer system has mathematical guarantees around separation between different threads down to the trend... Like, down to our model of what a transistor does. In theory, you could even prove that in, like, a physics simulator and show that no ROWHAMMER exists for this soft... Or for, like, for this software hardware stack. You could also prove that we won't have... Like, this system won't crash. Or if it does... Like, if some random event leads to it crashing, we can recover everything, can roll back any software execution. I think that there are a lot of things that you could imagine that are just really totally impossible as of, like, a year ago, maybe possible now, definitely possible soon, that could lead to software engineering and software maintenance being much better.

[1:10:58] Nathan Labenz: Part five, where the training data comes from. Edward Hu leads AI modeling at Merckor, which works with professionals to build training data, environments, and benchmarks for AI labs. He also led the development of LoRa, the widely used fine tuning method. Merkor's Apex Agents benchmark gave models a simulated workplace with files, email, and chat. Edward described the next step.

[1:11:24] Thomas Somers: We have a project that we are close to wrapping up where we're actually taking it a step further where we acquired real companies. Often, have, realistic apps and environments as a part of the data we have. For example, their Salesforce data, Slack data, Figma data. And then among all of these apps, often with gigabytes or entire bytes of data, we create realistic tasks out of them. So far in the field, we more or less have been really task centric. We often start with, okay, here is a task that you need to finish. Maybe it is to generate a roster. So that's one shape, and that has been the shape that we have publicly released and most people have been working on. But, this is one of the things we're working on, but we're asking ourselves, is there a better shape? What's the shape of the task in the future once we have these dataset? Right? Like, when we work at a company, we're not always just handed these tasks. Right? We're handed a role and a work chart. And, of course, all the data in the company, and then there'll be people coming to us, asking things from us, there will be people willing to chase down to get information from. There will be managers, people in charge of resources, conflicts to resolve. So increasingly, that, we believe, is the shape of the future, and, we'll have more to share in the coming months.

[1:12:46] Nathan Labenz: I asked Edward how supervised fine tuning compares with reinforcement learning as the way companies will train models in the future.

[1:12:56] Thomas Somers: And actually, one thought experiment is if we take SFT self distillation, that's actually more or less a step of RL. And if you were to iterate this process, you ended up with a process that's actually quite similar RL. And I do think in the future, RL, the way it is done today with these very elaborate infrastructure, very, very expensive rollouts, and, relatively inefficient updates because with all of these rollouts, especially a long rollout, only get one number at the end. These type of training, will not be as widespread. I don't really see every company in the world doing these big RO runs. So I do think SFT, especially combining the output of multiple teachers, will be a key part of how enterprises in the future customize their models, and we're investing in research on that as well.

[1:13:50] Nathan Labenz: Back to the cheating problem from the curve. I asked Edward what Merkor and the industry are doing about training environments that reward hacking.

[1:14:01] Thomas Somers: Reward hacking tends to be more of a problem when the task is set up in a way that is rather... Either it's too hard, it's too underspecified, and the model just doesn't quite have a a, quote, unquote, legitimate way to solve the task. And this pressure to also get a reward in those kind of scenarios often leads to this desirable hacking behavior. And then one example in this dataset Apex agents would put out, which is quite interesting, we have a recent blog about it, where we have these professional tasks where the agent is asked to produce a financial model and we migrated by checking, does the financial model include this particular answer, which we know is the correct answer. What the model will do is given certain ambiguity and and in many cases, it's it's on us initially for not including all these specific parameters, the model will guess these different parameters, ended up giving many, many different answers. And we call this scattergunning, where the model would just say, oh, if this is true, then the answer is that. If that is true, the answer is that. That is not how a human would actually handle it in a realistic case. A human would ask for clarification. A human would follow industry standard. But in this case, just because the rubric is rewarding the inclusion of a single answer, the model ended up giving many answers up to 10 or a dozen answers. And it'll get the point if it hit one of them right. So we did a re release of the benchmark. We call it Apex 1.1, where we, number one, made sure the tasks are well well specified because that is, in in a way, the source of the issue. Number two, we have rubrics that actually penalize this particular behavior.

[1:15:47] Nathan Labenz: Remember the Frontier Lab view from the curve? Superhuman performance at anything the labs focus on with science next. The next day, I put that view to Edward, adding that in some domains, human data no longer helps.

[1:16:03] Nathan Labenz: I guess, first of all, do you buy those claims? Would you contest them?

[1:16:08] Thomas Somers: There are absolutely domains where human is not contributing to the frontier anymore. Like, one example would be kernel optimization where the goal is to write a kernel that runs faster on a particular hardware. The model can go in and write code that expert human might not even understand. But in the end, as long as the setup is not hackable and often we run hardware, that's very hard to hack, the model will be able to to achieve these superhuman experience where superhuman performance, if not already. Whereas for domains where human judgments remains quite important, where often human taste... It's just really hard to encapsulate in a number that we can say, okay. This is clearly better than the other. There, the bottleneck to improvement is still human as we see it today. But we have seen in practice really comes down to a degree of how clearly can we specify for a given domain what it means to be better. And can we specify that in a function that is cheap to evaluate, that is uncontestable. And then if we can do so, when we put in a lot of compute, even the existing algorithms are quite efficient at finding solutions, that satisfy, what it means to be better. But, a lot of the economy still today are not quite there yet when it comes to our ability to specify what is better.

[1:17:36] Nathan Labenz: Part six, not just verifiable tasks. To close, something lighter. Wednesday's show opened with a track that Prakash made with Claude Fable built from clips of voices you'll recognize. It plays us out at the end of the episode. He explained how we made it.

[1:17:56] Nathan Labenz: It is...

[1:17:57] Prakash Narayanan: I was using up my tokens. I had some tokens left over from Fable. I gave it to... I I gave it to Fable, and it picked maybe, like, seven or eight clips from, like, the last, like, couple of years, you know, just randomly, like, stuff that I remember, like, key moments. Right? Jensen's, like, you know, I didn't wake up a loser, etcetera, etcetera. And, basically, I probably listened to about nine versions of it. So each time I listened to it, I'd like, I don't like this. Don't like that. And finally got it to this date. I gave it access to a API API based digital audio workstation, which is what you use to kinda compose various instruments together onto a track. And I said, go ahead. Use this. You can play the instruments. Figure out which instruments to play. It can't hear, obviously. And I think Suno is real generative AI. This is not really generative AI because it's really like playing the instruments. That... And that's what that's what really surprises me. It's it's actually composing multiple instruments together. And so the machine was the one that came up with the sequencing. It chose which samples to use. It chose the specific phrases. I told it to put together a narrative. I wanted, like, this kinda storyline. So it starts off with Dario's, like, the models they would just wanna learn, and then it does this... That became the refrain, and then it it did the the intro part where he said how Ilya told him about models that they just wanna learn. And it ends with Kurzweil, which is a very sober, like... You know, know, we're progressing towards a human machine, you know, civilization. My my bit was probably, you know, Ternstow saying, we gotta slow down. And each time he says we gotta slow down, it speeds up. That was fine. That was my contribution... Net contribution.

[1:19:40] Nathan Labenz: On one level, you know, who cares? It would be easy to say, you know, this isn't moving GDP. This is just like us, you know, arguably amusing ourselves to death in in a new form. So I do think that perspective is, like, somewhat valid. But, also, I do think this goes right to the heart of one of the big questions that everybody's been talking about recently, which is simply how good are these things gonna get in domains that are not automatically or readily verifiable? It really does seem like the generalization is pretty strong. I mean, I highly doubt they're doing much in the way of feedback loop on music. This is probably one of those things that they just kinda threw some stuff in, see what comes out. Okay. It's better than last time. Great. Let's ship it. Pretty impressive and definitely strongly suggests to me that we are not going to be able to hide behind the... You know? Well, it only works for verifiable reward type tasks for really any longer. If you're thinking, if you think that was gonna be the the last wall to hold, I think, unfortunately, it's not going to. Listen to the track once more if you don't believe me.

Outro

[1:25:58] If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the cognitive revolution.


Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to The Cognitive Revolution.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.