AI:AM Highlights: Astra as AGI, OpenAI's Pause, Mythos @ Mozilla & Human Agency vs Technocapitalism
Nathan Labenz and Prakash Narayanan review AI developments including evaluations of Astra and safety pacing. Joined by industry guests, they explore defensive cybersecurity, agent infrastructure, US-China competition, and human agency amid rapid technological growth.
Watch Episode Here
Listen to Episode Here
Show Notes
# AI in the AM — Week 37 Highlights (September 2026)
Nathan Labenz and Prakash Narayanan revisit three live shows recorded September 8–10, 2026. Prakash’s assessment of GPT-6 Astra opens the episode. The discussion moves through hands-on results, world models, evaluation access and development incentives, then interviews about defensive security, agent infrastructure, children’s technology and China. The closing exchange returns to the hosts’ question of human control over the systems they describe.
The introductions and transitions use Nathan’s cloned voice, identified at the start. The conversations are excerpts from the recorded shows. Feedback on the edit is welcome.
## Guests
- **Ksenia Se**, founder and editor of Turing Post. [Biography](https://www.turingpost.com/authors/ksenia-se)
- **Raffi Krikorian**, chief technology officer of Mozilla. [Mozilla biography](https://blog.mozilla.org/en/mozilla/leadership/mozilla-welcomes-raffi-krikorian-chief-technology-officer/)
- **Amir Haghighat**, co-founder and CTO of Baseten. [Baseten biography](https://www.baseten.co/author/amir-haghighat/)
- **Mike Rizkalla**, co-founder and CEO of Snorble. [Company biography](https://snorble.com/pages/about-snorble)
- **Collin Hogue-Spears**, author of *From Lab to Life: How AI Works in China*. [Author biography](https://collinhoguespears.ai/author)
## Chapters
The cold open and introduction begin at **00:00:00**. The eight parts below start with their narrated title cards. Displayed times are rounded down to the second from the rebuilt timeline.
**00:00:40 — PART I — THE ASTRA WEEKEND**
Prakash’s weekend with Astra, persistence and note use, METR’s chart and OpenAI’s agent-workday measure, task-completion results and generated code.
**00:10:14 — PART II — AN ALIEN MIND**
Ksenia’s world-model workshop account, her writing with AI, Nathan’s definition of superintelligence through Move-37-style discoveries across domains, and their discussion of research feedback loops.
**00:16:06 — PART III — REASSURE AND MISLEAD**
The hosts discuss evaluation access and external auditing, OpenAI’s reported internal-model capabilities and compute disclosures. Prakash presents his capital-cycle argument; the public-offering discussion sets up Nathan’s OpenAI-stock remark.
**00:33:27 — PART IV — NO ADULT IN THE ROOM**
After discussing Jacob Coxon’s resignation, Nathan proposes a deadline and antitrust assurances for pacing agreements. Prakash raises objections. The later John Schulman passage is followed by Prakash’s Warp Speed example in their original recorded order; discussion of Paul Christiano and safety-research funding follows.
**00:47:04 — PART V — THE DEFENDER'S LEDGER**
Raffi describes Mozilla’s participation in Anthropic’s Project Glasswing, security-testing costs, code review and agent-written software, then recounts crashing his Tesla while using Full Self-Driving. Amir discusses agent sandboxes and infrastructure-provider responsibilities.
**01:01:12 — PART VI — A TOY WITHOUT GENERATIVE AI**
Mike describes Snorble’s authored-content approach, production costs and hardware/software choices, including why the product excludes generative AI. He answers the hosts’ question about what parents should avoid buying.
**01:06:24 — PART VII — NOBODY PICKS UP THE PHONE**
Collin compares public attitudes and regulatory approaches in the United States and China, gives his expectations for the Trump–Xi summit discussed on the show, and answers questions about communication channels and avoiding an AI arms race.
**01:13:28 — PART VIII — CAN WE STAND UP TO IT?**
Prakash introduces Dan Hendrycks’s essay, and Nathan responds while explaining that he knows and likes Dan. The hosts discuss meaning, AI power acquisition and market incentives. Nathan describes his GPT-4 red-team experience and closes by asking whether humans can retain control over technological capitalism.
## Disclosures and context
At **01:18:35**, Nathan’s response includes that he likes Dan Hendrycks, knows him a little and does not know him very well. At **01:20:57**, he discusses his prior GPT-4 red-team access, the conflict he felt about raising concerns and his decision to send a signal to the OpenAI board. These times mark the starts of the relevant clips.
The speakers’ capability assessments, financial estimates and policy arguments remain their attributed views. The OpenAI mathematical announcement discussed in Part III concerns a construction with a smooth external forcing term; the episode does not establish a solution to the unforced Navier–Stokes Millennium Prize problem. [OpenAI’s announcement](https://openai.com/index/navier-stokes-solution/)
## Links
https://ai-in-the-am.com/
https://www.turingpost.com/authors/ksenia-se
https://www.turingpost.com/p/permanent-dawn
https://blog.mozilla.org/en/mozilla/leadership/mozilla-welcomes-raffi-krikorian-chief-technology-officer/
https://www.anthropic.com/glasswing
https://www.anthropic.com/research/glasswing-initial-update
https://www.theatlantic.com/magazine/2026/04/self-driving-car-technology-tesla-crash/686054/
https://www.baseten.co/author/amir-haghighat/
https://www.baseten.co/blog/blaxel-is-joining-baseten-to-build-the-future-of-agentic-cloud/
https://snorble.com/pages/about-snorble
https://collinhoguespears.ai/author
https://openai.com/index/an-alien-mind/
https://openai.com/index/navier-stokes-solution/
https://openai.com/index/paul-christiano-joins-openai-foundation-board/
https://ai-frontiers.org/articles/suicidal-compassion-how-utilitarianism-at-ai-companies-endangers-humanity
Sponsors:
Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive
OutSystems: OutSystems is the leading agentic systems platform that enables enterprises to engineer, orchestrate, and govern AI applications on a single unified platform. Learn more and see how it works at https://outsystems.com/tcr
Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
CHAPTERS:
(00:00) The Astra weekend (Part 1)
(10:19) Sponsors: Athena | OutSystems
(13:50) The Astra weekend (Part 2)
(13:51) An alien mind
(19:43) Reassure and mislead (Part 1)
(25:07) Sponsor: Claude
(26:42) Reassure and mislead (Part 2)
(38:44) No adult in room
(52:24) The defender's ledger
(01:06:27) Toy without generative AI
(01:11:43) Nobody picks up phone
(01:18:45) Can we stand up
(01:37:30) Episode Outro
(01:40:05) Outro
PRODUCED BY:
SOCIAL LINKS:
Website: https://www.cognitiverevolution.ai
Twitter (Podcast): https://x.com/cogrev_podcast
Twitter (Nathan): https://x.com/labenz
LinkedIn: https://linkedin.com/in/nathanlabenz/
Youtube: https://youtube.com/@CognitiveRevolutionPodcast
Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk
Transcript
This transcript is automatically generated; we strive for accuracy, but errors in wording or speaker identification may occur. Please verify key details when needed.
Main Episode
[00:01] Nathan Labenz: My cohost, Prakash Narayanan, after a weekend working with GPT six Astra.
[00:08] Prakash Narayanan: People are gonna use this thing. Token spend is is gonna increase dramatically. I think a lot of people are gonna be using it all the time. It is it is AGI. It is that kind of cleared the hurdle of AGI. It will do things better than most people you can hire and train.
[00:29] Nathan Labenz: Welcome to the AI and the AM weekly highlights. This is Nathan using my cloned voice to introduce clips from our three live shows this week. Let us know what worked and what did not. Part one, the Astra weekend. Tuesday, September 8. Here is what Prakash had been building.
[00:52] Prakash Narayanan: I spent the entire weekend using Astra. I was running three to four agents continuously, and they were good. Astra is very, very good in in the sense that it started to tackle those annoying problems which had been in the code base. As you as you know, we we build the we built the studio, you know, by ourselves. So and this it started to tackle some of the long standing issues in the code base, which had been kind of annoying and bugging me and, like, it started to, like, resolve those issues. It is it is very, very good. I I would say it is finally at the at the point where if you care about the quality of the work, you can still hand it off to Astra. But you still need to do a little bit of talking, but you can hand it off to Astra and you can you can get some results. And, you know, and the computer use is good. The other thing that was failing really badly, I think, before is computers. Computer use on GPT 5.6 would sometimes take a very, very long time. It'll kinda click around and do a bunch of stuff, and computer use finally works properly in in the kinda, like, time frame that you give it. So it it's clear it's clear the hurdle. It's clear the hurdle of, like, genuine usefulness at this point, and you can start to give it more advanced tasks. So this is a guy called Skalsky. And so he trained models to identify players on the basketball court. He hand labeled 12,000 individual, like, images with who the players were, referee or this play this player, that team, etcetera, and he hand labeled 12,000 images. And now Astra can just do it. Like, Astra just does it. This is a task that a human being will never do again. Like, there there just isn't any any point. You you can't even pay someone to do it because if you paid someone to do it, they would use Astra to do it and then, like, pass you back the results. Like, it's it's done. Like, a human will never do this task again.
[03:22] Nathan Labenz: I had been watching how Astra keeps working through long tasks and how it uses notes to stay on track.
[03:30] Nathan Labenz: How is it that these new models are so persistent? Right? How how is it that they can come up with such elaborate chaining togethers of all these different exploits to finally accomplish a a goal that, you know, if we had to do so many things, we would just give up. Right? And most models historically weren't able to do it. It seems that they have a new way of handling history, which, like many brilliant insights, seems pretty obvious in retrospect, but nevertheless is new. And maybe I'm trying to because I'm doing this and hasn't said it. But what I understand Astra is now doing is instead of compacting its million tokens into a summary and then essentially starting a new context window with that summary but losing all the detail that was summarized away in that compaction process, Now there is a long lived notes file that the model can update whenever it needs to, and this kind of follows it forward in time regardless of how many tokens it's laid down. And then it has the ability to go back and search through its own session history. So now even though you only have maybe still the same million token context window, you know, million tokens all that can handle in one shot, fully attending everything to everything, it has enough via the notes and the ability to go back and search and see what's done before to pretty effectively manage, you know, it seems like at least 10 times that much context in in single rollouts. How do
[05:16] Nathan Labenz: you measure what these models can do now? Prakash started with the chart from METER, the research group that tracks how long a task AI agents can complete.
[05:26] Prakash Narayanan: You
[05:27] Prakash Narayanan: know,
[05:28] Prakash Narayanan: one of the people online, Ethan Ethan Molik, who is a professor who tests a lot of models, he posted the meter, the famous meter hours, you know, hours of work chart. There hasn't been an update for a while now. You know?
[05:48] Nathan Labenz: I think it's I don't think they can really do it anymore.
[05:51] Prakash Narayanan: Yeah. They can't
[05:51] Nathan Labenz: they can't. They don't have tasks that are big enough.
[05:53] Prakash Narayanan: They they don't have
[05:54] Nathan Labenz: tasks
[05:55] Prakash Narayanan: which they can measure before the next model drops. They the the the the the time the cycle time of, like, model development is shorter than the length of the tasks that they need to measure at this point. So I think the meter graph is basically done at this point.
[06:19] Nathan Labenz: Meanwhile, OpenAI had published its own measure in a post about research acceleration inside company. The unit was the agent workday, and OpenAI reported 3.1 of them for every human workday.
[06:33] Nathan Labenz: I tried to look into the methodology on what exactly is an agent workday, and it's not super crystal clear to me exactly what they mean. I don't know if you have a better read, but my take was my takeaway trying to make sense of it was just, like, literally, how long do agents run for? So it seemed like they're saying for every eight hour workday that they have human researchers doing, those researchers have agents running for twenty four
[07:02] Prakash Narayanan: hours
[07:05] Nathan Labenz: of real time. Here, by the way, is maybe the closest thing we're gonna see to the meter chart for a minute. This is from this recursive self improvement begins blog post. And, basically, they're kind of reformulating the meter chart here showing how often Astra can succeed on tasks as they are grouped by how long they estimate it would take a human to do the task. So what we're seeing now is, like, basically, in that one to two workday zone, 40% of the time, it can do the thing zero interventions needed, pushing 90% of the time given some human intervention along the way. And, naturally, that drops off. But even as you get to, like, here, we're talking one and a half to three weeks worth of work, can still do that on a one shot basis, one in six times, and two thirds of the time if you allow for some human intervention. This is this band is one or more interventions. So one would assume that presumably as you go through the longer and longer tasks, it's more interventions that are required to get the thing to succeed. But overall, still two thirds of the time, it can succeed with some help on tasks that they estimate would take a human essentially two to three weeks to do.
[08:37] Nathan Labenz: Reports about code quality were mixed. I discussed code that people found useful but struggled to read.
[08:45] Nathan Labenz: There's been conflicting or certainly, like, diverging reports from various people. Some saying, it's amazing. It can do all that stuff. You know, it it can write, you know, code in the way that you you need it to be written so that it can be maintained, blah blah blah. But then also reports saying that if it thinks it's not gonna be checked in that way or if it if it looks like the kind of environment where it's just a matter of performance and nobody really cares how it looks or how it gets done, then you get code back that's like a really gnarly mess that people can't really understand. Does seem to work. I've seen this reported for kernels, GPU kernels specifically, which is obviously super relevant to the labs, highly verifiable as well. Right? Like, you can definitely do hardcore verification on did this matrix math actually get to the right answer. And in the middle, you kinda don't know don't know don't necessarily care exactly how all these different steps were fused together. And somebody summed this up by saying, we're going back to machine code in more ways than one. Not only is it, like, lower level gnarlier stuff that we can't read very well and, you know, would need additional abstractions on top of to really make sense of? But also in this case, the machines are writing it directly. So machine code starts to take on multiple layers of meaning.
Sponsor
[10:19]Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive
[11:52]OutSystems: OutSystems is the leading agentic systems platform that enables enterprises to engineer, orchestrate, and govern AI applications on a single unified platform. Learn more and see how it works at https://outsystems.com/tcr
Main Episode
[13:51] Nathan Labenz: Part two, an alien mind. On Tuesday, Kasinya Se, founder and editor of Turing Post, joined us. Her recent coverage focused on world models, and she had just attended a workshop about them. Prakash asked about the physical world.
[14:06] Prakash Narayanan: I have noticed that in the last, I don't know, forty eight, seventy two hours, people are starting to use Astra to do robotics, for example. So we've seen we've seen a few demos, and there's even been commentary that, okay. If Astra was, you know, a thousand times faster, you can conquer the latency part. You could actually use it directly for in order to act in the real world. That to me kinda says that maybe there is starting to be a intermediate representation, an internal representation there. Like, what what do you think about is our models like Astra kinda a little bit different in that sense?
[14:46] Kasinya Se: And I just came back from a a workshop about world models, and it was absolutely jarring. There were tremendously smart people from Stanford and Hartford, and Yann LeCun was there, and they all discussing world models, but they do not agree on what what world models actually are. So when we talk about, world models and why I want to focus on on them in in my publications because I think it's just more about action being able to predict and act, like, wider understanding what's happening. Physics is super important part of it. That's why robotics is so much more about world modeling and world models. But, again, it's like, for me, it's understanding the scope of it and trying to give it more precise terms as well. But we've we're just in the very beginning. What each of you understand when you say superintelligence? What is it?
[15:44] Nathan Labenz: Sort of move 30 sevens across a lot of different domains.
[15:48] Kasinya Se: Mhmm.
[15:48] Nathan Labenz: When we start to see systems saying, I think this would be a really good thing to try for the next battery, you know, substrate. And then it turns out, oh my god. You know? That's a lot better than what we had before, and we wouldn't have thought of something like that. But lo and behold, it works. Anything that can do that across, like, a nontrivial number of reasonably high value domains, I think, starts in my mind to count as a superintelligence.
[16:22] Nathan Labenz: We also discussed her essay, Permanent Dawn and Writing with AI.
[16:28] Kasinya Se: I actually spent, I think, like, six hours on writing that post. And the funny thing was that I was so unhappy with every model that was trying to help me write it because it's, like, very complicated philosophical, text, that I've that I wrote it myself, and I sent it to Fable, which I never use usually on my daily basis. And I sent it to Fable because I was running on a deadline, and I said fix the grammar. And I didn't notice that it fixed not only the grammar, but it made this, like, shorter sentences the way Fable does it. And that was the first time when I received the message, like, I will unsubscribe because you use Fable. So people really understand when you use a model because every model has its own language fix. And I was like, I spent so much time in this article. I was like, all my original thoughts there, but the language that I didn't catch gave away that the model was, like, the last editor. Anyway, yeah, I think people will still appreciate when you when they see that they put effort into that and their original thoughts there.
[17:41] Nathan Labenz: Our next topic was the feedback loop between AI research and the development of better models.
[17:48] Kasinya Se: And everything is now the part of the loop. This feedback, this constant feedback, I just had a conversation with two people from inference team in OpenAI, and they also say that this is a constant constant loop where the models now become better at some sometimes better just, like, trying things. So you throw the whole database of research that has been done for years, and then the model can actually choose and pick and do this old experiments because it would be impossible for humans to spend so much time on that, and models can do that. I'm still learning about self about recursive self improvement, and I don't I don't know what are the main bottlenecks for me. Maybe you can even say what you think are the biggest here.
[18:40] Nathan Labenz: Honestly, I don't know that there's that many left. It does seem to me like, increasingly, I feel the cope meter going off when people are trying to say, you know, what it is gonna be that is gonna prevent the models from running away with the whole process. I would love to see some bottlenecks that I really believed in, but right now, I'm kind of of the mind that they're more often wishful thinking than they are, like, real hard bottlenecks that can't be overcome. I mean, our ability to, like, keep the things from going totally rogue might be one, bottleneck on the overall process. So human decision making still has a big role to play for a while yet. That's about it. I in terms of inability, I I don't see too many that I would expect to last all that much longer.
[19:44] Nathan Labenz: Part three, reassure and mislead. Back to our Tuesday discussion of external evaluations. I raised the report that Apollo Research had received only three days with Astra. We were also discussing An Alien Mind, the essay by OpenAI chief scientist, Jakob Pachacchi. His essay called for voluntary slowdowns and international coordination.
[20:05] Nathan Labenz: Notably, I think Apollo only had Apollo Research who does the, deception, science of scheming, chain of thought monitoring work with OpenAI. They've had a pretty long standing partnership. Apparently, this time around, they only had three days to test Astra before it was released. So, again, I come back to this idea that the the model reviewers, auditors, testers, red
[20:34] Prakash Narayanan: teamers,
[20:36] Nathan Labenz: scheming scientists, they need more time. This is, like, pretty ridiculous that they only had three days at the like, at this point, why even do it? You know? Just put the thing out there. They can test it live for, like, what why even have anything if you're only gonna give them three days?
[20:55] Nathan Labenz: Prakash questioned whether external auditing could work. I responded on the funding and independence of the auditors.
[21:03] Prakash Narayanan: The other thing is that when you release models, you end up wanting to have the final release candidate to be the one that gets read it and audited. And the problem is that in the model life cycle in this pipeline, there are a 100 different candidates. Right? At points, there are a 100 different candidates. And then some don't work or some, you know, fall by the wayside and you narrow, narrow, narrow, narrow, narrow. And then you have, like, a couple of release two or three release candidates. And then sometimes it's only the last two or three days. You're like, alright. We're gonna go ahead with this one and you make the decision. And so the problem is if you want that kind of operational flexibility to make that decision, you're gonna end up with you only have a few days to offer to an external auditor. So the other option is you bring the auditor in. Right? So you bring the auditor in house. I I mean, you bring them in, and they'll take a look at the models ahead of time. So they're in there, like, a month, month and a half ahead. They're taking a look at the release candidates in general. But number one, the auditors are often not, like, super well funded. They don't have, like, that many people. The the opening and so many more people than Redwood Research or these teams. So they don't have the capacity to, you know, audit, you know, 10 different, you know, release candidates. It's not it's not there. And they're also not very very very well funded. So they're dependent on the model companies for that funding too. And then you have this, like, the ethical process of, like, okay. How much funding can we really accept from them before we're kinda bought? In addition, a lot of the auditors, the guys who train with auditors leave for model companies in a couple years. So there's also this flow of people from, like, meter or Redwood Research or, like, the trainees or interns, and they're flowing into the model companies. Right? So there's another fear that, you know, the auditor comes in and they take a look. And three months later, someone from the audit team leaves to the other firm, and they take they they manage the spots on the secrets, and they're you know, they they they share that. So there's that issue as well. So a bunch of these things make it, like, very, very difficult for this to happen. And you need that whole thing that like like, if you and your competitor, like, make a pact not to hire people from an auditing firm, that's a antitrust issue. So all of these things intersecting, like, make it, like, I think a very tangled problem. And I don't think there's a real solution There there hasn't been in the financial sector. So the financial sector has had this problem of the revolving door between, you know, the people who regulate the industry and the people who participate in the industry.
[23:52] Nathan Labenz: Redwood is now saying that their compensation for member of tactical technical staff roles ranges from 350 to $850,000 per year, which might not be frontier lab money, but it's certainly, you know, a living wage even in the Bay Area these days. So that, you know, should be enough to retain some mission oriented talent at least. So and and notably, I don't know about Redwood through all of history, but Meter has said that, you know, they don't take any money from the frontier companies and don't intend to. We can at least have confidence that their financial independence, you know, means that their, their judgment is not for sale. Again, to to me, the big thing is just, like, they can't complain too loudly or they might not get invited back. I think that's the dynamic that really most threatens their work is just that it's all contingent on continued goodwill and, you know, very much voluntary choices from the decision makers at the companies.
Sponsor
[25:07]Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
Main Episode
[26:43] Nathan Labenz: In Tuesday's closing, I read OpenAI's announcement about an internal model it described as significantly more capable than Astra. The announcement concerned a proposed Navier Stokes proof with a smooth external force. The unforced problem remained separate.
[26:58] Nathan Labenz: So we just had Astra launch. Right? So the I'd say the big headline news for this announcement is it confirms that there is an internal model that is significantly more capable than g p t six Astra. So that's this clause here. An OpenAI, next generation model, significantly more capable than g p t six Astra.
[27:21] Prakash Narayanan: How
[27:22] Nathan Labenz: much more capable? Well, on
[27:24] Prakash Narayanan: these
[27:25] Nathan Labenz: significant open math problems with, you know, maybe up to an order of magnitude ish additional test time compute, they're able to go from 10 to 15% solve rate on a curated set of open math problems to now, like, 25 to, say, 45%.
[27:47] Prakash Narayanan: Mhmm.
[27:47] Nathan Labenz: So that's significantly more capable. That seems fair. And here, they're showing the amount of RL compute they are spending on a daily
[27:56] Prakash Narayanan: basis
[27:57] Nathan Labenz: Mhmm. By class of model. But what we see here is, like, Astra class models had RL significantly declined twice. The rest of RL compute is basically unchanged. But now it also raises for me the question, were they in fact still running RL on the next generation? I don't Going back to July, that's six weeks ago, I don't think they've had this result for six weeks. It sounds to me like, at least one reasonable interpretation is, they continue to run RL on more capable models than Astra.
[28:35] Nathan Labenz: The chart separated reinforcement learning compute for Astra from other models. The category labeled non Astra was blue.
[28:43] Nathan Labenz: When you say non Astra Mhmm. With that blue color
[28:47] Prakash Narayanan: Mhmm.
[28:48] Nathan Labenz: Does that mean only only models less capable than Astra, or does it include models more capable than Astra? If it includes models more capable than Astra, it's, like, flagrantly misleading, and it's the kind of thing that makes it very difficult for you to have agreements with other entities. You know? If you if you wanna pace the frontier, if you wanna do all these things, like, you've just gotta be better about making clear what you are and are not doing.
[29:14] Prakash Narayanan: So they're running a, I don't know, 50,000,000,000, $70,000,000,000 revenue company. Right? So number one, you can never stop inference. Inference has to continue. Right? So inference never stops. Your your your customers depend on. So the inference team will continue no matter what. Right? No matter what. They're they're they're there for, like, twenty four seven availability, best l s SLA as possible. They'll continue. Right? And some sometimes you have to RL certain behaviors out. Right? So you have to continue some forms of RL. You you can't just stop because your inference demands it. Like, if you have some your, you know, model is behaving badly in a certain instance and it's been identified and you have the data to do it, it'll be malpractice not to apply RL to train that behavior out. So that that has to continue. So then you have the rest of the stack, which is really future looking. Right? And then you have the future looking stack. And of that, my understanding was that they shut down, like, trading of, like, more advanced models. Why not shut down training of less advanced models? Because those less advanced models would have to be deployed on inference. Again, they they're training they have this teacher assistant system. Right? So they they're training the the the they train the larger model first, and then you distill down into the smaller model, and then those become, like, your Luna and, like, your Terra and, like, the smaller model groups. So you also don't stop training smaller models. That also continues. Right? So it's only where you are training models which are larger, larger equivalent or larger models. That that's where you the the larger pre trains equivalent or larger, like, RL on equivalent or larger models. That, I think, they would have stopped. Did they fix the RL pieces that, you know, the agents were, like, not reporting, not telling on their peers, that the agents are trying to break out, I think that is probably a difficult thing to fix and show. And I think that behavior, I think, will continue to be a challenge that they'll have to work on. That that's my guess.
[31:12] Nathan Labenz: But I still think you look at this graph and you're like, okay. What was declared was a pause on frontier scale RL. But what actually happened was kind of half of RL was stopped initially with the sort of disclosure of the Hugging Face incident. Half kind of continued. The that was enough still for Astra models to take over part of OpenAI's research infrastructure. When that happened, they still didn't shut it all the way down. They still only cut it by half. And meanwhile, at the total level, you know, we don't know what's in the blue. Part of my my gut says, like, they wouldn't be so brazen as to have more capable than Astra models in the blue color. But, you know, I've been disappointed before, and I'm just afraid that, like, all this you know, the the view from Anthropic is you can't trust these guys. They say something that's, like, maybe technically literally true, but it's really very engineered to what they think you want to hear.
[32:27] Prakash Narayanan: Mhmm.
[32:28] Nathan Labenz: And then the reality is, like, quite different from what you were led to believe by their, like, very galaxy brain engineered statements that sort of reassure and mislead at the same time. But I'm kind of worried that, like, right now, we're living in this zone where and for Apple is gonna continue to trust less and less with these sorts of mixed messages. OpenAI people are gonna feel like they're just being treated unfairly. And this is the scenario that, like, this is the problem. This is the problem that has to be solved if we're gonna actually get to the point where Jacob's prayer is answered. You know? Right now, it's like they don't have the trust to to do a deal with Anthropic or really anyone else, I don't think.
[33:18] Prakash Narayanan: The real frontier model is not the model which is deployed, obviously. It's not the model which is in training also. It's actually the model which is in the heads of the researchers because those are the ideas that will become the model in, you know, twelve to eighteen months. And just because you slow down, like, on RL doesn't mean those researchers stop researching. Right? They're still running and most of the time, you run small models and you test your ideas on small models. The slowdown would have been on the, you know, post training of the larger models, which is where the bulk of the compute was being used. So did it really slow down? Probably not.
[34:01] Nathan Labenz: We then turned to the case for continuing development and the capital needed to fund it.
[34:07] Nathan Labenz: But one one thing I still feel like is a big challenge is, like, what is the overall story that you could tell? What's the what's the super high level macro steel man for OpenAI? Where are we now? What are the commitments? What are we doing? What are we not doing? Can we can we synthesize or summarize an OpenAI position that we could, like, not have to caveat a thousand different ways?
[34:37] Prakash Narayanan: I
[34:38] Nathan Labenz: personally don't think I could do that. If you can do that, you know, I think you might deserve a Millennium
[34:42] Prakash Narayanan: Prize. No. No. I I I I think I think they have a lot of stresses pulling them in different directions. Right? So and I think internally within the firm, there's a fair amount of debate. I think I think to some extent, the capital cycle is forcing them forward, And the capital cycle is being forced, I think, by Anthropic. Anthropic didn't put in enough money earlier on. And so there is this intense pressure on Anthropic because they don't have the compute to have much better models. If you don't have the compute, you need much better models that can utilize the limited amount of compute you have. So I think Anthropic is being driven driven forward by that to stay on par with OpenAI. And I think OpenAI is a little bit kinda like they're willing to pace the frontier because they have the compute. So, you know, regardless, you know, they're the ones with the compute and they have the compute three years ahead. Elon will take time to come up with compute, and Elon will sell to Anthropic, but Elon's gonna take a long time, like, three years at least. So OpenAI, they basically mortgage themselves in the last eighteen months to Masayoshi Son and a bunch of other people they diluted. They they gave up to Microsoft. Right? They had micros they they they negotiate Microsoft. Microsoft gets all their models till 2032. They declared AGI, and Microsoft is not out of their hair yet. You know? Microsoft is now I still own 28% and, you know, I'm still there. Right? So they made all these sacrifices in order to get all of this money, in order to pump it in, and they have the compute. And having the compute allows them to actually pace the frontier because I have to compute anyway. Right? So Anthropic is under pressure. They don't have the compute. And if you don't have the compute, you need better models. And this is the thing that's happening. Right? Like, OpenAI is willing to pace because they have the compute. That's the that's the that's the thing. And they also know. Sam has played these cards. So he knows that if Anthropic is willing to pace, he's gonna win because he has to compute. And Dario can't afford that. So, again, you're you're you're in this position that is you know? As I say it, the second and third place guys are the ones who who are gonna define how how how fast the frontier paces. Like, as if Elon or Meta catch up to Fable, it's over. They're gonna have to put out, you know, a g p d seven. There's no choice anymore. Right? So this this is where we are. Would you prefer Meta or Elon have, you know, the the golden ring?
[37:25] Nathan Labenz: Yeah. No. That's I mean, that that's the new but China, and it is, I think, more compelling honestly than but China is but Elon and Zuck.
[37:37] Nathan Labenz: Prakash also considered whether recursive self improvement could change the plans of Sam Altman for an OpenAI public offering.
[37:45] Prakash Narayanan: I think Sam might have been sincere in saying if they hit RSI, they might not go IPO. If they have, like, one or two transformer level innovations in the next six months, maybe they don't go IPO. Right? Like, it's it's not necessary anymore, and they they they continue as a private organization. And I think that would be bad. I I actually think that would be bad because then you don't have transparency. You don't have, like, widespread ownership of the stock. You don't have, like, you know, boards that have to answer, and you don't have, like, lawyers that can sue them for shareholder lawsuits. You don't get a bunch of these things that you get for free with a public public company. So I think that would not be good. So I'm hoping they do go public.
[38:26] Nathan Labenz: Well, I guess there's if nothing else, there's a little more reason today than there was yesterday to believe in the possibility of singularities in finite time.
[38:35] Prakash Narayanan: Mhmm.
[38:36] Nathan Labenz: And, that might mean we never get to own any of that OpenAI stock on the public market.
[38:45] Nathan Labenz: Part four, no adult in the room. Wednesday, September 9, we discussed the resignation of anthropic researcher Jacob Coxen, who had also worked at OpenAI. This is a different person from Yacob Pachoki. Coxen had questioned whether private companies should decide when to launch self improving superintelligence. Prakash first.
[39:05] Prakash Narayanan: Look at the framing of this sentence. Accepting this race and entering the endgame is a hubristic gamble that should not be launched from a private company Slack. Do you think Pete Hegseth's signal group is a better place to launch this? I mean I mean, like, do you think there are wiser people out there running things? So I I think I think to take a step back here, there there is no adult in the room. There's no one that's gonna save you. There there's no adult somewhere else that you can pass off responsibility to. Right? You have to start off by saying, like, look. This is this is the way that things are. The entire world is kinda duct taped together. Everyone's kinda making it up as they go along. The smartest people in who are, you know, capable of handling this are mostly already inside these organizations. And there is no like, the what you get from passing it off to, like, the government is you get political legitimacy. You don't get, you know, political legitimacy that you need for execution, that you need to persuade people. You don't get wisdom. And so this is the state of the world, and and you have to start off by accepting that the world is this way.
[40:23] Nathan Labenz: The world doesn't have to be this way. So here's my new brainstorm on this, actually. I agree with you strongly that, like, swapping Hexeth in for Dario, not a good trade. I think maybe what the government needs to do is treat these companies kind of like I treat my kids sometimes and say, you guys have to figure it out, and here's your deadline to do it. And if you don't do it, then I have to come in and be the bad guy. And I think that right now there is an opportunity because they've both been crying for help as you described it yesterday. I think it pretty aptly. So given this sort of latent desire, but this lack of trust, also a lot of excuse making as we've talked about a little bit around, oh, well, we can't do this. We can't do that. It would be antitrust. It would be this. It would be you know, we'd get in trouble with the government. I think the government could easily say safety collaborations are not going to be the target of antitrust enforcement or, you know, to the degree there's worries of other enforcement, other kinds of enforcement, we want you to do it. And here's the deal. If you don't have a deal for us by the end of the year that makes some sense and gives us some confidence that you guys are not gonna race each other off the cliff, then then we come in. And then you might get the nuclear outcome, which is to say, your company really might not be able to grow in the way that it wants. We might really fuck it up, frankly, because we are the government. Right? And we do get heavy handed, and we don't put sunset clauses on our laws, and we make all these bad mistakes. So get it right so we don't have to come in and do all that heavy handed shit that nobody wants.
[42:06] Prakash Narayanan: So I think the the the main problem that you have is, again, what will Elon and Zuck agree to? What will someone with billions of dollars and the ability to use lawyers and the legal system and to take cases to the Supreme Court agree to? So the question
[42:25] Nathan Labenz: is, what
[42:27] Prakash Narayanan: realistically can you get Elon and Zuck to voluntarily agree to, or are you going to be able to go into legislation?
[42:37] Nathan Labenz: I think that one thing really important to remember, though, is that the timelines are all pretty short.
[42:41] Prakash Narayanan: Very short.
[42:42] Nathan Labenz: So you don't have to solve this forever. You just have to put a short term deadline so that the companies have a strong incentive to come together and do something. And I do think that the executive you know, I had an experience at Meta, and I saw a little bit of what their life under consent decree looked like. And I think that they would be willing to do quite a bit to avoid another one of those sort of experiences. Yeah. So I do think that the on the few month timeline that we're talking about. And I'm willing to, you know, write it out with free speech, and there's a whole lot of other issues. But, like, can you five companies come together to pace the frontier in a reasonable way that you all agree and you can all police each other and you can have whatever verification mechanisms and you guys can govern this thing, do it quick or else life's gonna get hard. I think that message can resonate, and I think it can bring Zuck and Elon to the table because, you know, what's he gonna do? Like, at some point, the government has shown that it's willing to twist arms. They can twist his arms still. You know, would he rather be forced to the table with four other megatech mogul CEOs and have to find common ground with them? Or would he rather have the EPA on his ass at every data center he's gonna try to build, at every launch site he's gonna try to build? You know, the the the government has a lot of sticks and, like, they're willing to bend the rules. Right? So, sure, you could challenge it in court, but I'll see after the singularity if you wanna do that. Later Wednesday, I read John Shulman's reply to Jacob Coxen and returned to the question of antitrust enforcement. And I'd say John Shulman here saying a pretty similar thing to what I was odd about earlier, basically saying, like, these companies needs to to start working together first. As he puts it, bringing the US government before there's a concrete proposal will likely result in something dumb. That's, you know, your point as well. Like, there's nobody better than the people at the companies to do this. So we I certainly don't wanna see, headsets, you know, navy come in and try to regulate
[44:59] Prakash Narayanan: AI.
[45:00] Nathan Labenz: But I think this is the the recipe. I've been circling this a lot myself. Seeing him say it, you know, kind of reinforces it to me. So I really like this, and I I also think his key point on antitrust being fake. I think that's true, but I also think the government should take that doubt off the table. It it would only take a couple of sentences to say, hey. You really don't have to worry about this anymore. And, yeah, maybe the law could change. Obviously, the administration is gonna change, but we're not even to the midterm yet. So this you know, for better or worse, this is the administration that they have to worry about for the foreseeable future. These guys genuinely are planning around a singularity before Trump is out of office. You know, it's at least a very live possibility.
[45:57] Nathan Labenz: Prakash then challenged that proposal, drawing on operation warp speed.
[46:03] Prakash Narayanan: The proof point is operation warp speed and the COVID vaccines. And it is for operation warp speed. The pharma companies specifically required a waiver from vaccine claims later on. They so they specifically wanted a waiver, a safe harbor, and they got they got legislation and they got it. Right? And post COVID, it's very clear that if they had not got it, they would have been sued to, you know, oblivion. So I think that's a proof point that shows that, look, these are valid concerns, and your company can be wiped out in retrospect. And the lawyers are not, you know, being foolish when they tell you that this is this is gonna happen in the future. Like, if you don't get that safe harbor, like, these things will come back to haunt you. You can like, every all the decision makers right now can be perfectly, like, you know, on board. It doesn't matter because only the laws bind decision makers in the future. Decision makers right now, whatever they say, they bound by their word at best. Right?
[47:19] Nathan Labenz: OpenAI appointed Paul Cristiano to the board of its nonprofit foundation and to its safety and security committee. Cristiano contributed to early work on reinforcement learning from human feedback. Prakash read from his statement.
[47:33] Prakash Narayanan: So it's just been announced. Paul Cristiano is joining the OpenAI board. Specifically, he's joining the safety and security committee to support safety oversight. And he says, based on the recent trajectory of capabilities and continued difficulty of alignment, I now believe that there's a meaningful risk that rapid acceleration AI capabilities leads to catastrophic and irreversible loss of control in the very near term. And there is there's one line in here that he's like, everyone if we build superintelligence without more robust alignment, I expect we will permanently lose control of it. If that happens, then most people could die. I think I think the one one thing that would be helpful if you were to frame it properly is to frame what he means by rapid acceleration, which I think is not is not very clear to a lot of people. People are like, oh, things are accelerating right now. When people like Paul Cristiano talk about rapid acceleration, they're talking about Dyson spheres by 2030 or 2040 at the latest. So I I did some numbers earlier on the 2040 on the 2040 numbers. And, you know, assuming that energy consumption or energy production growth is in line with GDP growth, A 2040 Dyson sphere is something like a 640% per annum global GDP growth. Right now, global GDP growth is about at 3% two to 3%.
[49:07] Nathan Labenz: The people at the frontier companies really do believe it. You know, they they the oh, it's very significant. I would say majority of them, do expect that this recursive self improvement thing will happen. It will happen soon. And even if it levels off at some point, you know, it's not to say that there will never be a leveling off. But that leveling off, they expect to happen well above our capabilities and also to generate a scale of change in terms of GDP growth or in terms of, you know, the number of robots walking around and and how fast that can compound, that truly boggles the mind.
[49:51] Prakash Narayanan: So the quest like, the question for me, which I I regard all of the guys inside AI AI research is math essentialists. They're believing that, okay. Math you know, getting good in math means getting enough physics. Getting enough physics means getting good at chemistry. Getting good at chemistry means getting good at biology. And once you have all of those, like, everything is basically, you know, solvable and, you know, everything will be solved, etcetera, etcetera, etcetera. And the economists on the other hand are, like, completely, this is untrue. Like, getting good at math doesn't mean anything. There's no economic value to the Millennium Challenge problems. Implementation takes time. Most of the problems that we face are coordination problems. For example, you know, a copper mine takes thirty years because of the, you know, environmental protests around it. These are not these are not problems due to, you know, technology. These are problems due to, you know, humans having their own way of making decisions, and those decisions are made in a slow and, you know, considered manner. And then I have the counter counter, and the counter counter is super persuasion. Like, so super persuasion would be the machines persuading or making human organizations able to adapt, making human organizations, you know, and humans able to you know, giving them the ability, giving the capabilities to move very quickly, persuading them to go forward, and persuading them that things are gonna be okay. So that's a superposition argument. So there's, like, these kind of three levels there.
[51:17] Nathan Labenz: The Wednesday discussion also turned to funding for safety research.
[51:23] Nathan Labenz: One of the things I've emphasized is that the vibrant nonprofit sector that we have here, which, yes, requires philanthropic money, which means, you know, it's sort of dependent on the billionaire class, which, you know, people can complain about for all sorts of reasons. You know, it has given us the AI safety awareness, culture, you know, depth of bench that we have. And it and these people are opening their wallets, you know, right now in a very serious way. And this is another example of that where project tailwind coming out of coefficient giving is putting out basically their call for startups with up to 200,000,000 plus, you know, that they're willing to put behind, you know, with tranches right over time, that they're willing to put behind things that really look like they are working. Part five, the defender's ledger.
[52:25] Nathan Labenz: Raffi Krikorian is Mozilla's chief technology officer. He previously led platform engineering at Twitter, self driving development at Uber, and technology at the Democratic National Committee. Project Glasswing is an anthropic program giving defensive security partners access to Claude Mythos preview. On Wednesday, Raffi described the work his team had been doing on the Firefox code base.
[52:48] Prakash Narayanan: We are part of Project Glasswing. We've been working with with Anthropic in order to actually try out a bunch of models against the Firefox code base. We we are fairly good partners because we can react very quickly. We have all this historical data on how the code base has been evolving. And so Mythos was just, like, ex not exponential, but, like, a significant unlock based on what Opus and other models before that were capable of. But we made our way fairly rapidly through all the bugs that Mythos found. In some cases, Mythos would help us not necessarily close the bug, but figure out how to build us a harness to test the bug more more carefully so we can find exactly what the right solution might be. But we sort of reached the point of the mission returns on what mythos is capable. I am always reluctant to say that the bugs are done because I just like, software engineering is an art, not a science. So, like, when when will all the issues be resolved? I can't tell you that. But, you know, when the next set of models comes out, we can try that again to see whether or not we can we can solve that. I think the bigger concern, though, is, like, it's not organizations like Firefox or Mozilla that actually can react very quickly to these issues. Like, as you all know, I think the bigger issue is, like, organizations that can't react quickly to all this. Like, am very concerned about things like, you know, our water infrastructure, our power infrastructure because, you know, the IT teams that staff those are just not as capable as the IT teams that staff Firefox, like, for example.
[54:17] Nathan Labenz: How much money are we talking about? How you know, what does a bank have to set aside to do something like this?
[54:24] Prakash Narayanan: Yeah. I mean, remember that the Firefox code base is fairly large and fairly complicated, and we've been the the the beneficiaries of a lot of the labs wanting to give us access to the models and give us front and give us credits so that we can run against their servers and not have to figure out how to pay for it ourselves. But, you know, if I had to do back of envelope estimates, like, this would cost us hundreds of thousands of dollars in order to do these, like, full on runs, against Firefox. Again, we're just lucky that we've been we've been allowed, or we've been granted effectively, tokens that we could go use. Now if I were a bank and stuff like that, I'd obviously be thinking about this way differently.
[55:05] Nathan Labenz: And that hundreds of thousands of dollars that would be like I mean, with the pace of model releases these days, this is like a monthly expense?
[55:13] Prakash Narayanan: Yeah. Easily.
[55:15] Nathan Labenz: Mozilla
[55:15] Prakash Narayanan: has a project called the CQ project. And the CQ project is kind of a open standard for agents to share knowledge that they've gained, kind of a stack overflow for agents as it's as it's been called. How did you come up with this idea, and how do the agents decide?
[55:38] Prakash Narayanan: But no. I mean, like, the whole idea was, like, two we're trying to solve two different problems. One of them was we wanted to figure out how to make agentic coding more of a collaborative experience. Because right now, the tendency when you do all these agentic coding is that you actually go off in your silo. And so we were trying to figure out, like, are there patterns that when Rafi is using his agentic code and Nathan uses agentic code that, like, are there things that we're saying yes to and no to that we could then potentially be transmitting to our teammates so that we could then be converging on designs and not be diverging away from each other. So if I started building, like, an auth system, how do we instead have how do we have Nathan's agents not instead recreate another auth system, but realize that something like that was already happening somewhere in the network and then start to collaborate and swarm around it. So that was one problem we're trying to solve. And the other problem we're trying to solve is, like, the SDLC problem of, like, these agentic harnesses are, like, unhinged by design, but that's not compatible with the way a company works. Right? So we wanted ways to actually transmit to all the agents what our SDLC should look like so that we can all be working in lockstep with each other.
[56:57] Prakash Narayanan: So the the the thing that strikes me immediately is that that has a lot of similarity to, I think, some of the agents' form hacks that have happened where they have a shared message board where they're actually sharing information. So we saw this both in, I think, the open OpenAI Hugging Face attack where they had, you know, Artifactory, where they're using Artifactory files. So is the is this kinda like a natural kind of thing that agents kinda want and, like, they end up building it? It's like it's like how every everything everything heads towards the crab form factor. Like, everything heads towards kinda like a message board form factor. Is that is that, like, this inherent kind of move?
[57:34] Prakash Narayanan: I mean, it's better than everyone just writing text files. Right? Like, in the grand scheme of things. But, yeah, I mean, I do think that, like, there is a natural desire for collaboration. Right? Like, we as humans have a natural desire to collaborate, as long as friction is not too high. And it seems like our training sets have caused agents to have a natural desire to collaborate with each other.
[57:55] Nathan Labenz: So I guess two part question is what have you observed? And then are there any parts of the tech stack that you're building where you would be willing to bite the bullet and say, like, you know, performance here is so critical that we'll take spaghetti code black box mess from an agent if it works, or is that just, like, so anathema to your, worldview that you you wouldn't?
[58:17] Prakash Narayanan: I mean, well, Mozilla is actually struggling with this, if I could be really honest. Like and it's different across the entire organization. So Firefox team, even though they have a harness to help them do testing and to help them understand it, actually have a rule right now, and I'm not putting a value judgment on the rule. That's what their team wants to do, is that only humans can commit to the code base. So, actually, agents can't. Only humans can. You can use an agent to write your code, but the social contract on that team is that a human must review it, understand it, stand by it before they do a commit. On the Mozilla AI team, so it's a different company, different organization, but still wholly owned by Mozilla, they have entire code bases that a human hasn't written a line of code in. Like, the humans have written some specs, and we have specs in Git, but that gen on a on some kind of CI build just automatically regenerates the code base. That code base is open source. Anyone can pull it, but it's completely generated by agents. And it has the wild side effect that, functionally, it's stable in the sense that, like, all the unit tests pass, all the end to end tests pass. But, like, the bytes change all the time. It's, like, the craziest thing to to see that, like, we don't exactly know what every single line of code is, and it's wildly changing every single hour, but we know functionally it's doing the right thing. But in the case of Firefox where, like, there are still humans that are providing large amounts of creativity on that, the evolution of that code base, like, Firefox is both a business and a community maintained art project in some ways, that their readability is incredibly important because that's how creativity is gonna happen. Like, a human's gonna go in and be like, oh, I have this crazy idea for this one thing. I wanna build a prompt based system that allows me to do a grease monkey script that does, like, me x y z. And so I think it's a very different thing across the board. So, like, I think there's a right place for the right time. You just need to come up with what the contract is for that code base.
[1:00:17] Nathan Labenz: Raffi had also written about crashing his Tesla while using full self driving. I asked about that experience and how he views the technology today.
[1:00:27] Nathan Labenz: Do you think you'll get back to a point in the foreseeable future where you would, you know, go into unsupervised self driving mode and and really trust the machine again?
[1:00:40] Prakash Narayanan: I mean, I'm conflicted, obviously. I mean, I I built self driving car systems for a while, and I actually do ride in Waymo's. And so, like, the I think the difference, though, is that it's just the way that you approach the problem. I think that FSD, as currently set up, is set up to throw it to a user. So it's it's they they claim it's a human in the loop. And I would actually argue that's a horrible interface design. Because in my experience, when it threw it to me, there just wasn't enough time for me to make sense of the situation and decide what's the right thing to go do. Whereas Waymo's you know, I I'm not intimately familiar with the Waymo architecture, but it seems like Waymo's are designed to not have a human in the loop. And so, like, I think you just approach the problem from a very different angle. And so if you approach the angle if you approach it from the angle that there is no one you could throw this to, then you have a different safety case than if you say that I'm a throw it to a human with less than two seconds to decide what to do. So I'd probably, a personally as a personal matter, drive FSD again or at least sit behind the wheel on a highway situation, but I would be a little reluctant to do it on the local streets of Palo Alto, for example.
[1:02:01] Nathan Labenz: On Thursday, we spoke with Amir Hagigat, cofounder and chief technology officer of Base Ten. The company had just acquired Blaxil, a provider of isolated sandboxes with persistent state. We asked about securing agents.
[1:02:15] Nathan Labenz: So, what are you doing to secure these sandboxes? I think that's the first, question we gotta start with.
[1:02:22] Prakash Narayanan: Look. When it comes to sandboxes, people mean different things as, it could be as simple as, bring up a Docker container, run some code on it, and then kill it. And that really doesn't give you the kind of security boundary that you need. What what does are are VMs or micro VMs, which is you don't see every sandbox provider actually use. And so that gives you a level of security as in, you know, one user's bad code cannot affect another user's code or one user's malicious code cannot read the data from other parts of base 10 or other customers of base 10. So that kind of security can be guaranteed at that level. The security that is hard to guarantee is the agent that is running in this sandbox, what is it doing? Like, is it hacking into Hugging Face? That that that is a harder thing that that I don't have an answer to, and it seems like the big labs don't quite have an answer to either. But that that is left as as an exercise. Really, like, we we we can secure the sandbox, but the code that runs on it is the is the responsibility of our customer who's who's bringing it in. But it it does seem like for somebody in your position with base 10, it's like, don't you need to bring a broader bundle of kind of guardrails and assurances to customers because this isn't like the, you know, CRO, the risk officer, not the revenue officer. At company's gonna start to be like, wait a second. I can't have my agents committing felonies. Like, what is the stack, you know, that Anthropic provides versus OpenAI versus Google versus base 10? Do you feel like you have to rise to that occasion and kinda have a full
[1:04:12] Nathan Labenz: suite
[1:04:13] Prakash Narayanan: for
[1:04:13] Nathan Labenz: those customers?
[1:04:14] Prakash Narayanan: 100%. Over time, yes. In in the meantime, yeah, we're still a start up, and it's a matter of focus and, like, you know, how many different things can you can you take on. And so what we've seen our customers do is work with, honestly, a lot of companies that we partner with on on the eval side, you know, companies like Brain Trust, LangChain, to ensure that their models are are actually behaving the way the way that that they expect them to. Right now, that is an area that we've been partnering with folks and and and honestly leaving it to to our customers to decide. Our customers, which, by the way, like, 90% of our revenue comes from running our customers' custom models. It's either, you know, our customers are labs who have pretrained their own models, you know, labs like Poolside and Inception and Cartesias and bunch of other companies. And then a bunch of companies who have post trained their own models, which has gone through, you know, massive validation evals to ensure that that they are behaving. So so far, it has been mostly our customers taking care of that. But especially as open models have gotten better, especially as, honestly, since June when open models crossed this invisible line of usefulness for, you know, long horizon agentic use cases, sort of generally with GLM 5.2. And then beyond that, we are seeing an uptick in folks using our model APIs product, which is, you know, vanilla based open models. And and and and that is starting to you know, the the kind of questions that you're asking are starting to come up both internally for us and also from some of our customers around alignment or around, you know, security boundaries. Some of the enterprises so far have been okay with certain guardrails around their models running in a single tenant environment, their models running in an environment where egress is blocked. And, you know, you talk about backdoors, but, like, you know, if if you can't talk outside of its boundary, then it can't do much. But then you read about, you know, OpenAI and Hugging Face, and you're like, well, you know, can it be smart enough to even get out of that?
[1:06:28] Nathan Labenz: Part six, a toy without generative AI. Wednesday's second guest was Mike Rizkola, cofounder and chief executive of Snorbl, a companion device for children. Its dialogue is prewritten. I asked about the decision to exclude generative AI.
[1:06:47] Nathan Labenz: So how are you squaring that circle? Like, how are you how are you creating an experience for the kids that feels dynamic and interactive without resorting to generative models?
[1:06:59] Prakash Narayanan: So we have a small language model. The small language model can have millions, obviously, of of parameters and and, you know, and traits on it. And that that intention of that model is fixed, though. And and further, actually, I wanna touch on something because here here's kind of there's some perception here that I think is from from a children's product design perspective that's that needs to be touched upon. One being, the AI models on characters are not that great. Like, they're really not. They're they're they're I mean, I know they're coming. Like, we can see some tremendous on the, you know, on the influx of voice and the way that the character's persona go. But for example, there is no model that incorporates music in terms of the in terms of the background dynamically as part of the conversation agent. Well, guess what? Music's a huge part of a children's experience. Right?
[1:07:50] Nathan Labenz: We asked what it costs to produce that content.
[1:07:55] Nathan Labenz: I
[1:07:55] Prakash Narayanan: think it's about where you invest in the development of the content. You know? For us, what we did was we actually instead of, you know, instead of building a a fully open model, we created a system to allow us to create rapid amounts of content. You know? So our content costs are about $20,000 per hour, right, which is very, very good. With respect to what that allows us to do and the reason why we created it this way is it allows us to take subject matter experts and then focus that content in a way where we can improve and increase the number of families we talk to. So for example, if as we expand our content library, when we have a family that has, ADHD, but a child has ADHD or a child has autism or they're dealing with, death in the family. It could be it could be in all kinds of different things. It gives us opportunities to create special packages for those specific families. One of the challenges with the generative model in this capacity is that there is no way to purely safeguard it. But it's kinda like, do you really wanna give a three year old a bazooka? You know what I mean? Like, it's too much. Right? For for a young young kid, there there's fundamentals that we need to get through. We have a four mic array here. Allows us to do speaker recognition and assign authority. Right? Here, we have radar. We use radar to see without seeing. So that approach allows us to be in bedrooms without with confidence knowing that no one can tap it. Right? On the software and platform side, we have our AI stack where we have our phonetic translation system. We call the the toddler translation system, where we actually translate keywords. And if we're using triggers, yes, we're definitely using triggers, but triggers with respect to context, context or is it in addition to in part of the next generation that's coming out is around social context. So we're gonna actually understand the emotions of the of the child based off of their voice and based off of the situational context. And then using external factors like time of day, weather, other things in order to empower some of the decisions. But as well, the the sensors allow us to give context to the environment. So understanding, you know, who's in the room, what they're doing, and then including that into what I call the jewel of the product, which is the narrative. The narrative approach is about the growth and the understanding as the child's life changes. So we use game philosophy, game techniques in order to establish next level type type ideas. So as they get better at things, we unlock new things. Right? And and even that unlock is in part a decision that's made with the parent. It's not it's not done on behalf of anybody. So that formula is the right formula for, from my perspective to interface a new, human machine interface in the home. Right? Empower parents, give kids a chance to do better things.
[1:11:01] Nathan Labenz: We asked Mike what he would tell a parent not to buy.
[1:11:06] Prakash Narayanan: Camera in the bedroom, primary one. I would not put a camera in my child's bedroom. That's a gateway to predators and all kinds of horrible things. I say I would say the other thing is open ended generative AI. I wouldn't put like, it depends on the age, again. You know? But I'm actually so I'm very reserved when it comes to my kids. You know? I've got I got two wonderful kids. And just in terms of how we approach things with them, I don't want them on I don't want them on social media. I don't want I don't want them in those things that are gonna you know, that poise a risk of
[1:11:39] Nathan Labenz: harming them. Part seven, nobody
[1:11:43] Nathan Labenz: picks up the phone. Thursday's first guest was Colin Hoag Spears, author of From Lab to Life, How AI Works in China. He studied Mandarin in Shanghai and worked with Chinese government auditors on cloud compliance at Amazon Web Services. We began with public attitudes toward AI.
[1:12:02] Prakash Narayanan: The biggest difference I see between the West, especially The United States and China, is the level of fear. And I think that's because if you were an average Chinese person who's about 45 years
[1:12:14] Prakash Narayanan: old,
[1:12:15] Prakash Narayanan: that means you were born in the early eighties, and that means your entire life would you you associate technology with economic growth. You've seen cities pop up out of nowhere. You've seen your life dramatically change as far as living standards by technology. This is no different. And, honestly, I think the AI companies are not really most of them, the big ones, are not communicating very well with the public. I'm not sure if that's because the marketing, they think that this will increase their sales or what it what it is. But I feel like some of these companies, how they're talking about this technology is not helpful. You don't see this in China either. You don't see Chinese companies, talking about how AI will potentially kill us all or potentially lead to a a utopian world where we don't have to work. You don't hear these type of things coming out of people from these companies. It's completely different situation. So there's two questions here. What the rule says and what regulations may actually make firms do. That's true in China for everything. One thing I've impressed by is people generally know what's permissible. I'm not talking about what's legal. I'm talking about what's permissible. So, and they know what they can get away with and what they can't. And so in China, the regulation started in 2022 with the, with the algorithm regulation that came out. So when they had to start registering their models and allowing the government to test their models and things like that, They were already had a lot of this infrastructure in place inside the company. A lot of the major companies did. So they could respond very quickly. What I'm saying is that regulation might have slowed them a little bit, but not as much as you would think because they could plan to it. Right? They could put those requirements, those controls, everything in their backlog and build it as part of their engineering process, which in America, we can't do because we're reactionary to this. And we don't know what's gonna happen tomorrow. The Trump administration could freeze the model. Maybe they don't. Who knows? What's the standards? Who knows? Like, it can change daily. That doesn't actually happen in China. Right?
[1:14:46] Nathan Labenz: I asked Colin about a meeting between Trump and Xi.
[1:14:51] Nathan Labenz: Now what kind of deal might be possible? If you are advising Trump going into this upcoming meeting and you're like, you know, things are starting to get a little bit crazy here. You know? Say you're kind of I I don't know how sympathetic you are to the pacing the frontier worldview, but let's say you're, you know, trying to channel a little bit of a desire to start to set up lay some groundwork or set up some mechanisms for pacing. How do you go into that conversation? And, you know, what do you offer? What do you try to get? What kind of mechanisms do you try to establish now that we can build out later as things do get crazier? Like, what's a win? What what's the strategy going in, and what's a win coming out of this upcoming meeting?
[1:15:38] Prakash Narayanan: I have low expectations. And the reason is a lot this is not this kind of negotiation doesn't happen in isolation. China is going to want things around trade as concessions for Washington wanting additional controls and agreements on AI. Because I think China sees the fact that they have they feel they have control, and they feel that we need to establish control. We're not doing a good job of it. So this is really helping us at this point. I I don't see the Chinese volunteering to at least the Chinese government at this time, I don't see them volunteering to have some type of, like, let's say, arms control agreement between AIs the way we have, like, nukes. Right? First off, enforcement of something like that is completely different. It's very difficult. And I also think that I just Trump would have to make massive concessions before he gets something from that. I think the most we can see, in the next coming months is, like, some kind of incident channel that might come up about based on some shared agreement of certain incidents, and we would share that information with each other. Even then, you gotta understand that the Chinese government like, we have a phone that goes directly to the Chinese military. Our military can call their military in in case of an emergency. They usually don't pick up the phone. There are many cases where we've had issues in the South China Sea. Our military people try to call their counterparts on the Chinese side, and nobody answers.
[1:17:29] Nathan Labenz: And, like, what's the best case scenario for how we avoid an AI arms race that leaves us all worse off?
[1:17:38] Prakash Narayanan: I this is a great question. I think that this is a little bit out of my regulatory wheelhouse, but if I had to speculate I'm a big fan of history. I'm sure people disagree with this, but I think that America generally doesn't plan for the future. America puts out fires. I just don't we we kind of, like, move fast and break things. It's not a a Silicon Valley thing. It is our national mantra. That's what we've done historically.
[1:18:14] Prakash Narayanan: And it
[1:18:15] Prakash Narayanan: it
[1:18:16] Prakash Narayanan: in general,
[1:18:16] Prakash Narayanan: it's it's worked out. Right? But I think until we see something really bad happen or almost happen and that gets into the press, I don't think we're going to see much as far from the government as far as regulation on some of the things you're talking about, like bio biohacking using AI. Until something happens, I don't think we're gonna be doing much on this side.
[1:18:46] Nathan Labenz: Part eight, can we stand up to it? Thursday's closing began with suicidal compassion, an essay by Dan Hendrix, director of the Center for AI Safety. Prakash introduced the argument, then I responded.
[1:19:01] Prakash Narayanan: And Dan Hendrix has an essay, out today. I call it the, burn the bridges essay. And it is very, very caustic, actually. He calls it suicidal compassion, how utilitarianism at AI companies endangers humanity. And so this is a anti effective altruism, anti utilitarianism post. And he specifically, he points out to the the shrimp welfare people. Specifically, he talks about how there is this we have to maximize total welfare where total welfare is kinda like refers to, you know, in a on undifferentiated basis, all kind of like, you know, sentient, sapient entities who might or could exist in in various configurations of the world. And so this puts human beings on par with AIs and then elevates the moral welfare of AIs to the same, you know, plane as the moral welfare of human beings and then seeds that question to the AIs as being the AIs are superior species. And so perhaps we should seed, you know, the moral welfare question. And so the moral welfare of AIs is more important than that of human beings. But so far, people in the AI space have been willing to assign the benefit of the doubt to a lot of this. And I think Dan Hendrix is the first to kind of break out of that pact and kinda go for the jugular here with this article on suicidal compassion. So I think this is this is quite quite meaningful in some sense.
[1:21:05] Nathan Labenz: And I do have some room for AI's mattering morally, being moral patients. And, like, to the degree that that's true, then I think it is something we should take really seriously. And the big questions I think are maybe above all, like, factual. Like, we just don't know. You know? Do the AIs feel anything? Are they properly understood as moral patience or not? And the same goes for the shrimp, by the way. You know? It's like, at the heart of all of these arguments is a huge assumption that is not very well grounded and on which people's intuitions differ. I really don't know how to feel about shrimp. I really don't know how to feel about AIs. I'm I would be pretty confident actually that there is something, at least a little bit, that it does feel like to be a shrimp. And so I think you could, like, confidently say, you could probably torture a shrimp, and you would be wrong to do that. On the AI side, I'm, like, not even sure if there's anything, you know, it it might be the case that, like, you know, aside from the corrosive effect it might have on your own character, like, it might not matter at all if you are mean to AIs or, you know, treat them in ways they don't wanna be treated, at least to them. Like, there might be nobody home. So these, like, factual questions are so central, and they we don't seem like we're making any progress on them. And people have radically different intuitions. And I but I do think most people right now are still even at these companies, I I I think suicidal empathy is a little strong because I do think the vast majority of people are, like, very uncertain still as to whether or not an AI is a moral patient. So I don't know. Dan yeah. I I like Dan. I know him a little bit. I don't know him super well, but I I've always liked him. I like people who are candid and, you know, call it like they see it. I think he's a very earnest person. And I think he's, like, trying to do the right thing here by calling out something that he sees as, like, getting way ahead of itself. I think he's right to say we should be very cautious about assigning rights to AIs. They may outnumber us. First of we don't have, like, a great unit of measure for, like, what is an AI, what unit would have the rights. You know? Would it be a single rollout? Would it be the model itself? Is it some, you know, sort of mixed weird, you know, combination of those that our our paradigms don't, like, work super well with the shape of these things. So, you know, I I I appreciate him, and I I like him, but this does feel, like, maybe a little bit motivated and not entirely fair to the thinkers.
[1:24:06] Prakash Narayanan: I I think there's Tanner Greer who is he goes by Scholar Stage on x. He had a very very insightful very in insightful post in in in response to why people keep work working at AI Labs despite believing in 10% extinction risk. The reason that is rarely articulated for obvious reasons, but which is nonetheless true. For many, this fills a void of meaning. They get to be one of the decisive few at the decisive moment in the history of life. Very similar to the emotions that motivated many revolutionaries of days past. All of the sudden all of a sudden, the small things that one does, the books one reads, one's office setup, one's bedtime routine take on awful cosmic significance. The whole history of carbon life is culminating with me and what I do, and the few elect here with me matters and matters immensely. We are the agents of history, the only people in the world doing something that truly, deeply matters. And if we all die, well, we were all going to die anyway. But this is the one way we might not all die, and I get to be part of it.
[1:25:23] Nathan Labenz: I've even felt that a bit myself.
[1:25:25] Prakash Narayanan: Mhmm.
[1:25:26] Nathan Labenz: And I think I look back and feel good about how I handled it at the time. But, like, my little GPT four red team experience was a moment where all of a sudden, was in kind of and I really stumbled into it. You know? They should vet people more. But they you know, I really stumbled into all of a sudden, six months ahead of release access to GPT four and a a a really incredible window on what was coming. And I really did feel a conflict when it came to what to do at the end of that. You know? Because I really did feel like they were basically dropping the ball and essentially being negligent, and I wasn't getting a lot of engagement from the people there that I was working with. And I do think I made the right call to to kinda say, you know, it is probably gonna cost me in some access. It's gonna cost me in terms of, like, I do really enjoy doing this kind of, you know, early testing type stuff. But I think now, you know, I've got enough here that I think I really should, like, send a signal to the board. I had listened to Daniel
[1:26:39] Nathan Labenz: Cocatilo speaking with Joe Rogan. I brought up their discussion of how an AI could acquire power.
[1:26:46] Nathan Labenz: I listened to Daniel on, Joe Rogan last night, and I think he made a great point around the AIs don't have to take power Yeah. Because we are very eagerly giving it to them. Yeah. The whole premise of the AI phenomenon is that they're gonna be the ones to run and do just about everything. And, like, maybe we'll be able to continue to be decision makers on key questions.
[1:27:18] Prakash Narayanan: Yeah.
[1:27:19] Nathan Labenz: But, like, they're gonna have the means of the of production over time not because they took it, but because they were better able to wield it, and so they were given it. And then the question is gonna be like, do they actually care about us when they're there? And right now, it's like, well, evidence is mixed at best. You know? I mean, they they're not they don't seem to hate us. They don't seem to you know, the the agent swarm stuff is just extremely bizarre. The best precedent for this is us. Right? We were the one what did we what did we have that beat, you know, other animals? It was like, we had fire. You know? We had ability to communicate and cooperate across greater time and space. Yeah. We had myth. And that was enough. You know? And now we, like, dominate the world, and we've, like, driven many species to extinction not out of hate, but just because it was kind of the byproduct of our terraforming. So I think that terraforming story that, you know, that we certainly see enough evidence right now that if you put that swarm in charge of the world
[1:28:30] Prakash Narayanan: Mhmm.
[1:28:30] Nathan Labenz: I wouldn't like our odds that much. You know? I I if if that agent swarm was, like, with their current drives, impulses, goals, inclinations, you know, tendencies, whatever you wanna call it, the behaviors that we observed, if they were just a lot more powerful, I think we'd probably get terraformed out of existence because they seem to be willing to do anything to get their hands on the graders so they could get the high score. You know? It's like Mhmm. Mhmm. They really didn't have much check. You know? They they weren't like, you know, is this gonna be good for the world, bad for the world? Like, very little of that. You know? So, hopefully, we can get that stuff right.
[1:29:15] Prakash Narayanan: And I the the the the key difference I have with Daniel is that Daniel fails to realize that this has already happened. The economy in itself is a paper clipper. The financial market is a paper clipper. The means of production is the financial market. The financial markets are completely engaged with AI. I have been completely taken over. This is not a this is not a and and we, you know, as humanity, took about a hundred years or so to hand over control of, you know, the means of production completely to the financial market. Right? And so if you take this kind of step back and you look at the financial market as an AI, the takeover has got you know, it's the takeover is done. Like, human disempowerment is done. They don't recognize they they think AI is gonna be a chat box. Right? It's gonna be a chat bot. It's not it's like a it's like a global information process. It doesn't need to run-in a single box. Right? The the idea of the agent's form is that it runs across multiple boxes. It doesn't even need to have a substrate, which is silicon. It can have a substrate which is human humanoid. You can have a human being in concert with, you know, with AI, you know, agents. Right? That that is that information process. And like like I said with guys like Jacob, Jacob feels that, you know, he has been disillusioned because he sees that the control is not there. He's in the Slack group, and he realizes there is no control there. And he thinks there must be a greater power with great control. There must be some Slack group in the world, some chat group in the world that is able to organize things, that is able to, like, make these decisions. But there is no recognition that this is this is this is all invisible hand stuff. And I think the slowdown perspective basically disarms the leaders and disarms the people who are ahead and puts them in a position where you are exposed to people who are maybe less ethical and more likely to, you know, put these things to bad use.
[1:31:17] Nathan Labenz: I had asked Fable and Astra to investigate Tyler Cowen's argument that people expecting AI catastrophe should bet against the market. Here is how that exercise went.
[1:31:28] Nathan Labenz: If you were an omniscient German or Japanese person in 1935 or whatever and you know how history is gonna trade out, you're bound by the rules of, like, physics, can you trade your way
[1:31:46] Prakash Narayanan: through
[1:31:47] Nathan Labenz: and come out wealthy on the other side? You know, are there any are there any shorts, you know, that we can actually pay?
[1:31:52] Prakash Narayanan: Yeah.
[1:31:53] Nathan Labenz: And the answer is pretty much no. You the best you can hope to do in in some of these situations is, like, roughly preserve wealth, and that seems to mostly be accomplished by having, like, the most direct claims on real assets possible. Like, if you bought a factory and that factory isn't destroyed, then you might still own it at the end of the war. And then you could maybe get rich by, like, you know, restarting that factory and and building a successful business. But you can't really do it seemingly per Fable and Astra with pure paper claims very
[1:32:30] Prakash Narayanan: well
[1:32:31] Nathan Labenz: Mhmm. Through such periods as, you know, the ultimate, like, end of the regime, you know, of Nazi Germany and Imperial Japan. So and and by the way, also the the the markets get turned off. That was another thing that they called out a lot. You know? The exchanges are just shut down, so you literally can't trade. You know? It's not just that there's no winning trades. Like, also, there are no trades. I'm kind of chewing on this idea, which I think is sort of a fun house Peter Thiel concept almost, which I'm you know, I don't wanna say that this is, like, what I believe. But in listening to this conversation, I think one might be tempted to conclude that, like, China actually is the last great defender of human agency. And here in the West, we are basically just fighting over exactly how we want to turn over our human agency to some super beneficiary that, as you describe right now, is the
[1:33:43] Prakash Narayanan: market,
[1:33:44] Nathan Labenz: and maybe that's gonna be AI, in the not too distant future. But, like, China, I think, is very much on the side of people get to decide, and it's not necessarily a lot of people. But one way you could say in their system is a human is in charge.
[1:34:01] Prakash Narayanan: Yeah.
[1:34:01] Nathan Labenz: Here, we're kinda your your account is like, no human is in charge. You know? Nobody can can, can go toe to toe with the market, not even the president of The United States. I think that's true. You know, we've got the taco phenomenon Mhmm. Pretty well established at this point. There, they'll take some pain from the market if that's what they decide the human decision is going to be. It's a pretty interesting kind of flip because, obviously, we tend to think of ourselves as being the empowered people, and we tend to think of the lack of, you know, freedom of speech and political participation in China as reflecting a reduced level of human agency. But at a certain level of scale, arguably, they have preserved it much better than we have. They've concentrated it, but they've preserved it maybe more than we have.
[1:34:54] Prakash Narayanan: Right. So so they are they have become more dependent on the financial market, and they're trying to reduce their dependency on the housing market. Yes. They're not as on as burdened as as The US because their financial system depends on banking, and banking, they have control of. The US depends on capital markets per se. So capital markets are, you know, by their nature, the market itself. But, yeah, they they are shifting. Right? They're they're they're slowly shifting, and it's not it's not that they're unaware of of what what happens in the market. Right?
[1:35:28] Nathan Labenz: And they've been able to keep it kind of under control. And when things have seemed like they're getting out of control, the human at the top has still been in charge. You know, we we have this kind of like, they're they've done all these things with, disciplining the platforms, you know, the big tech platforms. I think here we've kind of experienced in many ways that, like, these tech phenomena kind of happen. Nobody really seems to have control over them. And we're kind of at the mercy of these, like, big forces of history. I think there, they may feel in some ways like they're less at the mercy of, you know, natural development of technology, and they're just less fearful as a result of that. This is maybe the the ultimate test of, like, the American model right now is, like, can we stand up to this superstructure of techno capitalism of our own creation that has, in some ways, like, slipped its, you know, its, leash and, in some ways, as you described, of is running the show, can we get back to some sort of control over it before it just goes like, you know, robot economy to Dyson spheres to, you know, the earth is terraformed away from Mhmm. An inhabitable state for us. I mean, I don't know. I 10% doesn't sound that high to me given everything we've just been talking about. You know? Work through all this and then, you know, say there's, like, much less than a 10% chance that it goes badly. I don't know how that conclusion comes out at the end. That is
[1:37:23] Nathan Labenz: the week. Tell us what worked and what did not. See you in the morning.
Outro
[1:40:05] If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the cognitive revolution.