AI:AM: Was Trump-Xi Anything? What Counts as Utopia? + AWS GPUs Cost 3X & AI Diagnoses Rare Diseases
Hosts Nathan Labenz and Prakash Narayanan examine potential US-China AI cooperation, the economics of GPU rental pricing with Steve Hou, and AI applications in rare-disease diagnosis.
Watch Episode Here
Listen to Episode Here
Show Notes
Real things are starting to happen. Nathan Labenz and Prakash Narayanan revisit the AI:AM conversations recorded Monday, September 28 and Wednesday, September 30: cautious openings for AI cooperation between the United States and China, the economics of compute, a novel co-written with AI, and Daniel McKinnon's account of rare-disease diagnoses in previously unresolved cases. The week ends with Nathan asking what slower progress costs families who are still waiting for answers, while continuing to support some pacing.
The introductions and transitions are spoken in Nathan's cloned voice, disclosed in the episode. Guest and host conversations come from the recorded live shows. Tell us what worked and what did not.
AI:AM and the hosts: https://ai-in-the-am.com/ · Nathan: https://x.com/labenz · Prakash: https://x.com/8teAPi
Part I — A slight thaw
Jeremie and Edouard “Ed” Harris of Gladstone AI consider the possibility of a channel for communicating about AI incidents between the United States and China. Nathan hears a slightly more hopeful note from two usually hard-boiled realists, but their optimism remains qualified. An encouraging speech or a meeting is not an enforceable agreement.
The discussion moves from an explicitly unvetted emergency scenario to the work needed before a crisis: verification technology, the time required to vet it, and possible cooperation among researchers and AI companies. Nathan presses for positive actions the United States could take, and the brothers discuss how reciprocity might become possible. Their proposed options and assessments remain their own; this episode does not announce an operational hotline or an agreed limit on AI development.
Guests and organization: https://www.gladstone.ai/about · https://x.com/jeremiecharris · https://x.com/harris_edouard
Primary context: the Chinese government's text of Xi Jinping's July 17 address to the World AI Conference: https://english.www.gov.cn/news/202607/17/content_WS6a5a1172c6d00ca5f9a0c46b.html
The Chinese government's account of the September 24 Trump–Xi talks: https://english.www.gov.cn/news/202609/25/content_WS6ab59fdac6d00ca5f9a0d742.
Part II — The economy of intelligence
Steve Hou, head of research at Silicon Data, explains what a GPU rental-price index can and cannot tell you. Executable quotes can carry information that a simple advertised price does not. Contract length, availability, location and bundled services complicate comparisons. And measuring what someone pays for a machine is a different exercise from testing how that machine performs.
The conversation then turns to model-token spending. A shift toward more capable, more expensive models can change observed spending even as the cost of completing a particular task falls. The distinction matters when interpreting what the price of intelligence is actually doing.
Steve and Silicon Data: https://x.com/stevehou · https://www.silicondata.com/
Silicon Index: https://www.silicondata.com/products/silicon-index
In a separate host conversation, Nathan and Prakash discuss OpenAI DevDay: Dots and the appeal of a personal agent that somebody else maintains; Prakash's first impressions of GPT-6.1 Sol and faster access to Astra; and why latency matters when you are trying to stay in flow. These are the hosts' assessments, not a controlled comparison of model quality or speed.
OpenAI DevDay: https://learn.chatgpt.com/docs/whats-new/devday-2026
Dots: https://learn.chatgpt.com/docs/dots
GPT-6.1 Sol: https://developers.openai.com/api/docs/models/gpt-6.1-sol
Astra Ultrafast: https://developers.openai.com/api/docs/guides/ultrafast-mode
Nathan also considers Sign in with ChatGPT and the cost of letting customers try an AI product, drawing on his experience co-founding Waymark. The feature lets eligible users bring their plan to participating apps; it does not make every subscription universally portable.
Sign in with ChatGPT: https://developers.openai.com/siwc/quickstart
Waymark: https://
Part III — The craft of co-writing
Joel Borgen explains how he wrote The Receipt Horizon with AI: what he planned himself, the chapter structure he supplied, the models' role in critique and drafting, and the editing that made the prose readable. In Nathan's assessment, the result is legitimately good even where recognizable AI habits remain.
Joel estimates that his editing cut roughly 14 percent of the text. That is his account of the work, not an independently measured productivity result. His story is about a human author making choices throughout the process. He also describes a viola-clef transcription test he tried on successive models, and the musical collaborator he would like to have.
Joel and his book: https://joelborgen.com/ · https://x.com/JoelBorgen
The Cognitive Revolution's introduction and first four audiobook chapters: https://www.cognitiverevolution.ai/what-is-utopia-presenting-the-receipt-horizon-by-joel-borgen-chapters-1-
Part IV — The diagnoses still waiting
Daniel McKinnon founded Gamow Labs to use AI to interpret genomes and revisit unresolved cases. He shares the family history that led him to this work. After his son Owen died, a human specialist identified a missing 91-kilobase enhancer region that the original analysis had missed. Daniel's later AI prototype rediscovered it. The order matters: Owen's case is not an example of AI finding the answer before every human expert.
Daniel's public founding account: https://www.ddmckinnon.com/2026/06/09/vibe-coding-my-way-to-a-healthy-family-introducing-gamow-labs/
Daniel and Gamow Labs: https://x.com/danielmckinn0n · https://gamowlabs.com/ · https://x.com/GamowLabs
From there, the discussion moves to clinical labs' limited time per case, the evidence required to interpret uncertain variants, new biological experiments, and the engineering around the models: tools, routing and evaluations. Daniel distinguishes reanalysis of symptomatic patients from screening healthy people. Cheaper sequencing alone does not solve the interpretation problem.
Gamow's July case study reports that its system reproduced 19 expert molecular diagnoses and added two solutions in a selected cohort of 26 affected infants and 20 healthy relatives; five affected cases remained unresolved. These are company-reported retrospective findings, not evidence that a treatment worked or a child was saved. Daniel's on-air account of additional current diagnoses is also his report.
Company case study: https://gamowlabs.com/sota-genome-interpretation-with-agentic-ai.html
RareBench is Gamow's own benchmark for identifying and prioritizing causal genetic variants. A benchmark percentage is not the percentage of all patients a system can diagnose. The public benchmark description is here: https://gamowlabs.com/rarebench-0-1.html
For the discussion of earlier genetic-testing companies, Myriad's acquisition announcement provides primary background on Counsyl.
Counsyl / Myriad: https://investor.myriad.com/news-releases/news-release-detail/19626/
The closing question is Nathan's: if AI can help people who are still waiting for a diagnosis, how should we count the cost of slowing it down? He retains his support for some pacing and treats that choice as a costly compromise.
Links
https://ai-in-the-am.com/
https://x.com/labenz
https://x.com/8teAPi
https://www.gladstone.ai/about
https://x.com/jeremiecharris
https://x.com/harris_edouard
https://english.www.gov.cn/news/202607/17/content_WS6a5a1172c6d00ca5f9a0c46b.html
https://english.www.gov.cn/news/202609/25/content_WS6ab59fdac6d00ca5f9a0d742.html
https://x.com/stevehou
https://www.silicondata.com/
https://www.silicondata.com/products/silicon-index
https://learn.chatgpt.com/docs/whats-new/devday-2026
https://learn.chatgpt.com/docs/dots
https://developers.openai.com/api/docs/models/gpt-6.1-sol
https://developers.openai.com/api/docs/guides/ultrafast-mode
https://developers.openai.com/siwc/quickstart
https://waymark.com/
https://joelborgen.com/
https://x.com/JoelBorgen
https://www.cognitiverevolution.ai/what-is-utopia-presenting-the-receipt-horizon-by-joel-borgen-chapters-1-4/
https://www.ddmckinnon.com/2026/06/09/vibe-coding-my-way-to-a-healthy-family-introducing-gamow-labs/
https://x.com/danielmckinn0n
https://gamowlabs.com/
https://x.com/GamowLabs
https://gamowlabs.com/sota-genome-interpretation-with-agentic-ai.html
https://gamowlabs.com/rarebench-0-1.html
https://investor.myriad.com/news-releases/news-release-detail/19626/
Sponsors:
Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive
Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr
Deepgram Flux TTS: Deepgram Flux TTS is a streaming text-to-speech model built for voice agents with natural tone, context awareness, and interruption handling. Try it free until September 12 at https://deepgram.com/keep-talking
OutSystems: OutSystems is the leading agentic systems platform, helping enterprises build, modernize, and operate mission-critical applications at the speed of AI. Learn more and start owning your agentic future at https://outsystems.com/tcr
Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
CHAPTERS:
(00:00) Weekly episode highlights
(03:44) US China AI diplomacy
(08:12) Verifying compute and compliance (Part 1)
(12:46) Sponsors: Athena | Parallel
(15:37) Verifying compute and compliance (Part 2)
(19:57) Reading diplomatic green flags (Part 1)
(27:46) Sponsors: Deepgram Flux TTS | OutSystems | Claude
(31:44) Reading diplomatic green flags (Part 2)
(31:46) The economics of compute
(34:06) Hyperscaler cloud pricing premiums
(42:02) OpenAI Dev Day takeaways
(47:30) Cowriting novels with AI
(52:02) Human judgment in art
(58:10) AI diagnosing rare diseases
(01:04:01) Testing uncertain genetic variants
(01:12:20) Model harnesses and evals
(01:22:14) Costs of pacing progress
(01:24:09) Episode Outro
(01:27:19) Outro
PRODUCED BY:
SOCIAL LINKS:
Website: https://www.cognitiverevolution.ai
Twitter (Podcast): https://x.com/cogrev_podcast
Twitter (Nathan): https://x.com/labenz
LinkedIn: https://linkedin.com/in/nathanlabenz/
Youtube: https://youtube.com/@CognitiveRevolutionPodcast
Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk
Transcript
This transcript is automatically generated; we strive for accuracy, but errors in wording or speaker identification may occur. Please verify key details when needed.
Main Episode
[00:00] Nathan Labenz: Real things are starting to happen. This week on AI in the AM. On AI risk and verification, we heard from Jeremy and Ed Harris of Gladstone AI. They remain hardboiled realists about US China cooperation, but this week I heard a slight thaw. Here, Ed considers what reciprocal transparency could offer the two countries.
[00:21] Ed Harris: Certain kinds of transparency can be stabilizing. So the the kind of transparency that goes like, hey. We're giving you enough vision into what we're doing to see that we are not doing the thing you fear most. That that sort of thing is is potentially useful. And additionally, may not actually be that costly to us to do depending on how we implement it simply because the Chinese are already all up in our systems. So, really, we're not giving anything away that they don't necessarily have already in in many cases potentially.
[00:51] Nathan Labenz: Steve Ho, head of research at Silicon Data, which builds GPU price indexes. We asked why renting apparently identical chips costs so much more at the big cloud providers. So in the case of hyperscalers,
[01:04] Steve Ho: indeed, you are observing correctly, they charge regularly, consistently, at least two to three times, sometimes more compared to a typical... A new cloud. The reason has to do with a long legacy of, you know, whether it's other type of products being offered on their platform for software analytics, safety, compliance, the fact that they already have this long established relationship with enterprise users that have been on board for a long time that don't have a certain stickiness for moving. It is being sold as a very much of a differentiated product.
[01:34] Nathan Labenz: Prakash, my cohost on AI in the AM. We spent part of the week on OpenAI Dev Day. Here, what changes for building software when the models get faster at the same level of intelligence?
[01:46] Prakash: With Ultrafast, you can... As you type, you can interact. It's an interactive kind of build, software build out, actively building games. I think I think that's really the future. I think the speed, the latency... Latency at same intelligence is probably something that is gonna be very important, especially as you clear these hurdles of capability. Like, the thing can build a game, but can the thing build a game with you in the moment while keeping you in flow?
[02:14] Nathan Labenz: Joel Borgan co wrote his novel with AI models. The text has a few AI ticks, but the book is legitimately good. Here, what he supplies as architecture and why the pros still needs him.
[02:27] Joel Borgan: I mean, I I started the project knowing more or less what I wanted to have made. I sketched it out myself. Especially the first half or so of the book was pretty well set before engaging the models. And then the models are good at certain things. They're getting better at everything. But as far as just prose writing itself, even if you tell it exactly what you want and you have a plan for a chapter, you often get something that's unreadable unless you know how to scaffold it and then edit it yourself.
[02:59] Nathan Labenz: Daniel McKinnon founded Gammo Labs, which uses AI to interpret genomes. He reports new diagnoses in children whose cases had gone unresolved. Here, what he found when he looked closely at a model's work.
[03:11] Ed Harris: And you'll see things like there's one particular case where Groc 4.6, which at that point was state of the art on our benchmark, in Groc build, where it missed one because it just renamed the gene. It just, like, was saying, like, oh, you know, n r f two or whatever is responsible, and then it just changed the name of the gene to something totally different. I'd never seen that before, and I was just like, this is dumb. And our harness and our tools prevent the agents from doing dumb things.
[03:44] Nathan Labenz: Welcome to the AI in the AM weekly highlights with these introductions spoken in my cloned voice. Please tell us what worked and what did not. Your feedback helps us make the next one better. Part one, a slight thaw. Jeremy and Ed Harris have been interviewing diplomats who negotiated with China. We discussed the Trump Xi talks and the possibility of a channel for communicating about AI incidents. Ed starts with what he makes of that contact.
[04:16] Ed Harris: Talking is always better than not talking. So that's that's a positive. And my understanding from at least the beginnings of the Trump g conversation and the stuff leading up to that is that one of the things that may be positive that came out of that relationship was, you know, the development of this... I don't know if you'd call it a red phone, but of at least some kind of theoretical line between the two governments on AI incidents and AI risks. The... What what... The the report that we came out with, which is really just a long newsletter, is informed by speaking to about a dozen state department diplomats who have dealt with China from the negotiating table and who've kinda seen how these things kind of develop in practice. And and one of the issues that they do see among many others is we have tried the red phone thing before in the context of nuclear, and by and large, they they don't always answer. And particularly in in critical phases, it's often exercised as a point of leverage saying, we're gonna take away this this phone line and not answer rather than as something that is, like, this this collaborative unified project that's... Makes everyone safer. Again, this is not to say if the CCP does, structurally take AI seriously, which there is, you know, decent reason to think that they may. We could have a whole different and much more positive level of engagement. All that we're recommending based on the experience of these folks is being realistic about it and having a backup plan.
[05:44] Nathan Labenz: I asked Jeremy, after a serious AI incident, what could The United States ask China to stop doing, and what could be verified with the capabilities that exist today?
[05:57] Jeremy Harris: Everything that that comes after the sentence obviously has not made contact with the intelligence community from a red teaming standpoint. So the true answer is we can't know deeply. I couldn't give you an answer to a level of detail where it would be like, okay. You know, that's executable. One easy thing, if I'm gonna caricature, data centers put off a hell of an energy footprint. You know, the thermals on those are really, really bright. Data centers are huge. They haven't, by and large, yet been built to be hidden, and it takes a long time to build data centers. Now this will change. AI twenty twenty seven talks about the timelines for this. We think it's quite plausible that that, the timelines could be a lot shorter for hiding data centers just based on conversations with folks in the industry. But, like, it... Whichever way you slice it, it... Like, you're gonna have an initial conversation where to first order for the eighty twenty that you really need then, it's like, so help me, god. If I see a cluster and that cluster is yay big, you know, that's, like, the kinda conversation that you're looking at. And how you quantify that is a matter of, you know, the sort of national technical means that The US currently has this So
[06:59] Prakash: so wait. Are you saying, having a cluster above a certain size would be a red line?
[07:05] Ed Harris: So I'm saying, yeah. Initial. So if you're if you're just in that panic moment, right, you're like, we have to do something. And you ask yourself, what is possible? Like, what is possible to do with the existing assets and infrastructure that we have today? Nothing else. Then you do you do get into a space not necessarily where there exists a cluster of this size because you... It's... Know, you can't ask them to tear down the cluster, but we have to see the heat signatures from this, like, go away. So if there is a running cluster above a certain size. And that is, to be clear, a tremendously expensive ask in either direction. The the depreciation on GPUs is is the major part of the OpEx cost.
[07:46] Prakash: So so so you're saying an incident happens first. Yep. And the response to that incident, the mitigation for that incident is, hey. Can you turn off this big data center?
[07:55] Ed Harris: Yeah. Like, turn off any cluster above a certain size. That's one one possibility because this As
[07:59] Jeremy Harris: of right now, to be clear, like, the the framing is basically, as of right now, that's where we're at. And so in some sense, we're gonna get into a a potentially a circular loop here where you you can see how big of an ask that is. That is an insane ask.
[08:10] Ed Harris: So if I kinda sketch out
[08:12] Nathan Labenz: the logic from beginning to end here, it's like we don't have that great of a relationship. AI capabilities continue to progress at a fast pace. We expect something crazy to happen. When something crazy enough happens, we're gonna find ourselves by default in a spot where we have to ask for some outlandish super high cost move, like shut down all your big data centers because we don't have any other mechanisms in place that allow for a lower ask, a better trust but verify type of environment because we haven't made those investments now. That leads me to the question of what should we be doing now to, a, ideally not end up in that situation, or, b, if we do end up in that situation, have better options available to ask for aside from shut it all down, which is obviously gonna be tough.
[09:12] Ed Harris: Develop. Basically, it's like, you know, you absolutely nailed it. It's, develop better verification and develop, better offensive options to ensure compliance in the event that, verification is is... Verification returns know they're doing it. The better verification stuff you can do, the faster. The less you have to rely on absurdly expensive things like this gigawatt of energy radiation shouldn't be visible from space. If we have techniques like this that are vetted by the intelligence community, even if they are even if they are just, like, 50% better than this, even if it's like, you can... You you have to shut down half your data... I'm making something up here. You have to shut down half your data center, and we can, like, kind of, sufficiently verify the other half or something equivalent to that. You are saving billions of dollars right off the bat. And so the ability for this industry to, like, continue to make large amounts of money is actually gonna be gated for that period of time by these little verification technologies. And there's already this community of little verification startups that's working on on this technology, and that's why this is so important. And, of course, I will also say, the the offense side of things is critically necessary. If you don't have those offensive options, you cannot assure compliance you cannot assure compliance. You can verify, and you can monitor the situation, and you can
[10:31] Jeremy Harris: say, oh, no. But, fundamentally, your your hands are tied. You don't have the tools to actually do anything about So both of those things are super important. There's this kind of bottleneck that I think a lot of these verification companies haven't really necessarily priced in. And and it... Like, so this is kind of where a lot of our our current work is focused. So imagine what happens when company a goes, I have the thing. This thing is gonna work. Right? And and it's a moment of crisis, and and Trump is casting about... Or POTUS, whoever it is at the time, is casting about for options to alleviate this, like, the trillion dollar bottleneck. And then they actually go, okay. The intelligence community has to now vet this. Because they're not gonna just, like, start using it, obviously. Right? So how long did it take similar technologies in the past to get used, to get vetted, and to become what's known as national technical means, NTMs? Right? And the answer is is years, depending on the technology, but very often years. We have to do it. China has to do it. We have to handshake on doing it. And, even if you remove money as an obstacle and so object, there's just certain things that take serial time to do. And so so a lot of what we've been doing is focused on saying, okay. Treaty... Or not treaty. Let's say, AI agreement verification or compute verification company x, you probably should be talking to IC element y about this. Because in a moment of crisis, you, a, want as much pre vetted as possible, and, b, anybody who's involved in assessing a potential national technical means had better have on speed dial. They better have the signal, the phone number, the email, whatever, of, like, the founders of, like, all the verification companies that they plan to use or they may end up having to use. You wanna cut down on all those those those barriers, the boring bureaucratic hurdles that nobody ever thinks about because they're boring and bureaucratic. But these are the things that we're just, like, trying to shatter right now so that when game time happens, things move more quickly. And today, those companies that potentially are pursuing research trajectories or agendas, in some cases, like, multi, multi, like, your $10,000,000 plus research agendas, that are just oriented in a way that, unfortunately, like, that appropriately placed person in the intelligence community would look at and be like, that's kind of a nonstarter. We want them to get that feedback as soon as possible so they can reorient on things that do have a chance of working.
Sponsor
[12:46]Athena: Athena matches you with a dedicated, top 1% executive assistant to handle your inbox, calendar, and daily workflows so you can save an average of 15 hours a week. Get matched with your EA today at https://athena.com/cognitive
[14:18]Parallel: Parallel provides enterprise-grade web search APIs for AI agents, offering the optimal balance of quality, speed, and cost. Get started for free at https://parallel.ai/tcr
Main Episode
[15:38] Nathan Labenz: Still with Jeremy and Ed on verification, I turned to work that could start among the AI companies themselves. How important do you think it is that we remain on the same fundamental tech tree or AI paradigm across US and Chinese AI development. My sense is that, like, we're in, in some ways, a very fortunate position right now because we're basically building the same tech in the same way, and we're sort of encountering the same surprises along the way. So that's one thing. And then the other question is, I feel like if there's anything good to be found in the OpenAI Anthropic adversarial dynamic, it would be that maybe they can be a test bed for techniques that might later scale to a US China dynamic. And so if I was the president, I would say, you two have to figure out a way to police each other. And maybe we expand that circle to a few other frontier companies. But, like, you two, you're the ones setting off all these alarm bells. I need you guys in a room, whatever technology you need to develop, whatever access you need to give one another, it's on you to figure out a way that you can trust and verify one another. And then maybe we can kinda scale that up to a transpacific dynamic that could work similarly.
[17:02] Jeremy Harris: I think that's that's not that's not insane. Obviously, there are big differences between, you know, what Anthropic and OpenAI respectively have on each other. Although poaching of personnel does a decent job of mirroring the kind of access that China clearly has the Frontier Labs anyway.
[17:17] Ed Harris: But there There's also a basis of trust, I think, between two fundamentally US companies with fairly close to similar values and blah blah blah versus, like, fairly radically different. Not to say that this is, like, totally a mess or whatever, but, like, there there are gonna be some differences as well, some similarities. But I think the similarities might, might be worth mining. Yep.
[17:37] Jeremy Harris: Yeah. And it... You know, it's it's also the case that, like... So you're talking about, the idea of the importance of the stacks being aligned. The hardware lottery does a lot of really good things, in this space. It... I think the most crucial thing is The US and China clearly don't trust each other in terms of the the motives that bring the... Bring each respective side to the table. Right? So when China sees The US come to the table and raise issues like, so, you know, slowing down the air or safety guardrails, whatever, the interpretation that we've heard consistently from people involved in, like, the track two or track 1.5 kinds of of dialogues... And and, Nathan, I mean, you've you've been kind of in that ecosystem or touched it as well, is, you know, the Chinese kind of view it as an attempt to curtail their own development because they see themselves as being behind, and they're justified, therefore, in doing things that even wouldn't be appropriate for America to do in their eyes just to catch up because they're in second place. Being in second place does also mean... Being in second place with respect to scale does also mean that you don't see the warning shots with the same resolution. However, they seem to also be more public than at least I would have expected, and so the spillover is something that's nice that China can just verify themselves directly. So that that might be a mitigator if you see the same kinds of failure modes emerging from, whole brain emulation or any... You know, if if some completely wacky other branch of the tech tree were to become dominant in China. So I do think it's good. I think we we kinda get there by default. It's hard to imagine alternatives that, like, really shake things up at this point. Quantum machine learning, if you wait long enough, I just don't think that's gonna be relevant on the time scales that matter that could could radically reshape algorithms. But what we're seeing right now is an industry that's more or less doubling down on transformer MOEs with some bells and whistles and some some variations here and there, but everything's kind of a transformer. And, and that's that's what seems to shift. So, yeah, I I expect that we will have that... The benefit of that. It also comes with the benefit of being able to share safety technology as The US did with Russia during the height of the Cold War at times. And so in principle, that does mean that we can work on each other's safety stacks, and that might be the source of some trust building measures, though that term also is problematic for China, and they've also pushed back on attempts to do that sort of thing in the past. So, I mean, it... It's a... I think it's it's a good thing. I don't know how far it goes.
[19:57] Nathan Labenz: With the Harris brothers, the discussion returned from verification technology to the diplomatic evidence for cooperation. I want to not be naive, but I do want to notice and give appropriate weight to positive signals as I see them somewhat developing. I think Xi's speech at the WAIC was, like, pretty friendly, pretty conciliatory. He, like, gave credit to The US for inventing AI. You know, he certainly didn't call for an international arms race. He warned against overstretching the national security concept. We can dismiss that as just nice talk. Probably should have at least some weight on that possibility. But what are the, like, meaningful things that you are watching for the decision points that will update your thinking on? Are they inclined to at least try to control AI for their own narrow self interest? Are they inclined to meaningfully cooperate, or are they inclined to seek some sort of domination as as we often project that we are interested in doing?
[21:04] Ed Harris: So in terms of what signs to watch out for, I think anything that looks like a positive sign is at least a positive sign to some degree. So conciliatory speech is good. It's good. At least it's not a hostile speech. Right? It could be worse. Everything is kinda tempered by the fact that, like, rhetoric is is is often used for strategic purposes in this way and to kind of shape the battle space in terms of the narrative and so forth. The the the point is not, like, these are not positive signs. It's just, like, we have to weigh the evidence in the context of the credibility that has or has not been established by this entity over the past span of time. In terms of what are maybe slightly more unfakable signals that they could give off that that you would make me go like, oh, woah. Okay. This is... This looks legit. You can imagine maybe, I don't know, something like DeepSeek and, like, Zifu coming to some sort of pacing agreement because they, some crazy thing happened over there. They're equivalent to the Hugging Face incident. You can imagine, the the cyberspace commission that that has... Or the commission that has jurisdiction over the... Whether models can be released and, like, are they, you know, properly ideological and stuff, actually putting the brakes on something because, there was some misalignment thing. And they're actually, they... They're finding that a model that was properly ideological in testing suddenly is is not being ideological in the wild or something like this. I would say indications where that that seem genuine that that they are starting to be on the receiving end of these incidents at the level... The same level of detail that our own Frontier Labs. Like, I think that begins to to make us... All of us as humans go like, oh, you know, maybe the the thing we should be concerned about is the giant the giant, like, shaga thing that we don't understand and that we're growing in the labs, at at accelerated pace. Like, that that would make me feel a little safer. Yeah.
[23:03] Jeremy Harris: And and maybe procedurally too. Like, so if if you take... We have a list of, I think it's, like, five different historical traps that we've seen in US China diplomacy. The... These are essentially, like, a a brief catalog of the ways in which China behaves when they're full of shit. At least, like, by by the assessment of of a lot of the diplomats we spoke. I should be clear, actually. There there was a, a dissenting diplomat
[23:25] Steve Ho: who
[23:26] Jeremy Harris: felt that some of these things were much more sincere, including the use of the the language of the... You know, the objections over language, arms control, this and that. And that itself is the epistemic problem that we talked about earlier. But, basically, I would say take each of those red flags and then flip them over, and you get the corresponding green flag. And so if you if you don't see arbitrary, concerns raised about language that seems random, that's a green flag. If you see engagement... This is actually really important. If you see engagement from empowered people, actually, arguably, as we did, right, with she, though, again, you gotta calibrate everything with, like, the... We've seen this before in other contexts, but it's certainly not a red flag. But in terms of, like, concrete commitments, that's the sort of thing you look for. Empowered people, the lack of sort of capricious arbitrary objections, things that actually look like they're making progress qualitatively is actually a surprisingly good sign because you can contrast it directly with how things have gone in the past, which is not very good. I mean, the contrasting point is is actually that low that it can be genuine signal.
[24:26] Nathan Labenz: What are the things that we can do that are not so costly to us but are still credible signals to them that we are not going to try to use AI to gain a decisive strategic advantage and ultimately make them an offer they can't refuse. If indeed that is not what we're going to do, which I'm a little worried we might actually be about to try to do that. If we were on the path of trying to seek a Pax Robotica where we can all benefit from the abundance that AI, especially in its Chinese manufactured embodied form, might provide for us, what would be the steps that you would prioritize next on our side?
[25:11] Ed Harris: Well, there may be some stuff we can do that's not functionally really even that costly. Generally, as as I think, Jared, you guys maybe as well highlighted earlier, certain kinds of transparency can be stabilizing. So the the kind of transparency that goes like, hey. We're giving you enough vision into what we're doing to see that we are not doing the thing you fear most. That that sort of thing is is potentially useful. And additionally, may not actually be that costly to us to do depending on how we implement it simply because the Chinese are already all up on our systems. So, really, we're not giving anything away that they don't necessarily have already in in many cases, potentially. So maybe some kind of, you know, level of visibility into here's what we're doing. Here's what the Frontier Labs are doing and so forth. The problem is, like, that's kind of not necessarily something you wanna be doing unilaterally. That would be something that that would come as part of, like, trust building measure, dare I say, between two powers in the wake of a a a moment like this. But stuff... I guess, goodwill type stuff we can do now certainly would be more interactions between the verification communities in The United States and in China. This kind of thing is already happening, actually. There's some there's some quite good and positive interactions between those communities. So, yeah, the the kind of linkages that you get at the level of academic to academic, startup to startup, all kinda trying to solve for the same mission, these are very, very positive things. They're also like the the the... It's true that right at the political levels, the two countries have started separating out and and and and even at the level of kind of big companies and stuff like that that you see the the Chinese kind of steel leader designs and turn... The the classic, like, Chinese, like, spin off story and all the stuff. But there still are, like, real genuine linkages between the two countries where... Especially on the academic side and in number of other levels, there's, like, there's a bunch of sincere people just talking to a bunch of sincere people about, like, yeah. This is a problem, and this sucks. Yeah. I agree. Like, let's try to solve it. So the more of those linkages there are, the better. Like, the more people there are on on both sides of the ocean who have the ability to talk to their own kind of domestic leadership and say, look. I've spoken to them. Like, they're not they're not evil. They're just trying to do this or that. To whatever extent that's true, that kind of moderates the more extreme tendencies on both sides. It's very, very hard to do that completely because the actions of these countries are also constrained in a number of ways, but, like, it really does help, I think. Part two, the
Sponsor
[27:46]Deepgram Flux TTS: Deepgram Flux TTS is a streaming text-to-speech model built for voice agents with natural tone, context awareness, and interruption handling. Try it free until September 12 at https://deepgram.com/keep-talking
[28:17]OutSystems: OutSystems is the leading agentic systems platform, helping enterprises build, modernize, and operate mission-critical applications at the speed of AI. Learn more and start owning your agentic future at https://outsystems.com/tcr
[30:09]Claude: Claude is the AI collaborator for problem solvers, helping with writing, coding, financial models, strategy, and more. Get started with Claude and explore Claude Pro at https://claude.ai/tcr
Main Episode
[31:46] Nathan Labenz: economy of intelligence. Steve Ho leads research at Silicon Data, which builds price indexes for rented GPU capacity and model tokens. Turning compute into a measurable market means deciding which prices can be compared. Steve starts with the inputs to the GPU rental index.
[32:06] Steve Ho: And we use a combination of both quota prices and transaction prices in our calculation, making those distinction clear. Right? So it's... This is a large, you know, sort of... So to speak, a a normalization process, machine learn... Using machine learning to to help us make the contracts apples to apples comparable to the extent that you're looking at a single chip, chip, let's say, or h 100 being rented from different parts of the world. So one question that jumps to mind right away is that, for example, in the case of a apartment rental index, you will never consider to use just an offer price. Someone who lists a number on the front of a building, say, you can rent an apartment for $2,000 a month, you wouldn't expect you to just trust that number. You can walk in. But we do use quota prices. Why is that? The reason is because when... Unlike, you know, an apartment, you know, you cannot click an API... Unpass an API and just get hold of the apartment. You have to walk in and talk to somebody and negotiate, and that probably may not be available. In this case, very often, right, with API executable, you know, GPU rental, you can actually get ahold of the GPU just same way you can, you know, buy something on Amazon. Right? That being said, we also have a transaction, so we... Comparing them so that if someone who has three nodes of GPU that rented out, say, all all three at $2.75, I will not presume that I can go back to the same merchant to run another one to... At $2.75. It could very well be the case that the next one is not available anymore. Or if they run out two out of three for $22.75, the next one will not necessarily be $2.75 again either. It could be $4 or $6 or $1 depending on what everyone else is quoting. The quote person will be crazy if everybody is quoting at $4 and they continue to run that 275. You would think, I... This... They're gonna change... Raise price or there's something wrong with the price. Right? So this is the reason why I gave you a long answer again, but what we do is that we want to provide as broad a set of coverage of the market as possible and normalize everything so that we're capturing the market as it is, as faithfully as possible.
[34:07] Nathan Labenz: One big difference I noticed in prices just browsing the Silicon Data website is Mhmm. Between the neo
[34:13] Ed Harris: clouds
[34:14] Nathan Labenz: and Yes. The hyperscalers. These are not small differences. These are, like No. Multiple differences in prices, like, seemingly two to four x. That's a that's a pretty big delta. Why does that delta exist?
[34:31] Ed Harris: Mhmm.
[34:31] Nathan Labenz: Is there no way to arbitrage it? And Yeah. Does that imply that, like, GPU hours are being wasted or that they're... That when I buy from a hyperscaler, I'm sort of preempting their internal work, and they're using everything that's not sold Mhmm. At runtime? Give me a little peek behind the
[34:49] Steve Ho: So you so you can buy... Forget about... Again, like, I like to use analogies. I gave you analogies in apartment... You know, single bedroom apartment. I'm not gonna use an apartment again, although I can. You can... By the way, let's say you imagine we're doing a single patty burger index. Right? You can buy a burger from a burger stand off of the street Of New York, or you can walk into a high end steakhouse and order a burger. Believe you me, like, those two burgers are gonna cost very different. Right? They're both burgers that... You know? And what we try to do is we try to, as much as possible, find the marginal price for a unit of a compute or burger in this case that is sort of have the same relative, the same feature. I wouldn't want to compare the price of a three, you know, a a three piece, you know, a a three patty burger with a single patty burger. Right? But once I make those adjustments, there are some adjustments I cannot reasonably make because it's capturing a different type of premium from product bundling or product differentiation. So in the case of hyperscalers, indeed, you are observing correctly. They charge regularly, consistently, at least two to three times, sometimes more compared to a typical a new cloud. The reason has to do with a long legacy of, you know, whether it's other type of products being offered on their platform, for software analytics, safety, compliance, the fact that they already have this long established relationship with enterprise users that have been on board for a long time that don't... That have a sort of stickiness for moving. It is being sold as a very much of a differentiated product. So going to that single patty burger index analogy I gave you, I see people who sell single patty burger with fries and milkshake. I can try to strip out the prices of those two elements and isolate what what I think will be the price of a single patty burger from that merchant. But if they, let's say, they blended up the both fries and the and the milkshake and inject it into the into the patty and say, this is a brand new product and charge two times, three times the price, I can't very easily, you know, sort of strip it out. Right? At which point, I say, okay. I'm raising my hands. Okay. You guys are a little bit different of a beast, and let me put in you a different category. Right? And that's how we have so far handled it. We, you know, we believe that the way... This this way, you are going to get into a more... A much purer form of a single unit of compute. Right? You know, that is actually be getting closer towards this idea of fungibility to the extent that things are direct substitutes to each other.
[37:07] Prakash: The same chip can perform differently at different places. And I think according to your site, you
[37:12] Joel Borgan: have
[37:12] Prakash: some software that runs on the chip.
[37:14] Steve Ho: SiliconMark. Yeah. Mhmm.
[37:15] Prakash: Yeah. That you that you benchmark the chips with.
[37:18] Steve Ho: So
[37:19] Prakash: I guess every single price in your index has been benchmarked by this... By the system?
[37:25] Steve Ho: No. So we have a physical benchmarking service called SiliconMark that actually visits individual GPU at the EUID level to try to assay... To assess the GPU health, the performance, you know, the throughputs, you know, t flops, and so on. At a moment, that physical benchmarking and physical spec performance does not enter into our pricing. Right? We do not see a strong relationship, at least, at at enrollment given the nature of the market, a a relationship between how the GPU is specifically performing. Even though we do see that in the cross section, you can actually have a bit of a variance, right, in... From GPU to GPU. This is not, I think, surprising or new to anyone who is, you know, sort of in this space, the GPU lottery. Right? But over time, if you have a cluster that probably averages out and and and do all flash numbers kicks in. So we don't use that. We have six features, including geolocation, you know, sort of CPU memory and and and various things and term and provider, but this is... But the physical spec is not one of them at this at this very moment. Eventually, when we head towards a scenario where we potentially could have physical delivery because inference, let's say, based on my, you know, sort of thesis that could make compute more interchangeable, we... That we could enter into pricing scheme. By the moment, it does not.
[38:40] Nathan Labenz: Another price that I noticed having moved on the website that caught my attention is the proprietary LLM index. Yeah. There, you you break down just token pricing close. Yeah. I know overall. Yeah. And so the the proprietary one, have up 11 and a half percent over the last seven days. And I guess I'm wondering, how are you measuring that? Because it seems like, you know, quality adjusted, everybody would say prices are coming down. Right? Even, like, dramatically so. And, like, the retail, you know, posted API price hasn't changed, right, except when they kind of introduced new models. So what are you measuring on a day by day basis that allows you to say what the proprietary cost is doing at such a fine grain of resolution?
[39:31] Steve Ho: Yeah. So first of all, I want to give a little bit of context for what our token indices mean because they've been, I think, I think, misunderstood. I think, part of it has to do with the unfortunate naming. We called it, you know, the expenditure index, and then people maybe thought that it either means price or total volume when it does actually mean either. It's actually an expenditure weighted price index that's normalized to a million. And so that can show trends based on usage mix, which I'll come to in a second. When you point out the proprietary LLM model, which is sort of the frontier model, having recently bounced a little bit higher, Should put it in a context that since, like, maybe late June through, like, basically, you know, middle of this month, it has been on a very sharp downward trend of going from some $4 to, what, you know, 1.6 or something that's more than 50% drop. Right? During this time, we have seen a lot of not just frontier, you know, leading labs cutting prices on their, you know, sort of newer variants by releasing cheaper variants of powerful models, new generation models. But, also, we've seen other proprietary models coming out like Meta and Grok. Don't forget these are proprietary AI models as well. Right? And they've been very aggressive on the price front. So more recently, I think that the bounce could come from a a variety of sources. Right? You know, basically, if people decided to say, okay. I actually quite like, you know, the the more expensive power model from, Anthropic, and they use more of it, that could actually drive up the expenditure weighted index... Price index. The the analogy like to give people a little bit is, like, forget all these are, again, LMs. Forget this. Imagine these are, like, cars. Right? You know, if you have, like, you know, a a Mercedes have, like, sort of cheap variants and expensive variants, depending on which cars are being sold more and people are liking more, that could affect the average price of the car sold. Right? In the same thing here, right, in the market of tokens, there are two things happening. You know, token model prices are changing, but usage behavior is also evolving. And to the extent that we observe volume from a handful of these public inference platforms, learning platforms that allow you to look at how people use different type of models, I think this most recent balance of the front... Frontier proprietary element index, I don't know. I haven't looked into details, but suspect has more to do with usage mix than anything else.
[42:03] Nathan Labenz: From measuring compute to using it. The next conversation is just the two hosts discussing OpenAI Dev Day. I start with Dots, the personal agents OpenAI presented, and who might want a system they do not have to maintain themselves. It was funny. I thought when they led off with the whole dots thing, I was like, well, I think I have all of this, and I'm pretty sure I'm gonna continue to prefer my version that I've gradually evolved over the last nine full months now. So that was definitely... It didn't... That didn't really feel to me like a developer product. I thought that that was a little bit muddled, because it very much felt to me like that's a consumer product. Right? Like, that's for JGPT users to use, and it wasn't entirely clear how that would be used by developers if at all. I I I haven't been down every last, you know, breakout session video, so there's possibly some more that I missed. But certainly at the keynote level, it just felt like that's a product that they are offering on a first party basis to their users, and I do think it will be really useful for people. I I guess I would say... My guess is that people are really gonna love these dots. But certainly for, like, my mom, on the other hand, you know, I would say, go for it. Just use that. You know? It's it's probably pretty easy. You don't have to worry about taking on all this stuff yourself, managing your own database on your computer, troubleshooting when things go wrong, even though the models are getting so good at that on their own. I do think this higher level and and more polished abstraction will probably be really good for a lot of people who don't care to learn a bunch of new tricks and just want to have this thing that they can delegate to. Prakash then turned to the models OpenAI presented at Dev Day. These are his first impressions of their capability and speed.
[44:02] Prakash: They had GBD 6.1 Soul and GBD six Astra Ultrafast. G p d 6.1 Soul is about the... I think the same price or or slightly lower price than I think g p d six Soul, but it's basically Astra for the price of Soul, which is what they're calling it. It seems to be a very competent model. I've used it. It's better... It's definitely better than g p d six Soul. G p d six Soul was only released, like, don't know, two weeks ago, two, three weeks ago. So cadence cadence of releases is stepping up. G b d six Astra Ultrafast. So now you have... They used to have a fast mode, which is two times speed. Ultrafast is a eight times speed, and they demoed how Ultrafast works. With Ultrafast, you can, as you type, you can interact. It's an interactive kind of build... Software build out, interactively building games. I think that's really the future. I think the speed, the latency... Latency at same intelligence is probably something that is gonna be very important, especially as you clear these hurdles of capability. Like, the thing can build a game, but can the thing build a game with you in the moment while keeping you in flow? I think that's I think I think that's what's coming up next. Still in our dev day conversation, I turned to sign in with ChatGPT, which lets eligible users bring their plan to participating apps. My example comes from Waymark, the video company I cofounded, and the cost of letting a new customer try an AI product.
[45:34] Nathan Labenz: So the classic try one, try seven days, whatever, those are tried and true tactics that have become very difficult in the context of, oh, but I gotta spend a certain amount on tokens for every new user to give them that decent experience, especially if you have something that's, like, involves a decent amount of setup or profile processing. With my company, Waymark, we're not, by any means, like, the most token hungry business. But the first thing that we do when you sign up is we make a big profile of your small business so that we can then use that profile as an input to make video content later. And that profile creation process, it's come down in price. I don't know exactly what it is today. At one point, it was, like, a dollar per customer. And it was like, okay. Well, this does start to become material when you think about the all in cost of customer acquisition, and can we make this flywheel work, and how fast do we get paid back, and all that sort of stuff. But I think this is a this is a great value add to your ChatGPT subscription that you can now take it around. And, of course, developers will need to implement this, but it shouldn't take too long to tell your coding agents to implement this. So that's great for ChatGPT, great for users, great for the SaaS companies. I think it's a huge win, and I I do think it's, like, pretty pro ecosystem. It seems to me that they this is, like, one way in which they can actually follow through on the promise of, like, not trying to eat the world and instead trying to, like, empower people to to build cool stuff. Part three, the craft of cowriting. Joel Borgan is the author of The Receipt Horizon, a novel set in a world shaped by advanced AI, which he wrote in collaboration with AI models. We discussed what he supplied as the author and how the collaboration changed the work.
[47:31] Joel Borgan: I may have written the book eventually, but the activation energy that... That's required to get started when you have, you know, kids and a part time job was limiting. I mean, I started it well over a year ago, and the tools at the time were not as developed as they are now. There were a lot of limitations, especially around context length. And you have the thread where you're trying to work on a a chapter or part of the book, and you're constantly having earlier context drop out and being to manage that carefully. I think it's getting easier and easier to make use of the tools, to do something collaborative and still kind of your vision. I mean, I I started the project knowing more or less what I wanted to have made. I sketched it out myself, especially the first half or so of the book was pretty well set before engaging the models. And then the models are good at certain things. They're getting better at everything. But as far as just prose writing itself, even if you tell it exactly what you want and you have a plan for a chapter, you often get something that's unreadable unless you know how to scaffold it and then edit it yourself. So I've been working on a follow-up book, actually, and the process has changed quite a bit as as the models have become more powerful. So the book is definitely a lot better for having done it alongside the AI, but I'm cognizant of the fact that, you know, AI writing will be controversial for a lot of readers and able to engage with it. And I think the process made the novel a lot better, and it also had weaknesses as well that I didn't entirely, like, account for.
[49:11] Nathan Labenz: I had asked Joel about building a positive vision of the future into a story that still needs conflict and about the scaffolding he gives the models.
[49:21] Joel Borgan: Utopias often want to become dystopias in narrative form because you need to have conflict and friction. And, it... It's kind of dull just if everything work out and everybody's happy, of course. So the fact that this was set, you know, after a a big AI war and humans are sort of presumably locked into what they're able to do, you see both incredible flourishing in terms of the technology that's available and the options people have and also some, like, really major downsides. So I do think so much of what, you know, you've talked about, and I agree, is trying to paint, like, a positive vision for the future, and that is difficult to do in, like, narrative form. So I didn't want any... Anything in the book or any character to be, a mouthpiece for me. It's more... A lot of the stuff that I want to see created in the world is a part of the story and a lot of... I I present, you know, trying to steal man the that have... Would have different opinions about that as well. So as far as the the scaffolding, I often will will write out kind of a large scale architecture for the story and a lot of beats and then go back and forth with both the major frontier models. I mean, I've primarily used whatever the flagship Claude model is and then the Pro model from, like, five on because that was the most... That was kind of the workhorse. It did the most kind of detailed work. Claude used to be a lot lazier as far as, like, how much you would follow-up on. And I'm doing it mostly through the chat interface, which is probably... Has its advantages for me and also probably limitations. But, you know, context length, like I said earlier, was a huge unlock as that got longer. You could just... You could give them the novel the entire book. Because at one point in drafting the first one, and I would... You know, every time the new model come out, you'd you'd use that to to try to make it better and up your game and allow it to do more. But as far as the the actual scaffolding for the story, you you kind of go in chunks often, and I'll I'll draft chapter by chapter. And then, you know, when I was done with it, I I had the advice to to hire professionals for copy editing and art and layout. And I did that, and I learned a lot from doing it. It was very expensive and time consuming and kind of slow, and there were definitely frustrations involved. But I've tried to kind of extract what I've learned from that and what made the book work better. And I've, you know, I fed it to five or six now pro and developed my own workflows for, like, doing the same process as the second book. So I think that'll be an interesting experience to see how all that works.
[52:03] Nathan Labenz: With Joel, we moved from drafting to judgment, where he still wants a human author making the choices and how he edits the prose.
[52:11] Joel Borgan: But as we move into a world when... Where the, the AI systems can do a lot of what we do, a lot of what we thought was valuable, thought it was like a human contribution. You think about the abilities it has now. Like, the reason why I found the book valuable to write is that it was a lot of me in it. It it took a lot of effort on my part still. I know you can get... I I think I've I've heard a podcast episode where there was a... Somebody writing books on Amazon in the romance area, and they were mostly automated and AI written. And I think there were something like a couple 100 per year. So they were playing kind of a scale game. You know, a few of them would make a few dollars, and overall, it was profitable. And you can certainly do that, and the models are getting so much better that you can probably have customized artwork of the same quality or better than the book that I created over many, many, many months on demand. And then what happens to the art at that point? Thinking of it as an art consumer, if there was a new a new movie by Ingmar Bergman or Stanley Kubrick or something in movies that I've spent many hours thinking about and consuming and are deeply moving to me. Like, what would that world look like? Well, there's just a huge plethora of art on demand, and it's pretty much custom just to you. You know, we already have a shared dropping off of of a a cultural space. You know, people are are more fragmented in what they consume, and there's less unification, I guess, across the across the cultural landscape. So, like, is that art still valuable at that point if it's if it's not something that we're sharing with other humans? I I don't know the answer to that, but I think we're in a kind of a sweet spot now where in order to get, you know, a work that you're proud of working with AI, it still requires a great deal of human judgment and work. And I don't know how long that will be the case given kind of where things are headed. I could envision taking some of the the scaffolding I've I've I've... So I I had the... I had my my book edit. I edited it down about 14% or so mostly manually and through AI helping me deciding what to cut. Often, can take a novel... A chapter that's so so and actually just take away some of the the stuff that AI does particularly badly, the text, the overexplanation, and improve it in that regard. So I I think you could you could architect something now that that could create a pretty decent book with having, like, the right kind of the kind of feedback loops. A lot of the the stuff that I've developed is to to take my own preferences where I want it to be be able to give the entire novel to a model and have it come back with, like, an analysis and a list of, like, these are things you should consider. Because doing it manually slowly is extremely time consuming, as you can imagine. So bringing things to your attention like an assistant and saying, like, here's something you would consider, and then you can decide for yourself kind of point by point, I find is a fairly satisfying way to work.
[55:22] Nathan Labenz: Joel is also a musician. To describe a capability he had watched arrive, he turned from prose to musical notation and a viola clef test he had kept trying on successive models.
[55:35] Joel Borgan: I've had this held out test for well over a year now, maybe two years even. Very simple. A simple, like, high definition capture of five notes on the viola. And I give it to every... Give it to the the reasoning models. I'm like, okay. What is the clef? What are the notes? What are the the note values and so forth, the time signature? And not a single one of them got it right until Astra OpenAI put in some... Presumably put in some actual musical scores. And so it went from basically not being able to do anything to... Now I gave it a a complex, you know, twentieth century score that's almost painful to look at given how complex it is. I mean, this, you know, double sharps and accidentals and ties everywhere. Is... It's it's, like, difficult. And it it did an almost perfect job turning it into a machine readable form via codex. So the score that, like, you just basically had a PDF of before, now you can turn it into something that you can manipulate and evaluate. And it sort of went from kind of zero to 99 overnight. And I think that is an interesting way to turn some of the the scores that aren't machine readable into something that you could use to train a system. So I would love to see Suno move in the direction of adding a lot of classical stuff to it. Because I think in addition to it making interest in classical music... And classical music will have... It it requires a different a different level of understanding of the form because it's much larger scale often. You know, like, a three hour long opera that has some internal structure or even a a twenty or thirty minute single movement or something that has a lot more going on than a pop song. So I would love to see see that happen. As far... Again, it it does tie into the same element of... If you're making just kind of private art just mainly for yourself to listen to. I still think it's valuable. But, again, like like with writing a book, I would want to do it in tandem with the model, this hypothetical model in the future, and create something that's, you know, a large part of my efforts and taste goes into it as well. I don't know how long that that will last where you have a role for humans and AI to create something together. We use it as a tool, but it still requires a lot of you and your judgment. But I look forward to experimenting with that once the the models improve.
[58:11] Nathan Labenz: Part four, the diagnosis still waiting. Daniel McKinnon founded Gamo Labs to use AI agents to interpret genomes and revisit difficult unresolved cases. His son Owen died from a rare lung disease. A human specialist later found the genetic deletion the original sequencing analysis had missed. Daniel subsequently built an AI pipeline that rediscovered it while his family was seeking answers during another pregnancy. That experience helped lead him into this work.
[58:45] Ed Harris: And it really leapt out at me last summer now when... I mean, we we blessedly have a healthy 13 old right now. But because of our history, we are being monitored very, very carefully. And there was just something a little bit questionable in the sixteen week anatomy scan that our own maternal fetal medicine doctor, who is an absolute wonderful person, said, normally, I wouldn't even flag this, but because it's you guys, we need to look carefully. And we did a whole genome of the fetus at that point. It came back negative. This is actually the third negative whole genome I've seen in between losing Owen and having our son Warren. We actually lost a second pregnancy due to genetic reasons very late. So, you know, we've had basically five years of heartbreak before bringing Warren into this world, and all of it kind of genetic mystery related. And so just as a patient, I've started to learn... Or a family of a patient. I start to learn a lot about, like, the failings of this system. And it was that moment where I said, I'm heartbroken, but I'm mad, and I'm gonna do something. And I have tools to do something. Fast forward to last summer, you know, I got this result back. I said, this is unacceptable. This is the number one thing that I want in my life. I really wanted to have a family. I wanted to understand what would happen. I... You know, I didn't believe these labs. And I basically called all the labs, and I said, give me my raw data, which you can get due to our HIPAA rights here. And I, you know, vibe coded my own interpretation pipeline, which... I mean, now vibe coding is, like, crazy. Right? Like, right now, this would be so easy. You could just say, Cloud Code, you know, make me this thing. But back then, I mean, it was still, like, kind of, like, autocomplete original codex. Like, it did actually require quite a bit of work on my end. And I just did it to see if there's any comfort that we could get around the pregnancy, and I was just shocked that it also outperformed on these other cases. And, you know, most clearly, is it diagnosed Owen when, you know, one of the best or, you know, some might even say the best prenatal sequencing lab did not. And I knew at that moment that this is something I had to contribute to. I didn't know it would be a company.
[1:01:05] Nathan Labenz: The deletion in Owen was 91 kilobases long and removed an enhancer, a region that helps regulate a gene. The specialist had already identified it before Daniel recovered it with AI. We asked Daniel how the original analysis missed that enhancer.
[1:01:22] Steve Ho: How
[1:01:22] Ed Harris: you miss it is this enhancer is a megabase 1,000,000 bases up upstream from the gene. And what Rady did, what many other labs do, which is a very reasonable assessment, is they have a filter that for structural variants, which tend to be quite messy in what's called next gen sequencing, which is how sequencing is done today where you have basically roughly a 150 base pair fragments, and you need to piece them into this clinical puzzle. They said, we're gonna have a filter. And anything more than one kilobase up or downstream of the gene, we're not going to consider. And this is a totally reasonable trade off if you have humans looking at all of this stuff. But my moment was... And this is before Codecs, before Cloud Code, was can you put the o three model, which was the model I used at that point, kind of first, like, real agentic model? There was o one, but o three was, you know, really first breakthrough. Can you put it in some kind of loop and just have it keep looking? And it basically looped, and the first loop was like, is there anything wrong with the coding elements, which are kinda obvious things? And it's using these kind of bioinformatics tools. It's very crude compared to what we had today. But then it misses something, and it goes through another loop and another loop, and it just keeps working. And when you are a clinical lab, whether you're a profit, nonprofit, like, whatever your structure, ultimately, you gotta move. You know? You have to spend some amount of time on each case. And if you don't come to a conclusion, you know, you say this is nondiagnostic. And this is very common. Most... All genomes, even with infants suspected to have genetic disorders, come back nondiagnostic. And I really think it's one of these, like, meat space problems is if we can export these problems onto a machine intelligence that can work nonstop and in parallel, then we will be able to see many more kids, treat many more kids, and do much more kind of interesting analysis on top of the basic things that human are just pressed for time. And I don't wanna claim I've, like, reinvented this. Like, people are trying to build software to accelerate genomic interpretation. People have been doing this for a long time. But I think what I probably relatively uniquely identified early was that this is a really great task for agentic AI. And from, you know, improving on it, it's also... It's long time horizon. It's agentic. It's verifiable, and it's unsaturated. So I suspect that this will be a task like math where we can just generate these, like, you know, Olympiad level problems and just keep hill climbing on that until this problem is, like, basically solved.
[1:04:02] Nathan Labenz: Daniel McKinnon then turned to the unresolved cases his team is analyzing now.
[1:04:06] Ed Harris: Sometimes
[1:04:07] Nathan Labenz: the remaining obstacle is a variant whose effect is unknown. We asked where the new biological evidence would come from.
[1:04:16] Ed Harris: And on the backside, once we've done the interpretation, the majority of clinical cases we see right now... And to be clear, we are only seeing hard cases. So this isn't, like, you know, most generic labs. We are seeing cases that need reanalysis that did not previously have an answer. The majority of them end up in what's called a US state. So it's a variant of uncertain significance, and these are scored according to, like, a very standard rubric developed by the American College of Medical Genesis or ACMG. And you need to get six points or more to be bumped into the likely pathogenic category. And that unlocks a lot of treatment options, insurance options. Like, just... Let's say it's, like, good to be either benign or likely pathogenic. It's very bad to be in this kind of intermediate stage. And if you are in this intermediate stage in a VUS, you can do two things to get a diagnosis. One is you just wait, and you can get more points if more patients emerge who have a similar phenotype or a similar condition as you and the same genetic variant. So this is often what happens when, like, you read in the news. You're like, oh, this kid has, like, had epilepsy for 10. It finally got a diagnosis. And if they're lucky, oh, and since we know that's a diagnosis, you know, we work with pharmaceutical company, and there's, like, some off label use of a drug, and it actually can, like, help their condition. That's typically because they just waited until somebody else had the condition, but this is not scalable and takes a long time. Another fork is you can convince a university lab to kind of care about this problem. And what they'll do is they'll do what's called a functional study. So they'll make a cell line that's, like, emblematic of your particular condition. So, for example, for alveolar capillary dysplasia, that commonly used cell line is called IMR nineties. It's a fetal lung cell line. And this is very, very important because this particular gene is only expressed from week sixteen to week twenty of development. Like, you can't just put them in a cancer cell line. And then you can go in, and you can do something. Like, you can edit the genome, or you can insert a plasmid, or you can do a number of different functional studies to say, okay. In this physiologically relevant cell line, if I have this genetic mutation, is this gene broken in some way?
[1:06:33] Nathan Labenz: Daniel's team had just taken over an existing biology lab to test uncertain variants. He described the experiments starting that day and the feedback loop he hopes to build between real biology and the models making predictions.
[1:06:47] Ed Harris: So we basically bought what was remaining of Arpeggio, hired four people on the team, and very quickly pivoted to creating this basically, like, RL data from real world with biology experiments. And we're three weeks in. We're running our first experiments actually today, and I'm really looking forward to being able to close that feedback loop. And I think a lot of clinical genetics labs are as well as this service is not commercially available. We can not only sell kind of, like, a lookup table version of this. Like, oh, I have a patient with this variant. Can you help me get a couple more ACMG points to get this up to likely path? Because you can get two to four points. Remember, you only need six points, so two to four points is a lot of points for a functional study. But, also, it lets the machines learn. Right now, the machines aren't learning. Like, you get to a VUS. It's like, that's the end. Now it's like, you get to a VUS. You do the biology study. You say, oh, this region, the genome, is actually quite important for this particular disease. We're gonna do that. We're gonna check these predictions, and then we can improve these predictions over time.
[1:07:52] Nathan Labenz: We had been discussing earlier genetic testing companies, including counsel. Daniel returned to the gap between sequencing a genome and interpreting it. The scores he cites are from Rarebench, a benchmark his company built to prioritize genetic variants.
[1:08:07] Ed Harris: I was like, oh my god. We sequenced the genome twenty five years ago, and, you know, Bill Clinton and Tony Blair got up on stage and said this is gonna revolutionize human medicine. And I look back, and there's this, like, graveyard of genomics companies. Like, you know, you mentioned Council, which I actually would not say is a graveyard. I think they were, like, you know, modestly successful. Successful. But, But, like, like, is all of health care not based on precision medicine? Like, everyone in this space is like, it should be, and there are examples. It's just too challenging. The interpretation is too challenging. So I actually look back and I see... And, you know, we kinda talked about counsel ahead of time, is I think that what was missing prior to now was really that interpretation layer. It's very, very complex. Like, you know, genome sequencing has gotten what I would call cheap enough maybe five years ago. I mean, below 1,000 a genomes, maybe 500 a genome. There are high throughput labs that are doing this for less than 100 a genome. You you see announcements on Twitter saying, oh, I can do a genome for less than $100. There's many people doing this right now. Like, the sequencing is not the problem. It is given your problem, what insights can you derive from that? And that was only possible as of, like, last summer. Like, I'd say o three is the first example of that. And it's not just the variant interpretation. Right? It's explaining to the provider why it's important. It's ranking variants in a nice way. It's explaining, you know, how can you treat this person. It is ranking different patients for ASO eligibility. There are many, many, many other things beyond just scoring the variants. Although I will say we we benchmark a very annotation in Rarebench and other benchmarks we have, and the best performing, like, I'd say, machine learning based approach in terms of variant ranking is called Lyrica. I think it scores something like 10% on our benchmark. And right now, Opus 5.5 in Cloud Code, just a vanilla thing, scores something like 50%. So there's... These traditional tools are also getting blown out of the water by, you know, this this newer approach, but then they can do much more. And you basically need to make a very, very cheap, easy thing that people who are not surrounded by fancy clinical geneticists, genetic counselors, you know, top tier hospitals can use, and that's the problem that we're trying to solve.
[1:10:28] Nathan Labenz: Daniel distinguished screening a healthy person for possible future problems from investigating an illness that is already present. His team focuses on diagnostics and reanalysis of cases that have gone unresolved.
[1:10:42] Ed Harris: There's already product market fit for sequencing. Every baby in top NICUs is getting sequenced. Every baby in lower tier NICUs would get sequenced if they had the resources. And the problem is much better scoped. It's baby has pulmonary hypertension. Figure out why. Not is there something obscure that could be wrong with this baby now or in the future? And that is where we are working with, you know, various hospitals, various families today. This is... There's two forks of this work. One is what I call online, and this is, like, case comes in. Physicians do not get good results from the traditional labs. Patient consents for a reanalysis. They send the data to us. We reanalyze it. And sometimes, but not always, we find, you know, additional things that can help them make decisions around either pregnancy or a newborn. The second fork is bulk scale reanalysis. So this is rare disease centers that have thousands of genomes, and they say, there's probably kids that we can diagnose or even treat in that database, but we have a handful of genetic counselors, a handful of bioinformaticians, and we just don't have the resources. And because it's a machine that does most of the work here, you know, we can get a kit... You know, we can get a set of 80 or a 100 or 200 or whatever and turn them around in a week. And that's not... This is kind of six months or a year's work that they really can't prioritize. So those are the two forks right now. And inter... I mean, we're four months in. We're prerevenue. We're really not thinking about what is the, like, exactly appropriate business model here. We're just trying to diagnose more more sick kids.
[1:12:21] Nathan Labenz: For those cases, Daniel described what his team adds around the underlying models, the tools, the routing between models, and the evaluations of the whole system.
[1:12:31] Ed Harris: And so then, you know, know, stepping one step up, it's like, what do we do? I mean, we're basically a a harness, a tools, and an eval's company. And I think that you'll see this a lot in vertical AI in general. You know, I've I've worked in these frontier labs. I started working on LMs. Actually, the first LM project at Meta, which was OPT one seventy five. And unless you are somebody who is very famous with very deep pockets, I don't wanna bet against the NeoLabs by any stretch. I think it's possible. You know, you are not gonna compete at the model layer, and the intelligence is progressing so fast. You know, we ensemble models for sure, and that helps with both cost, and it helps with performance. If you look at the confusion matrix of Rarebench, we also publish this, you'll see that the cases kind of cluster. And, like, Claude is good at this, and Astra is good at this, and even Gemini. Right? Gemini is not on the frontier, but they're kinda good at this. And you can kinda ensemble these things together and get better performance. So I'd say that's, like, one thing we do. And ensembling, there's, like, very dumb ways of doing this, but, like, there's also smart ways of doing this. And I don't have to go into all the details here, but, like, I would say figuring out, you know, unique routing intelligence per model is an important edge that I think we and many others are doing. And then I think another key thing that people kinda forget about, and this is what I worked on both at Meta and at Google, is Evals. Is unless you have very structured ways to measure the performance of the system, it's not just the model at this point. It's the whole system. You don't know it's a hill climb. So we do, like, pretty robust, like, trace analysis of how models perform with and without our harnesses. And it's it's... I mean, this sounds stupid, but, like, not that many people actually look that closely at data. Like, this is a meme on Twitter, but it's very true. And you'll see things. Like, there's one particular case where Grok 4.6, which at that point was state of the art on our benchmark in Grok build, where it missed one because it just renamed the gene. It just, like, was saying, like, oh, you know, n r f two or whatever is responsible, and then it just changed the name of the gene to something totally different. I'd never seen that before. And I was just like, this is dumb. And our harness and our tools prevent the agents from doing dumb things. I think there is a world where if I worked at OpenAI or Anthropic and had access to infinite compute and I could do millions of rollouts that the agents would, with enough compute, kind of, like, all converge on, like, some answer and, like, maybe mitigate all of this work, especially because they can write their own tools these days and everything. But, really, it's, like, consistently and repeatedly and, you know, honestly, affordably getting to the right answer. I mean, for $10 a case, like, who cares or even a $100 a case. But if you would need to do a million rollouts and all of a sudden this is, you know, a thousand dollars a case or $10,000 a case doesn't make sense. So we're doing a lot more research around, like, how those, you know, cost curves work and how everything works. But, ultimately, our job is to take the smartest intelligences in the world and mix them together and give them access to the best tools and then get answers for our patients. And that's, like, the core of what we do. And and and be able to measure if it's working. So I'd say that's, like, the core of what we do. So would you say... Like, when you talk about
[1:15:46] Prakash: vertical AI, you are looking at benchmarks where you combine an intelligence with the tools that you have. How do you evaluate the strength of the tools on their own, you know, excluding ex... Because you... You're gonna upgrade model families over time. Right? How do you how do you measure the performance of what you're building that that layer in between in particular?
[1:16:09] Ed Harris: Yeah. That's a good question. And I don't... I honestly don't have, like, a great answer to that question in that we we are codesigning the tools with the models. And we have seen cases where... I mean, this is one... Like, again, it's a very dumb example. And if you are getting deep into this and you have a specific eval, you will start to see this. Like, if you're not deep, it's very easy to be like, oh, I just had Cloud Code, like, one shot this video game. Was it good? Like, I don't really know. Like, it seems amazing. And I don't wanna communicate that these models aren't amazing. They absolutely are. But with one tool, it was around a specific type of literature search. And the model... I think this was Astra, actually. Actually, okay. Here's a good example for Astra after this. But Astra said, oh, yeah. Okay. These are all the papers you tagged for me to read. And then it got... The the answer was in the paper, and it didn't... It missed it. And then we're, like, looking through the trace, we're, like, looking through the context window. It's like, I don't think you actually read these papers. Like, there's lots of just, like, little things like that, and it's little paper cuts. And that's what these tools do is it's like, oh, you miss a case here because of this. You miss a case here because of this, and you kind of put it on these better rails. And, like, actually, Astra is a good example. The first time we ran the eval for Astra, it just stopped because some of the tools we had, it just, like, wanted to think about too much. It's almost one of these things. Like, you see this on X where people are like, oh, I asked I asked the model to do some research for me, and I came back. And thirty minutes later, it's like looking at flights from, you know, Dubai to Calcutta on, you know, on, you know, Emirates airline. Like, why are you looking? And they just kinda do stuff like this. So I think that, you know, we try to co evolve the tools with the models, and we try... You know, we have, like... You know, coming back to something that many people might find boring, but I find fascinating is, you know, evaluating a model is very hard, and we need to just have very robust evals that can catch these things early. And in that case, we caught it early, and we updated how Astra called some of the tools, and it works again.
[1:18:09] Nathan Labenz: Prakash asked Daniel what he most wants the Frontier Labs to improve.
[1:18:14] Ed Harris: Well, I want them to hill climb my task. Right? Everyone wins when these models get really, really good at clinical genetics is probably we have less to do at the harness layer, which, to be honest, I'm, like, both bullish and bearish vertical AI in that I think that there's, like, lots of value in kind of owning a customer relationship. I think open evidence has shown this, but there's also... Like, that layer is getting thinner. When I first started this, like, vanilla o three would not do this task. In fact, o three scores 0% on my benchmark. I rebenchmarked it just for fun. So it was like, oh, I needed to do all this work and scaffolding to get these things to work. And now, like, you know, there's, like, your twenty, thirty percentage points or something. I mean, it's definitely, like, shrinking. So... But that said, like, my my goal is to build, you know, this, like, great AI native clinical diagnostics company, and I want the best person to do the work.
[1:19:08] Nathan Labenz: Prakash also asked how Daniel hopes to make those biology experiments cheap enough to run at scale.
[1:19:14] Ed Harris: But it's it's really an AI and robotics thing is you have an arm there that is doing 380 wire, four experiments at a time in, you know, a bigger plate. You have a cartridge over there with thousands of plates lined up. You have AI that's assisting with primer design and experimental design. And, actually, this is one of these things where you say, like, oh, you know, can you use AI for this stuff? You have to get special permission, but you can. And then, you know, at very high throughput, you can design, order the primaries, design the experiments, and run the experiments. And then it comes to analysis as well as you end up with, like, very large scale datasets that you need AI to poke through. So it's not... I don't have some, like, perfect explanation. Like, there's just this one invention, but it's a confluence of all those things.
[1:19:59] Nathan Labenz: Daniel also described the response to his diagnostic work from people at the frontier model companies. Here, he reports new diagnoses in other children whose cases had been unresolved. What was the process of talking to frontier model companies about your bio use case experience like?
[1:20:21] Ed Harris: It was really positive for me because it's extremely obvious why I'm doing this, and I've never met anyone who is like, why are you doing this? Every single person I've met at any of these labs has been like, I wanna help. Whether wanting to help as a person translates into, like, I can pivot this trillion dollar company to work on your problem is, like, a different thing, and I understand that. I've worked at these trillion dollar companies too. But it's been incredibly supportive, and I would be pretty surprised if at least the top frontier labs did not spend some amount of time on this problem because, again, it's meaningful and it's good. And, you know, even selfishly for them, there's a really negative AI narrative swirling around right now. And you hear things like in Dario's recent tweet where he's like, oh, we're gonna cure cancer in five years. We're diagnosing kids today. We have many examples of kids who are undiagnosed who we have diagnosed at a small company four months in. And that's a great story, and it should be told by them even if, you know, even if it's just for selfish reasons. Like, hey. People of The United States Of America who are not happy with data center build outs or uncomfortable with AI taking jobs, here's a very concrete way where we are helping people today. And then I also think it comes from the individuals. It's like people find meaning in their work, and I think that there are a lot of people, especially, like, mid career in tech who are like, I've been doing some kind of ads ranking, data munching, whatever for my whole career. And what you're telling me is that you have a very concrete way where you can help kids, and I wanna do that. And so I think that there's, like, the combination of, like, the business reasons and the personal reasons that are driving people to this task.
[1:22:14] Nathan Labenz: After that discussion with Daniel, I came back to what slower AI progress could mean for families still waiting for answers. The 50% figure I refer to is a score on Rarebench, the company's variant prioritization benchmark rather than the share of patients diagnosed. Intellectual honesty demands, recognition that this is probably still one of the things that we will be
[1:22:41] Ed Harris: trading
[1:22:42] Nathan Labenz: off when we pace the frontier in the near term to whatever extent that actually happens. And I still think that's probably a good idea and probably worth it. I don't think the frontier model companies are at risk of, like, not being able to grow business, not being able to grow revenue, becoming surpassed if they pace their frontier efforts. But it is important to keep in mind that there are very real problems in the world that AI is, like, climbing that hill on now but has not finished climbing on. And this is one, you know, obviously, that everybody would love to see solved sooner rather than later. And, it's important for me to stay honest with myself and, everybody else that, like, that is a, you know, very real cost when you're talking individual people and their families, and it's that 50% number that Claude gets to today. You know, that's, like, obviously leaves half of the questions unanswered, and those questions are extremely meaningful to people. So I I do think we shouldn't take that lightly even if it is a a trade off that, on balance, I think, probably, we should be willing to make some compromises on. It's it's a it's a costly compromise. Thank you for listening. Please tell us what worked and what did not. See you in the morning.
Outro
[1:27:18] If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries either via our website, cognitiverevolution.ai or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of a sixteen z where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the cognitive revolution.