Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics

Nathan reports from a two-week trip to China, examining the common but China objection to U.S. AI safety policy. He compares Chinese and American safeguards, regulation, research institutions, and debates over open weights and service-level risk.

Nathan Goes to China – Part 2: AI Safety with Chinese Characteristics

Watch Episode Here


Listen to Episode Here


Show Notes

AI Safety with Chinese Characteristics

There is a move that ends about half of all AI policy arguments in America. Someone proposes a duty, an obligation, a pause, a standard — and someone else says: but China. China doesn't care about AI safety. China will never slow down. Anything we do is a gift to Beijing. Nathan calls this the "but China" endpoint, and this episode is a two-and-a-half-hour attempt to dismantle it with evidence rather than sentiment. He spent two weeks in China in July 2026, attended WAIC in Shanghai, sat in on an AI safety hub launch in Beijing, and had many conversations with researchers, company leaders, and ordinary residents. The whole episode is delivered under Chatham House rules — no names, no institutions, no attribution to anyone he met — with one exception he states on air: published papers, reports, and public speeches are fair game, and those are all linked here. He also notes, unprompted, that he paid his own way; the only hospitality he accepted was about half a dozen meals.

He leads with the part that cuts against his own thesis. As of today, Chinese models and services do not have safeguards as strong as the leading American ones. That much is real, and he states it plainly. But the shape of the gap is not what the headline numbers suggest. The American average is being carried, almost single-handedly, by two companies — OpenAI and Anthropic, which he half-jokingly proposes merging into "OpenAnthropic," the duopoly leading on both capability and safety. Subtract those two and compare the rest of the American field to the Chinese field, and most of the reported differential evaporates. Two independent sources on opposite sides of the world agree on the ordering: Concordia AI's airiskmonitor.net, which plots capability scores against safety scores for 47 frontier models across cyber, bio, chem, manipulation and loss-of-control risk, and Adam Gleave of FAR.AI, whose jailbreak-robustness work lands in the same place — OpenAI and Anthropic hardest to break, Gemini and Grok easier, Chinese models easier still. Concordia is based in China and publishes this in Chinese and English, on the open internet, without apparent difficulty. That fact is itself part of the argument.

The organizing idea in Chinese AI safety is the "45-degree line," introduced by Zhou Bowen, director and chief scientist of the Shanghai AI Lab, in a 2024 WAIC speech: capability and safety should rise together, slope of one. Nathan finds it a genuinely useful frame — and finds China visibly below the line today, which the Chinese ecosystem does not much dispute. He also surfaces a real conceptual disagreement rather than papering over it. American safety discourse asks what an open-weights model can do in the worst case, in anyone's hands. Chinese regulation attaches to the service, not the model, and the assumption behind that is partly empirical: these are trillion-parameter systems, nobody in psychological crisis is spinning one up on a laptop, so what matters is the business that wraps it. Nathan thinks both views are legitimately right about something, and that the American discourse blows past the Chinese point a bit too fast.

The deepest structural difference is not technical. China does not have the AI safety nonprofit sector the United States has — the permissionless one, where a few weird people convince one wealthy patron and go chase an unpopular idea for a decade. Nathan is emphatic that this is an American strength worth not taking for granted; it is where essentially the entire Western AI safety field came from. In China, the analogous work lives in the universities. That means the field moved into academia faster than American academia moved, which is to their credit, and it also means the people doing it are institutional people on institutional career tracks. Go looking for your counterparts and you may not find them. In one of the episode's better moments, Nathan asks a professor who has published extensively on AI safety whether people trade p(doom) numbers over lunch and whether they read AI 2027. The answer was no on both counts — AI 2027 is "too political" — followed by a genuinely curious "do Americans do that?"

Rhetorically, though, the signal comes from the top, and it is stronger than most American listeners will expect. Nathan reads three passages from Xi Jinping's WAIC opening keynote, sourced from Matt Sheehan's annotated full translation: on how humans should coexist with machines that think; on how "the more rapidly safeguards against loss of control must improve"; and on building legal, monitoring, early-warning and emergency-response systems to "keep AI under human control." He engages the cynical reading — that "loss of control" is just social control by another name — and argues it is incomplete rather than wrong. Then he offers the comparison that does the real work: set that speech next to the most safety-forward remarks from any prominent American politician, or next to what J.D. Vance said in Europe, and Xi comes off looking like the hawk.

The research picture is where the convergence becomes almost eerie. A couple of days before WAIC, Nathan flew to Beijing for the launch of an AI safety hub at Tsinghua University's College of AI — an event with, as far as he could find and as our own searching confirms, no English-language coverage whatsoever; the only source is Tsinghua's own Chinese-language announcement. The hub's founding board includes a European professor relocating to Beijing. Speakers repeatedly named Constellation and LISA in London as explicitly what they aspire to be, with Singapore's SASH also mentioned; they counted roughly sixteen such hubs worldwide and want to join the top tier, sending students out and bringing international residents in. The research presented that day carried the logos of Apollo Research, METR, Palisade Research, Redwood Research and the UK AI Security Institute as prior work and inspiration. Nathan's counterpoint is sharp: a Chinese big tech company runs an agent that reads American AI safety discourse daily and files a report on it, and meanwhile no American outlet covered a safety institute opening at China's premier university.

Volume backs up the anecdote. Concordia's aisafetychina.com database shows Chinese AI safety output going from a couple of papers a month in 2023 to 50–60 a month by mid-2026 — a more-than-tenfold rise, against a US/Anglosphere figure Nathan estimates at 50 to a few hundred. Multiples apart, not orders of magnitude. And the papers themselves read as near one-to-one analogs of Western work: Frontier AI Systems Have Surpassed the Self-Replicating Red Line (2024) against Palisade's self-replication research; Evaluation Faking (2025) against eval-awareness work out of Anthropic; R²AI, whose supervising last author is Zhou Bowen himself and which cites the Guaranteed Safe AI agenda — led by davidad — in its opening paragraph; DeceptionBench against Apollo's science of deception; and a robotics-flavored pair, When Alignment Fails and its six-months-later mitigation follow-up STRONG-VLA. Interpretability is rising too — Mechanistic Origin of Moral Indifference in Language Models opens on the gap between surface compliance and internal unaligned representations, which Nathan calls a total LessWrong cultural victory, and SafeSeek goes after the generalization of safety-circuit attribution. The most uncanny item is one he could not find online at all, presented at a WAIC-adjacent event under the title "Toward Decoupling Capability Growth from Risk Growth: Isolating Hazardous Capabilities in Mixture of Experts," promising on the slide that harmful experts could be switched off or removed at inference. That is, essentially, GRAM — the gradient-routing work from AE Studio with Anthropic (plain-language writeup here, building on the 2024 Gradient Routing paper) that Nathan has been enthusiastic about for months, arriving independently on the other side of the world.

On governance the comparison flips: the Chinese government is unambiguously doing more, whatever you think of the doing. The Cyberspace Administration of China is the agency companies actually deal with — weekly, sometimes daily — and it maintains a registry of reviewed and approved AI services, with a provincial review preceding the national one. The precedent Nathan keeps returning to is that this government has already slowed its own AI industry: for roughly six months in 2023, Chinese companies with ChatGPT-moment models were simply not allowed to launch while standards were written. The broader pattern runs through the social media era — rules on recommendation algorithms since 2022, protections for gig delivery workers, anti-scam rules aimed at protecting the elderly, limits on price discrimination, AI-content labeling — and, while he was in the country, a new AI companionship regime with anti-addiction measures and an outright ban on minors. Add a Politburo study session where Xi discussed technological loss of control, a draft cybercrime law that would require monitoring for bulk generation of malicious code, a security warning issued about OpenClaw within weeks of the craze, and an emerging labor-market posture that reportedly extends to forbidding layoffs justified by AI. Nathan's honest read on the last one is that they probably can't hold it — but that the willingness to try is the point.

The one place he thinks Chinese thinking has a genuine blind spot is the externality. Officials and researchers seem comfortable releasing open weights partly because, inside their own borders, they believe they can put the genie back in the bottle: pressure the cloud providers, delist the model, and — he thinks plausibly — find an unregistered inference operation by its electricity draw. "We think in the West, 'oh, you can never take it back — the internet never forgets.' Well, the Chinese internet does forget." That confidence does not travel. No other government has that capacity, and a weights release that is recoverable in China is permanent everywhere else, which matters most for bio risk that could eventually blow back on them. He also treats the unfulfilled Seoul commitments — Chinese signatories who never published frontier risk frameworks — as a real failure, while noting the mirror: Anthropic's own Responsible Scaling Policy shed its hard if-then commitments for what Zvi memorably called a "trust us" regime. His best guess at the Chinese explanation is a division of labor that has settled in and that both sides regard as legitimate: standard-setting is the government's job, and it would be something like overstepping for a company to publish its own.

He closes on a gap he went looking for and did not find. If this is AI safety with Chinese characteristics, is there such a thing as alignment with Chinese characteristics? The prompt came from using a Chinese AI at the Temple of Confucius in Beijing and learning that Confucius's descendants, 79 generations on, still identify as such and still perform rituals in his honor. If we could project our values through 79 recursively self-improved generations of AI, we would be doing extraordinarily well. Could there be a Confucian constitution the way there is Claude's Constitution — an alignment target built on a wisdom tradition an AI grows into, rather than a rule set it complies with? He asked repeatedly and got nothing. One professor's answer: we are probably the generation in all of Chinese history weakest on the traditional philosophy, we are all engineers, none of us studied this. The Chinese AI safety community sits firmly on the corrigibility side of the corrigibility-versus-character debate — clear rules, reliable compliance. Nathan thinks the other path is wide open, and that in a future of multiple powerful AIs held in some ecological balance, having them rooted in different wisdom traditions might be a very good idea indeed. Right now it is greenfield. Part 3 turns to the US-China relationship itself.

Topics covered

Timestamps are approximate — see the note at the end of this page.

  • 0:00 — What Part 2 is, and why it's separate from Part 1
  • 1:52 — Chatham House rules, and a disclosure: he paid his own way
  • 3:32 — The "but China" endpoint of every AI safety argument
  • 6:20 — Leading with the uncomfortable part: where the two ecosystems actually stand
  • 8:00 — "OpenAnthropic" is carrying the American average
  • 10:49 — The intellectual history of Chinese AI safety: pragmatic, not sci-fi
  • 12:10 — Zhou Bowen and the 45-degree line
  • 13:44 — Concordia AI, airiskmonitor.net, and where the models actually plot
  • 16:30 — Model vs. service: the monitor that sits on top and cuts you off
  • 18:58 — China regulates services, not models — and why that's a real disagreement
  • 22:05 — Do the companies care? Five of ten publish safety evaluations
  • 23:37 — Big tech versus the startups, and who has more to lose
  • 25:05 — The Meta/Llama analogy: "it's already been out there for a year"
  • 34:42 — Past the three Ts: everyone in China is talking about agents
  • 37:38 — China has no AI safety nonprofit sector. Academia is the substitute
  • 42:35 — Who you actually meet: professors, not LessWrong
  • 43:35 — A professor asks: "Do Americans really trade p(doom) at lunch?"
  • 46:59 — Xi's WAIC keynote, in three quotes
  • 51:39 — What does "loss of control" actually mean in that speech?
  • 54:44 — Beijing: the AI safety hub launch at Tsinghua's College of AI
  • 58:30 — Apollo, METR, Palisade, Redwood and the UK AISI, name-checked on Chinese slides
  • 1:01:28 — The palm-on-the-podium launch ceremony, and being less too-cool-for-school
  • 1:04:27 — aisafetychina.com: from a trickle in 2023 to 50–60 papers a month
  • 1:07:23 — A tour of the papers: self-replication, evaluation faking, R²AI
  • 1:09:28 — DeceptionBench, and the vision-language-action robustness pair
  • 1:12:05 — Interpretability in China: moral indifference and safety circuits
  • 1:14:56 — The uncanny one: isolating hazardous experts in a mixture of experts
  • 1:16:28 — GRAM, gradient routing, and having the open-weights cake
  • 1:23:01 — Governance: the Chinese government is simply doing more
  • 1:25:36 — The Cyberspace Administration of China
  • 1:28:23 — Regulating platforms: gig workers, elder scams, price discrimination
  • 1:34:07 — AI companionship rules, Doubao, and anti-addiction measures
  • 1:38:30 — The registry, and how a service actually gets approved
  • 1:42:34 — 2023: six months of held-up launches
  • 1:47:35 — Weekly, sometimes daily contact between companies and regulators
  • 1:51:28 — Politburo study sessions, a draft cybercrime law, and OpenClaw
  • 1:55:34 — What would an OpenFace-scale incident look like in China?
  • 2:01:10 — Putting the genie back in the bottle — and the world outside the border
  • 2:05:33 — "You can't fire people because AI made them redundant"
  • 2:08:56 — The Seoul frontier-safety commitments that never arrived
  • 2:17:03 — Is there such a thing as alignment with Chinese characteristics?
  • 2:18:20 — 79 generations of Confucius's descendants
  • 2:22:44 — Closing: what to say the next time someone says "but China"

Resources

Concordia AI

  • Concordia AI
  • State of AI Safety in China (2026) — July 2026 edition, covering July 2025–June 2026, by Gabriel Wagner, Erik Lindblad, Kwan Yee Ng and Brian Tse
  • airiskmonitor.net — capability score vs. safety score for 47 frontier models across five risk domains, in Chinese and English
  • aisafetychina.com — interactive explorer over Concordia's database of Chinese frontier AI safety papers and research groups

The 45-degree line

Beijing, WAIC, and Xi's keynote

Papers named on-air

The Western work being echoed

Governance

  • Cyberspace Administration of China
  • ⚠️ CAC generative AI service filing system (beian.cac.gov.cn) — could not be verified. The fetch failed with a socket hang-up, apparently blocked to automated or non-China access, so we could not confirm its contents match the registry Nathan describes. CAC also publishes the filed-service list periodically as announcements on its main site (for example the list published January 2026), which is the safer link.

Quotes worth pulling

"If you were to subtract those two companies from the mix, and you were to look at the rest of the American companies and how they compare to the Chinese companies, honestly, a lot of the safety differential that gets reported would disappear, and I think the story would look a lot more muddled."
"They are following American AI safety discourse on a daily basis with — get ready — an agent that goes out and surveys American AI safety discourse on a daily basis and gives them a daily report of what is going on in AI safety in the United States. … Certainly you don't have too much of that going in the reverse."
"He also said, 'No, p(doom)? No, we're not trading p(doom) numbers at lunch.' He kind of looked at me and was like, 'Do Americans do that?'"
"If you compare this speech from Xi to what J.D. Vance said in Europe not that long ago — Xi comes off looking positively AI-safety-hawkish by comparison."
"Amazing to go all the way to Beijing to sit in this launch of this AI safety hub, at one of China's top, if not the top, university, and hear Constellation and LISA name-checked as, 'This is what we strive to be. These guys have done great, and we want to be like them.'"
"We're talking about the difference between surface compliance and internal unaligned representations that could lead to long-tail risks. That's a total LessWrong cultural victory, if I've ever heard one."
"Overall, I'd say — with a smile — the Chinese AI safety community, they're just like us."
"The careful way to say it is, if a human did what OpenAI's and Anthropic's models have reportedly done, I believe it would be a felony."
"We think in the West, 'Oh, you can never take it back — the internet never forgets.' Well, the Chinese internet does forget."
"If we could project our values through 79 recursively self-improved generations of AIs, we'd be doing really well."
"It's an interesting idea, but we are probably the generation in all of Chinese history that is the weakest on this traditional philosophy. … We're all engineers. None of us really studied philosophy."

Notes on sourcing

This episode was recorded under Chatham House rules. Nathan deliberately does not name any person or organization he met with, and these notes reproduce that anonymization exactly. Everything named and linked above is material he cited on the record as public: published papers, published reports, institutions, public speeches, and public figures.

⚠️ Timestamps are approximate. The source transcript for this episode shipped without timecodes, so the chapter marks above were derived by mapping position in the verbatim transcript onto the 144.2-minute runtime. They are reliable as an ordering and good to roughly a minute early in the episode, with more drift later. Re-derive them from real timecodes before using them as YouTube chapters.

Sponsor: Anthropic.

Sponsor:

Claude:

Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr

CHAPTERS:

(00:01) China safety framing

(05:37) Safeguards and benchmarks (Part 1)

(15:04) Sponsor: Claude

(16:33) Safeguards and benchmarks (Part 2)

(21:39) Company safety incentives

(28:17) Research cross pollination

(36:07) Nonprofits versus academia

(45:13) Xi safety rhetoric

(52:00) Tsinghua safety hub

(01:00:28) Research paper boom

(01:18:32) Governance track record

(01:31:37) LLM regulatory process

(01:42:57) Monitoring and enforcement

(01:57:45) Risk framework gaps

(02:05:08) Confucian alignment questions

(02:11:49) Episode Outro

(02:15:36) Outro

PRODUCED BY:

https://aipodcast.ing

SOCIAL LINKS:

Website: https://www.cognitiverevolution.ai

Twitter (Podcast): https://x.com/cogrev_podcast

Twitter (Nathan): https://x.com/labenz

LinkedIn: https://linkedin.com/in/nathanlabenz/

Youtube: https://youtube.com/@CognitiveRevolutionPodcast

Apple: https://podcasts.apple.com/de/podcast/the-cognitive-revolution-ai-builders-researchers-and/id1669813431

Spotify: https://open.spotify.com/show/6yHyok3M3BjqzR0VB5MSyk


Transcript

This transcript is automatically generated; we strive for accuracy, but errors in wording or speaker identification may occur. Please verify key details when needed.


Main Episode

[00:01] Hello, and welcome back to the cognitive revolution. This is gonna be part two of the Nathan goes to China series, and we're calling it AI safety with Chinese characteristics. Part one, if you haven't heard that, is up on the feed. It's been up for a few days. And in that part one of the series, I really just laid out the tech setup for going to China, what you would need to do if you wanna get a cell phone ready, the apps that you need to download. I also shared some of my experience using Chinese AIs on the ground there and then just shared a bunch of observations based on my two weeks in China and many conversations I had. I appreciate the kind comments that I've got in response to that one, including at least one from a listener in China, which was the probably the one that mattered to me the most. I think if you've been to China in the last few years, you can probably skip that one. But if you're interested in this, you might also be interested in that, though I do think they should be pretty self contained. So part one is a little bit closer to a travelogue and tech review, and this episode is gonna be really focused on the Chinese AI safety ecosystem and trying to just describe it and as much as possible understand it on its own terms. As we saw last time, there are a lot of similarities. Just as there are major similarities between American big tech and Chinese big tech, there are a lot of similarities between American AI safety and Chinese AI safety communities and the work that they're producing, but there are also some differences. And I think it will definitely be helpful if we have a better understanding of those. Before I really get into it, just a couple quick disclaimers again. I will be, again, following a Chatham House rule for this episode, not because I was asked to do that, but just because I wanna keep things simple for myself and make sure that I'm protecting everyone that I talk to from being misrepresented by me and put in an uncomfortable position. So I won't be naming any names or attributing things to the the organizations or institutions that people are affiliated with. The one exception to that for this particular episode will be when I'm citing published papers or reports, then those I can give you the name because those are out there in the public domain. And there are some good sources that I'll mention and we can link to in the show notes of this episode for anybody who wants to go deeper, which, of course, is always recommended. The other point of order or clarification, just in case anyone is wondering, I have no financial conflicts on this matter. I took this trip paying my own way, flights and hotels, all that stuff. The biggest thing that I did accept from some of my gracious hosts were a number of meals, probably half a dozen meals over the course of the two weeks. And I think if you have listened to me this far, you can probably feel pretty confident as I do that accepting those meals has not overly colored my take on what is going on in the Chinese AI safety community. So with that, let's get into it. I think so often, and this has faded, I think, in recent times as the situation arguably has just gotten a lot more real real quick, especially with things like the recent Open Face incident where it's starting to become undeniable that we have some real AI safety problems on our hands. But still, you do hear from time to time and you and used to hear it, I think, more that we could never possibly slow down our AI race. We could never regulate. We couldn't impose any any duties or or obligations on the frontier companies because why? China doesn't care about AI safety. China will never slow down. I used to call this the but China endpoint of so many AI safety and regulatory discussions. And I think if nothing else, I hope that this episode serves to really disabuse people of that misconception. I think it should become clear by the end of everything that I'm about to take you through that China does care about AI safety. Not it's not the only thing that they care about, certainly, but they do care. China has, at times, slowed down their AI companies in the name of safety. Not necessarily existential safety, but safety as they understand it, domains that they care about. And so I think that coming out of this, we should have clarity at least on that basic point. Right? That there is a there there, that people in China do really care about these things, and that the Chinese government is willing to take action when it deems it necessary. So that doesn't get us all the way through to an international treaty, obviously, and that's gonna be the subject of the third episode, an analysis of The US China relationship, obviously, through the lens of AI issues and what we might ought to try to do about bringing the relationship into a more productive state than it is today. Today, we're really just gonna focus on what is going on in China with respect to AI safety.

[05:05] But I think this this foundation is really critical, and, hopefully, this will become an artifact that people can share with those who, if they're open minded enough to listen to something, but have the misconception that China doesn't care or China will never slow down or whatever. We'll be handing the race to China if we do anything other than just race full speed ahead into recursive self improvement. I think that really, in plain terms, does come from a position of ignorance, and, hopefully, we can we can eliminate at least a portion of that over the course of this episode. Okay. To start off with, I wanted to just do a real quick factual rundown of, like, where the two ecosystems are today when it comes to the actual safeguards that they have in place on their frontier models. And I think this is important just because it is grounding and I and generally, someone who comes off as a China dove and, like, conciliatory and tries to find the positive in things in general, not just in China. And so people might accuse me of burying the lead or shying away from contradictory evidence. So I figured I would just lead with this upfront. I think it is fair to say that as of now, Chinese AI companies and models do not have as strong of safeguards protecting the public against the potential for misuse of the models as the American companies do have. So that, I think, is important to just state very plainly and be quite direct about. It's the difference is more than we might like to admit. The difference, I think, really is driven by two companies in The United States. I just did an episode not long ago with Adam Bleeve, which talked a lot about jailbreaks and robustness to jailbreaks and all that kind of stuff. And what we saw in that episode was OpenAI and Anthropic, which I'm maybe gonna start calling Open Anthropic as they become the the duopoly that is leading both in capabilities and fortunately, happily, also in terms of their safety measures. Those two companies are really doing a lot to bring up the American average. If you were to subtract those two companies from the mix and you were to look at the rest of the American companies and how they compare to the Chinese companies, honestly, a lot of the safety differential that gets reported would disappear, and I think the story would look a lot more muddled. I do think The US companies would still have a bit of an edge, but it would be a pretty slim edge, and it we it certainly would not give the American ecosystem some sort of obvious moral high ground from which to proclaim that the Chinese are not doing a good job. So that, I think, is, again, really important to understand. There are in each country, depending on where you wanna draw the line on frontier, if you draw a permissive line on what it means to be a frontier model maker, there's something like 10 to 12 in each country. And, that is quite permissive. Right? That's most close watchers of the American AI ecosystem would not count 12 companies as being at the true frontier. But if kind of say near frontier, there's like 10 to 12 in each country, and it really is the top two companies in The US that are doing a lot of work of bringing up the average, especially the weighted average in terms of what it is that people actually use on an ongoing day to day basis in their lives. And Google with Gemini is, like, not doing as much or as well to implement safeguards as OpenAI and Anthropic are, but they're also, like, ahead of the pack at least a little bit. And, again, kind of would bring up the average or if we didn't have OpenAI and Anthropic, we might try to hold Gemini up as a standard. It wouldn't be a standard that's, like, so much better than the Chinese companies, but it would be at least a little bit better. So, again, to state it plainly, on average, especially if you take a weighted average by what products people actually use, the American companies are ahead in terms of having better safety practices, better safeguards, also better funnel through on their model cards and their commitments to publishing safety evaluations. We'll get into that a little bit more later. But the American ecosystem really is being carried by a couple of leading companies. And beyond that, the mixes the relative positions are much less differentiated than we might like to think or we might look at if we just see the headline numbers where we are including OpenAI and Anthropic in those measurements. Okay. So just a little bit of kind of intellectual history of Chinese AI safety. Broadly speaking, I would say it is a pretty pragmatic bunch of people with pretty pragmatic ideas, pretty grounded, pretty normal seeming people. They're not like the crazy sci fi people. They don't they're not so influenced by sci fi tradition.

[10:05] They're really just looking to make things work in a kind of practical one generation to the next sort of way for the most part. A good source that I could point you to is the director and chief scientist of the Shanghai AI Lab. That's a major institution based in Shanghai, of course. His name is Zhao Boen. Again, I apologize to everyone and especially my Chinese listeners for being just terrible with Chinese. Joke about that over the last couple of days has been that, in general, I feel pretty young for my age, and I'm, like, pleased with how agile I feel like my mind is. But when it comes to getting Chinese pronunciations right, the neuroplasticity is low, and I'm struggling. So, anyway, I apologize for that. But Zhao Boen, he is the director and chief scientist of the Shanghai AI Lab. And at the twenty twenty four WAIC, that's the same Shanghai conference that I just attended, but obviously two years earlier, he introduced this concept of the 45 degree line. This speech is available. It's on government websites. You can go find it. But the idea of the 45 degree line is that capabilities and safety measures should grow together. Right? That's the the 45 is kind of the y equals x, slope of one. As your capabilities rise, your safety standards and measures also need to get stronger and stronger. And as long as those two grow in tandem and the safety measures are up to the challenge presented by the capabilities at any given level, then you're good. And probably these things in this telling kind of naturally should evolve together and both should develop in tandem. The idea is that this is important. Right? We don't wanna we don't want the capabilities to get far ahead of the safety measures. But I think also implicitly is like, there's not really much point, and it might be sort of a waste of time or perhaps even counterproductive to try to get the safety measures to be super robust relative to capabilities that don't exist yet. So the 45 degree line has been at least like one pretty broadly accepted guiding principle, I think, how the Chinese AI ecosystem at large thinks about AI safety. Now we can ask, of course, how's it going? And I already kind of spoiled the answer that the American companies are doing better, but we can dig into that in quite a bit more detail. Concordia AI, which is led by past podcast guest from about a year ago, Brian Say, they maintain a website called aimonitor.net, where they run a bunch of evaluations and plot them. And you can see just a really, honestly, almost overwhelming amount of detail on all the different models that they've tested, on all these different benchmarks. You can really dig into it in quite a bit of detail. The graph that I found to be the most informative was the one that compares proprietary API models, which are mostly American, to open weights models, which are mostly Chinese, and finds that for all the big risk categories that matter, the closed source, the proprietary API only models are at or maybe a little bit above that 45 degree line. Of course, we have questions of, like, the metrics. Right? They're they're kind of plotting a capability score and a safety score. And what exactly do these scores translate to in terms of model behavior or what it will and won't do? I haven't chased all these things down to ground truth, but this is these are composite scores. And what you do see in general is that the proprietary closed source, mostly American models are kind of at or above the 45 degree line, whereas the mostly Chinese open source open weights, I should say, least models are kind of lower and tend to be below the 45 degree line. So they're definitely not doing as well. This is an organization based in China reporting this. Notably, this is a public website. They've evidently feel comfortable doing this reporting and calling it how they see it online in a website that is available in both Chinese and English. So this is not like a secret or something that the the Chinese community can't handle or would or would, I think, particularly fight back against. It's pretty much just the facts, and those facts are also definitely echoed by Adam from far in the conversation that we had just a few days ago as well. He said, yeah. We got OpenAI and Anthropic at the top. Like, it's it's hard to jailbreak them. Gemini and Brock are, like, a lot easier. The Chinese models are even easier still. Like, this is a very consistent story across these two organizations on opposite sides of the world asking the same question coming to the same answer.

[15:04]Claude: Claude by Anthropic is an AI collaborator that understands your workflow and helps you tackle research, writing, coding, and organization with deep context. Get started with Claude and explore Claude Pro at https://claude.ai/tcr

Main Episode

[16:34] Now I think there's one thing that is worth keeping in mind for multiple reasons here, which is that sometimes I think there is a little bit of confusion between the open weight model that a Chinese company might release and the actual service that it provides to the public. I talked last time in my review of the Chinese AIs about how sometimes within the Chinese apps, I experienced this on multiple different Chinese AI apps, If you ask a sensitive question, you will sometimes see that the answer is coming in, and then you'll have that old school Bing experience where all of a sudden the answer that you were starting to read disappears and you get refusal. Sorry. I can't help you with that. So clearly, this is a multipart system. Right? It is not just the model. It's a model, but then it's also some monitor that sits on top of the model that's classifying or otherwise reviewing the output and can come down and say, nope. We're gonna cut you off right there, even though the model itself was happy to answer. So I bring this up because I think this can sometimes muddle the results and make the differences look a little more stark than they are. I did go into the methodology for the Concordia report, and they did say that wherever possible, they are testing the API direct from the company. So whatever systems they have in the API would be included. Do they have the same systems in the API as they do in their first party consumer app? Not always clear. And then in other contexts, when you see results like this, like I think when Anthropic does some of their testing and reports on the Chinese models, I think they are generally, if I understand correctly, just testing the model itself without whatever surrounding additional measures it's deployed within practice in the Chinese economy. Why does that matter? I mean, I think as always with these things, like, both data points are interesting. The most hawkish AI safety line of thought would be like, well, you put an open weight model out into the world. Anybody can use it, so we need to understand the worst case scenario. I think that's, like, totally valid and definitely is important for us to understand. And I think the Chinese might underestimate the importance of that in some ways because they tend to think a lot more about services than just the model. My broad sense is that what is regulated in China is a service. You are offering an AI service to the public. That is the kind of thing that gets regulated. You're putting a model out into the public domain. Okay. Who's gonna use it and under what context? They seem to kind of have the mental model that, like, when a company releases an open weights model, that mostly what's gonna happen with that is, like, other businesses are gonna pick it up, and they're gonna build their own services around it. And so if that is another Chinese company building a new service around an open weights model, then again, that will be regulated at the service level. And I think they sort of expect that, like, people around the world or countries around the world will probably function similarly where it was mentioned to me multiple times that, look. Like, these new models are trillions of parameters. Right? Like, this is not the kind of thing that you can run on your laptop. It's pretty far from it. You need some serious hardware to do any inference with these models really at all. So they were kinda like, we yeah. It's out there as open weights, but it's not like any random deranged person can download it to a phone or a laptop, do harm with it. It's really gonna be used in the context of other services. And so we should look more at, like, the context in which that model is ultimately used than just the model itself in isolation. I do think this is one way in which American AI safety discourse and Chinese AI safety thinking just kinda see things a bit differently. And I think both have quite valid points. I think the Chinese point is, like, legitimately apt. Right? I mean, it is hard to run these giant models. A random person, certainly somebody who's, like, having an episode or is experiencing some sort of psychological distress or psychotic break. Right? I don't know what the right terms are for these sorts of things. But this is not somebody who's gonna go set up the infrastructure to run the 2,800,000,000,000 parameter k three. Now that doesn't all the risk, but I think it's a valid point that it's not easy to run a model in isolation. And so how is it gonna be run by whom with, like, what surrounding stuff? That is a that is an important question that I think is maybe a little bit too for practical purposes, maybe a little bit too quickly long past in the American discourse. Although, again, the worst case understanding is also really important for us to have. So I think both perspectives are valid, and I do think there's a bit of a disconnect on that particular point.

[21:40] Now we could still ask, with all that said, how much do the Chinese companies care about AI safety? Heard, again, as I said at the top, like, they don't care. I think the Chinese system as a whole definitely cares. I did hear somewhat different reports on the companies themselves, especially the younger kind of LLM or AGI chasing startup type companies specifically. The there's a few different data points that I could share. One, again, Concordia who does a they also do a real good state of AI safety in China report, which is, again, public. They just updated it for right around WAIC, and so there's a July 2026 edition, which I do refer to or have referred to pulling the information for this episode. They said that they had reviewed 10 major Chinese companies safety disclosure practices. And they found that five of those 10 companies have done at some point recently a safety evaluation, which they then published along with the release of the model. So five of 10 have at least done something like in the spirit of your classic Anthropic or OpenAI model card. However, obviously, that means five have not done that. And even the five that did, they said that they're not necessarily doing it on every single release of a new model. So that definitely leaves something to be desired. I get the general impression that big tech companies, these are your sort of your Alibaba's, your Ant Group's, your Tencent, maybe to some extent your ByteDances at this point. I mean, ByteDance's a big company at this point. These companies that are very established, that have existing businesses, that are, like, making a lot of money in their existing businesses, that they are more apt to focus on this AI safety stuff. Why is that? Maybe they just feel like they have more to lose. They they really don't wanna get on the wrong side of the Chinese government. Maybe they just feel like they have the luxury of paying for it. Maybe it's just kind of institutional culture has matured over time, and they they do this kind of they have these kind of practices. Right? Certainly, you're a fintech company and you're moving money around, like, you're gonna be extremely careful about upgrades to that system and thinking about how AIs could run amok in that system. You just have a lot to lose. So I couldn't pin down a precise explanation, but my general sense is that those kind of big established incumbent big tech companies, they are more inclined to do this sort of safety work than the startups who are, for the most part, like, really just trying to race to catch up and be relevant. I think this is not too dissimilar from, like, where Meta was a couple years ago when they were open waiting llama two and llama three. I think the story that they told themselves at the time was and they do have a they do have a policy at Meta, and they do testing. And I did actually a brief engagement with Meta as a bread teamer for one of the models, which was, for logistical reasons, kind of a total failure in the end, at least my contribution to it was. But I think that they sort of said to themselves, look. Like, we're a year behind the frontier or something. And by the time we bring a model forward with a certain level of capability, OpenAI and Anthropic have already had that model out in the public for a year. And all the jailbreaks, there have been enough opportunities for people to jailbreak it and see what it can do. And we know that these defenses aren't super robust. So, like, if something really bad was gonna happen at g p t four level or g p d four o level or whatever level they're chasing at any given time, they kinda feel like that probably already would have happened. And therefore, they can put the model out on an open weights basis, and they probably won't be moving things too much. And I think at least so far, that has been fair enough, and history has kind of proven them right. It's not like we've seen crazy stuff happening as a result of llama 3.370 b that was at whatever GPT four o or whatever exact level it was at. It seems like it's been fine. And so I think there's some part of the Chinese kind of startup lab type companies that feels similarly where they're like, look. It's already been out there for a while. If nothing bad has happened, if we don't have any incident reports, then we can probably follow-up to that level of capability and not have to worry about it all that much. Will that change now as we are starting to see actual serious incidents being reported out of the frontier companies on the American side? We'll have to see, but I think there's, like, definitely a very plausible story that a very one might. So, obviously, there's a lot of different takes. There's a lot of different I can't go company by company.

[26:51] This is abstracting away a lot of detail, but I think it's safe to say that the 45 degree line, which has the idea as its core, the idea that safety measures should grow step for step with capabilities, I think that's still a bit aspirational in China. I think they could do better, and I think they have a little bit of a reason to believe that it doesn't matter so much because the American companies have already explored what happens at any given capability level that they are following into. But I also think it's definitely fair to say that AI safety is still very much aspirational here in The United States as well, and we should definitely be keeping in mind that if it weren't for the two companies really at the leadership at the frontier of both capabilities and safety measures, then the two clusters would look a lot more overlapping, and the question would be a lot more muddled than it is today in terms of, like, who's really as an ecosystem doing a better job with AI safety and safeguards. Okay. That's the current state of, like, deployed AIs. There's obviously a lot more going on in terms of research and the role that the government is playing in China, and so I wanna get into those next. On the research front, it is very clear that Chinese researchers are increasingly engaged with and concerned with and focused on and actively shipping research regarding AI safety issues of all kinds. I would say this was probably inevitable just because, at least for now, both AI ecosystems are developing essentially the same technology and gonna hit essentially the same problems and naturally gonna reach for similar solutions, both because those are natural solutions and because there's the opportunity to look at what the other is doing and try to copy the best of it. I think there has been some really good citizen diplomacy that has gone on as well. While in China, I did, especially at WAIC, I did bump into a number of people who were there to participate in various fora and track two dialogues, some of which are closed door and off the record and give people a chance to be, like, very candid with one another. And you see some of the people that you would expect to see there, people who have been arguing very clearly and forcefully in the American discourse that we have to get an international treaty to get this stuff under control, that cooperation with China is on the critical path. I think you can rest assured that those people are acting on their stated beliefs. Like, I saw some of the same people saying that stuff online. I saw some of them in China. They are doing the work. And it does seem like it really has at least, again, I think a lot of this probably could have been expected to happen organically anyway, but I do think their efforts borne some fruit. I wasn't able to participate in all those kind of track two dialogue sort of things as somebody who's like kind of journalist coded. There were a couple times where it was like, I think everybody will be more comfortable if we just don't have somebody there whose main credential would be a podcast. So I wasn't quite privy to all of the most kind of candid off the record conversations, but there's just so many moments where very recognizable ideas or even like American or British AI safety organizations were just name checked directly that it's clear that there is, like, meaningful cross pollination of ideas. And mostly, it has flowed from from the West to the East in this case. It's not, of course, like the Chinese scientists or researchers who are receiving these messages are, like, taking them uncritically or just doing whatever they're told by by whatever Westerners roll in. It's not like that at all. Somebody actually told me that there are people who view AI safety as some sort of Western SIOP designed to slow China down or prevent them from catching up or whatever. So it's like we have our cynical voices. Apparently, are cynical voices in China too that would express concerns along those lines. I didn't meet anyone who seemed to believe that or who said anything like that. And, again, my guess is that those kinds of notions are going to be fading relatively quickly as the clear and present danger of some of the latest models capabilities becomes more widely known. But at least there's been some of that out there. I I thought that was worth mentioning, if only because it it does make clear that Chinese people are perfectly capable of thinking for themselves. They're not just taking whatever whatever Westerners come to tell them.

[31:57] I think what really is happening is that they're pretty open minded, and the ideas are pretty compelling. And, again, the examples are becoming more colorful and more real all the time. Interestingly, one big tech company where I had the chance to, along with several others, to meet with really a pretty impressive leadership roster, they brought out their CMO, their head of communications, their general counsel, their head of AI security, more people beyond that. But, I mean, this was a really pretty senior group of leaders at this big tech company. They said a couple of things that were producing. One, they just first of all, it, like, almost pounded the table and said, we really do care about catastrophic risks. Like CBRN risks, we absolutely do care about that. So they were very adamant that, like, at least for their part, they care. They're aware and they care. And they also said that they are following American AI safety discourse on a daily basis with, get ready, an agent that goes out and surveys American AI safety discourse on a daily basis and gives them a daily report of what is going on in AI safety in The United States. I thought that was pretty interesting. Certainly, I don't have too much of that going in the reverse. So, again, I think bottom line, engagement has worked. Probably some convergence could have been expected over time, but the people who have been saying we need to work with China have in fact also, at least some of them, been doing the work, and it seems like that work has been going at least reasonably well. And, again, I think evidence will mount as I as I continue forward through this outline. One big thing to emphasize too is it's not just content safety. Content safety in the Chinese context is like, again, the three t's, Tibet, Taiwan, Tiananmen. Everybody knows that Chinese models, at least Chinese services. Right? Often the model will be more inclined to answer your question. It's then some other part of the service, some monitor or whatever that will shut you down on those topics. But everybody knows that, like, that's sensitive in the Chinese context and, like, companies kinda have to get that right according to the Chinese government if they're gonna be operating. But then some people will say, oh, that's all they care about. All they care about is censorship. And, again, this is, like, definitely not the case. The big trend right now, I would say, at this point, really, it feels like they feel like they've got the content thing pretty well figured out. All these companies have launched. They're they're in the market. Things are happening. They're doing business. They're not that worried about this sort of content safety today. What they are worried about, what they are talking about nonstop is, like everybody else, agents. Agents taking autonomous actions, what could happen, can we keep them under control. Of course, there's questions of exactly what do you mean by control. Agents, agents, agents, that's what people are talking about everywhere. The big thing that they just kind of, in a very plain spoken way, motivate AI safety discussion with is like, AIs have gone from answering questions to actually taking actions in the world. The digital world still mostly, of course, but, like, they've got a big emphasis on robotics. Right? So they fully expect that these agents are going to leave the digital world and find their embodied successful commercially viable selves and be out there in the physical world doing things as well. And this raises all sorts of questions of like, jeez, if these AIs are empowered to take action autonomously, we better make sure they're taking actions that are good and that we like and that are not causing big problems. So this is like very down the fairway, again, practical, grounded motivation for these issues. But you hear that, I would say, everywhere I went. It was like agents and jeez. Agents, boy, they can take action in the real world. So we really gotta start to get a handle on that. A big thing that I think is is very different about the American and Chinese AI ecosystems, and this is, I think, an important one to understand, is that China doesn't really have the same kind of nonprofit sector that The United States does. There are nonprofits in China. You can set one up. But the scope of what you're kind of allowed or expected to do, as far as I can tell, is much narrower. There's obviously no appetite for political activism, so that's, like, right out. And it seems that most of the nonprofits are, like, service organizations that are there to address some obvious down the fairway legible social problem.

[37:02] And the money, the philanthropic side, also, again, seems to be a lot more down the fairway, a lot more just kind of conservative, conventional, doing things that, like, everybody can agree is good to do. Much less speculative stuff than goes on in The United States. I think actually that this is a huge strength of The United States that we should really not take for granted. Like, the fact that we have this civil society where anybody can go set up a nonprofit and have their crazy ideas and go they don't have to get permission from the government to go try to chase down an agenda. They just need to convince one wealthy patron that they have something worth chasing. I think that is, like, a great strength for us, and it's it's given us among many other things. Really, the whole of the AI safety community that we have today. Right? It was, like, extremely fringe when it got started, but there were a few philanthropists who took it seriously enough to help people keep the lights on and help them do the work that they wanted to do. Sure enough, here we are. And so many of the predictions that have been that were made many years ago are coming sort of true or at least, like, true enough to be scary. And it's, like, really good that we have this ecosystem, and we we just wouldn't have had that if it had to be all approved through the government. And so China, because their nonprofits do have a much more sort of heavy and restrictive government approval process and generally just much more narrow and conventional scope of action, like, they don't really have that. They don't have they've never had the opportunity to develop this sort of ecology of AI safety organizations the way that we have here. As a result, most of the AI safety research that you find in China is actually coming out of the universities. And to some extent, the companies, and there are definitely academic industry collaborations as well. But the number one source seems to be the universities, the sort of analog for the nonprofit sector in The US. When it comes to who is producing the bulk of the AI safety work, it's academia in China. And so I think you can compliment them and say, wow. You're I think it's safe to say that their academia has moved faster than our academia has moved to take up AI safety as a research area, and that's cool. But it's been slower than our nonprofit sector has. And I think because of the kinds of people that tend to be professors, and this is true across both countries. Right? You do have your iconoclast professors that are typically older these days. And it feels like now we kinda have a more conventional profile, people that have been kind of careerist. And I don't mean that in a dismissive way, but, like, people that have gone one rung up the ladder at a time through the PhD and the postdoc and getting the the first professorship. These are, like, fairly institutional people. They're people that do value creativity and research and new ideas, but they do it in a pretty conventional way that, like, doesn't push the boundaries too hard, that tries not to seem too weird, certainly that, like, doesn't borrow the aesthetics of less wrong or talk too much about, like, super unlikely tail risk scenarios. It's just because it is academia and because people have kind of followed something more like that career path, and I I can't say I know exactly what the ins and outs of the Chinese academic career path are, But you clearly get the vibe that these professor types are more like American professor types than they are like the the moral weirdos, if you will, that were, like, first sounding the alarm about AI safety years ago. But this is where the as far as I can tell, this is where a lot of the early and best work has come from in the Chinese context. So if you go over there as a a sort of whatever, I don't wanna be too silly about it, but a sort of blue haired, polycule, AI safety hawk member of the American community and you're looking for your peers, the reality is you just might not find them. You you won't necessarily find people that kind of look like you, have the same attitudes as you. What you can find, and I think the the most strategic and effective of the American AI safety is when they've gone to do their their bridge building and try to have meetings of the mind with people as much as they can. I think what they've ended up connecting with mostly people in the academy. So that is pretty interesting. And again, I think you do just feel that in a bunch of ways.

[42:07] The AI safety work in China, it's it's less speculative. It tends to be more focused on reliability. It tends to be a little more focused on, like, protecting minors. All good things. You don't hear p doom talk. Actually, one of the more interesting moments of the entire trip was sitting at a table with a professor who has done a lot of AI safety work. And I asked him, like, what are the lunch conversations like? Do you guys, like, trade p doom numbers back and forth? Do you read AI twenty twenty seven? What what's the kind of vibe? And he said, well, we don't read AI twenty twenty seven, really. It's too political, and there's kind of US China dynamics and whatever in there that may scare people off from wanting to talk about that too much in the Chinese context. He also was like, no. P doom? Like, no. We're not, like, trading P doom numbers at lunch. Like, he kind of looked at me and was like, do Americans do that? I was like, yeah. Oh, yeah. Definitely. Like, if you come to the the bay and you go to any number of venues and have a lunch conversation, like, yeah. People will at this point, maybe it's a little passe, but, like, it it's definitely in the water that people think about this question in a very live way. He seemed to find that a little bit surprising, actually. Just seems like a little bit and this is somebody, again, who's done a lot of work on a lot of different aspects of AI safety, put tons of papers, but seemed a little bit kind of surprised that, like, wow. That's, like, that's pretty far out conversation to be just considered normal in the at least from his perspective in the Chinese context. So I think those are those are interesting and kind of important comments mostly because, you know, in some ways, we do find these mirror image structures in China like big tech. Right? Like, again, I talked last time about how going to have lunch at ByteDance felt almost exactly like going to have lunch at Google. But this is a little bit different. You don't have quite the same type of people leading the effort. And so I do think that creates some some risk of miscommunication. And we've obviously seen that, like, the vanguard of the AI safety community in The US has at times had trouble communicating with our own government. It certainly has some inroads into academia, but not as much as probably they would have hoped or expected by this point. Imagine how tricky it might be for them to go engage Chinese academia, let alone the Chinese government. I think it it it does create some potential for disconnect, some some potential for confusion, what have you. But the activity is there. It's just coming from a different institutional context, coming from people with quite a different personality type just based on the the kinds of career paths that they have chosen. Now why has the Chinese Academy moved faster than the American Academy has when it comes to getting serious about AI risk. I don't have an answer for that really, but I think one candidate answer is that it comes from the top. So I have a few quotes here from Xi's speech at WAIC that I think would probably surprise many American listeners. Certainly, should surprise those that would say China will never care. This any regulation we do is a gift to Xi. I would think twice about that. Here are some quotes from and this is the opening keynote, the first twenty minutes of WAIC, this big event. He's there to headline it. I wasn't in the room. There were only a few 100 people in the room. And I did I did talk to a couple of people who were in the room. You know, security was obviously super tight. They they were in the room for, like, a couple hours before he showed up and took the stage. So really big deal. And they also they think really hard about these speeches. He's not somebody who's, like, out there winging it. They he chooses his words very carefully. They also choose their words in the translation very carefully. They put an English translation out. If you a couple events that I went to, there was live simultaneous translation, including of a couple panel discussions that were happening in Chinese and then being translated through the earpiece live for me and and others in the audience who didn't speak Chinese. Very nice of them to do that. Right? I mean, you would not get that sort of translation as kind of a standard expectation if you came to a similar event in The US. You'd be on your own. You'd be expected to speak English or figure it out for yourself. There, they obviously don't expect that we're going to speak Chinese. They want to welcome guests. They do go to the trouble of doing this simultaneous translation. Nevertheless, if you're talking a panel discussion, at least in my experience, the simultaneous translation of panel discussions is rough. Like, I there were there was a one that I remember especially where I was like, I have no idea what they're just talking about, really. I you were getting the kind of big themes that they were talking about, but, like, really, what were they saying?

[47:16] I found it very difficult to really get that from the live simultaneous translation. Now they don't do random translator decides what the English version is going to be on the fly when she gives a speech. He's got his speech prepared, and they've got an English translation of that ready to go. And what you are getting in the English translation is, like, gonna be, again, pretty well and carefully thought through. So I sourced these following quotes from and there I think there are a couple different versions still somehow. I'm a little bit confused honestly by why there are multiple different translations. I guess the the Chinese government gives an official one. Of course, people are free to do their own translation, and they might translate certain things a little bit differently. They might think they have a better better way of understanding what was said in Chinese and how we should understand it in English. The version that I'm pulling from here was from Matt Sheehan's blog where he just provided a full English translation and then did a bunch of commentary on it. But here are three quotes that I think should cause anybody with an extreme position on what TriNet will never do to at least soften up a little bit and begin to reconsider. Here's the first one. Quote, how should humans coexist with machines that think? How can safety be protected when algorithms participate in decisions? How can governance keep pace when technology challenges ethics? That's all. That's big questions. I think the right kinds of questions from president Xi's opening speech from WAIC. Here's the second quote. The faster AI advances, the more firmly its direction must be anchored toward human benefit, the more precisely governance must be calibrated, and the more rapidly safeguards against loss of control must improve. And then he closed by saying that countries should, quote, strengthen risk awareness, confront AI's inherent and downstream risks, build legal, technical monitoring, early warning, and emergency response systems, prevent misuse and malicious use, and keep AI under human control. I believe that's the final section of the speech. Okay. There's a lot of echoes there of big ideas from the American AI safety discourse. And I'm not saying he sourced them from the American AI safety discourse, but you might call it instrumental convergence where people that are worried about making sure that these transformative technologies go well seem to, one way or another, land on very similar ideas. Keep AI under human control is, I think, a pretty heady idea. Right? This is not somebody who can't think about the big picture like it that is the big picture, I would say. I invite people to compare. Of course, you do have the question, and this is important as well, what exactly is meant by loss of control? A cynical read which I'm not really fully qualified to parse myself. But I again, I think there's a lot more supporting evidence. But a cynical read is like loss of control just means content safety. Right? It's like it's social control. It's the government's ability to control the people. I don't think that's the right read of this. I think that is definitely part, obviously, of what the CCP wants to do. But I don't think that is the full story, and I I think that's kind of unwarranted cynicism. And I would just invite anybody to compare that speech to whatever you think is the most pro AI safety speech, whatever is, like, the the the highest situational awareness, AI safety related comments from any prominent American politician today. I think, like, Bernie Sanders has kind of said some interesting things. He's also said some things that are a little bizarre from my point of view. Love Bernie for who he is, what he is, and how sincere he is. I don't think he's he's pretty far from power these days. Pretty far from actually pulling the levers of policy. What else have we got? Right? I mean, if you compare this speech from Xi to what JD Vance said in Europe not that long ago, Xi comes off looking positively AI safety hawkish by comparison. So right there, again, I think people should be softened up a bit to believe that, like, hey. Maybe there is actually some willingness to take these issues seriously in China, and maybe we're not gonna just seed the future to them if we do anything on our own because maybe they are actually open minded to doing more than we would have assumed. With that, let me get into some of the experiences that I had and some of the some of the research specifically that I think, again, should further support the notion that there is really a there. Actually, a couple days before WAIC, I originally flew into Beijing.

[52:17] And and the reason that I flew to Beijing aside from wanting to go see the Forbidden City and the Great Wall and a few things like that was that I had an invitation to attend the opening of an AI safety hub at Tsinghua University's College of AI. It was launching just a couple days before WAIC began. And interestingly, I don't think there's been any international media of this. I haven't been able to find anything by searching for it, and I haven't been able to find anything really at all on the English Internet. Like, Claude can't find anything for me. I do have a link which we can put into the show notes to a Chinese media source. And, again, this just goes to show, like, they understand us a lot better than we understand them. Right? That they this big tech company that I mentioned earlier has an agent trolling AI safety Twitter and putting together reports. And as far as I know, the American media has not managed to publish anything on the launch of an AI safety hub at either China's premier university, Chinghua, or one of its very top tier universities. It was a full day event. There were some research presented. There were some statements made about the aspirations for this AI safety hub and what they wanted it to be. And couple things really quite interesting about it. One is that they're really trying to be very international and collaborative. There are, like, five founding members of the of the institute, kind of the board. The the leadership group was five members. One of them is a European professor who is actually taking a position at Tsinghua to help develop and lead this institute. So right off the bat, they've got an international person on their leadership board. And then multiple times from multiple different people, they specifically cited other AI safety hubs around the world as inspiration. They specifically cited Constellation and Lisa in London as what they want to be in Beijing. They wanna be meeting place, a place where people can come for a time and do their best work, where ideas will be exchanged, and again, highly international flavor. Some of these speeches were actually in English, others were in Chinese with the simultaneous translation. But really amazing to go all the way to Beijing to sit in this launch of this AI safety hub at one of China's top, if not the top university, and hear Constellation and Lisa name checked as, like, this is what we strive to be. Like, these guys have done great, and we wanna be like them. I thought it was a really, really interesting and, like and quite telling in in important ways. The other thing that happened at that day was a bunch of research was presented. And, again, here, you would just recognize the research very readily. Just honestly, if you're just a a listener to the cognitive revolution over time, the organizations that were name checked, like, literally having their logos on slides as, like, previous work, inspiration, again, like, organizations that we think are excellent and we wanna be more like or wanna be on their level. I heard Apollo Research, METER multiple times, Palisade, the UKAC, I think Redwood Research as well, all of these organizations called out by name by Chinese researchers as they were presenting their own research. I thought this was, like, extremely impressive in terms of, like, how aware they are of what's going on in the rest of the world, how open minded they are to taking inspiration, and just how, like, there there's not a sense, at least at the AI safety level, there's not this sense of, like, we have to do it all from scratch or the sort of not invented here bias that has a lot of organizations, including, I think, often like American culture at large, rejecting good ideas from other places. This was a very open, clearly desiring to be collaborative. And, also, they're gonna be they're gonna be bringing people in and sending people out. That's a big part of their mission too. They're looking for international applicants to come spend a period of time in residence in Beijing doing research at the hub and cross pollinating ideas there. And then they're also gonna fund their own students to go abroad to other AI safety hubs around the world, Constellation, Lisa, and probably a bunch more as well. So and I I think they said that they had counted 16. If I if I recall correctly, it was it was definitely in the teens number of AI safety hubs created all around the world. We didn't get too many name checks on those past Constellation and Lisa. There was also a mention of Sash, the Singapore one. Those those kind of most prominent few were called out by name, and they wanna be in that top tier. They want the AI safety hub at the Tsinghua University College of AI to join that upper echelon. And it seems like the resources are there.

[57:26] I didn't really understand entirely where the money had come from. They said it wasn't government money, wasn't university money. I guess it was kind of philanthropic. This maybe begins to complicate or contradict a little bit of what I said earlier around these these sort of nonprofit sector being like more narrow in scope. But I think also it's it's about 2026. Right? And this is just happening now. So I think we are at the point where the Overton window, even within academia or official nonprofit realm within China, can see that this is, like, actually something that that is a real issue that's not some Western psy op that really is worth taking seriously and putting some resources into. So that was cool. It was a really it was a really neat experience. There was a moment that I thought was quite endearing. I'm always a fan of when people are not too cool for school, when when people are pre is is strong, but, like, when people are less self conscious and more willing to go with something that's kinda silly just because it's, like, the thing to do in the moment, They had this sort of official moment of, like, okay. This is now the moment when we're gonna launch this hub. And the five board members were all on stage, they brought, like, a little podium in front of each one. And each member was instructed to place their palm on the podium in front of them, and then there was this, like, big sort of graphics package that kind of erupted on this giant screen behind them on stage. And it was, like, in one sense, like, pretty cheesy, I think, objectively, or at least objectively through an American cultural lens. But it also sort of hearkened back to me to, like, reading, like, a Teddy Roosevelt biography or something where they used to do these break a bottle on a ship ribbon cutting ceremonies. And I think they were just less focused on looking cool and more inclined to kind of get excited about that kind of moment. I felt that a little bit in that moment in China. I I honestly thought it was very endearing. One of the things I probably like least about American culture is when we are so concerned with how we're gonna look or what the perception will be or am I trying too hard in this moment to be cool that we kind of shrink away from those moments or, like, perform them in a sort of disinterested way and or or try to create some ironic distance between ourselves and what we're doing. And that didn't seem to be present in that moment. Like, it seemed like this was kind of the thing where the board members weren't, like, expected to do jumping jacks or do anything, like, theatrical. And they didn't, but they just they did their part, and it was like, yeah. This is the moment. It's official. It is launched. Round of applause from the audience. And I found that an endearing moment if only because it kind of contrasted with our, I think, sometimes overly image conscious culture in The United States. Anyway, that was cool. What does it add up to, though? Right? We saw a number of papers presented at this launch. Again, recognizable subjects, recognizable inspiration in the form of Apollo meter, Palisade, UKAC, etcetera. But how much of this work is, like, really going on? For this, I would go back again to Concordia. They have this website that monitors all these papers. It's aisafetychina.com. So if you go to aisafetychina.com, and then you can go into the research thing. You can drill down into, like, all these papers that they compile and classify and so on. In just raw numbers terms, they have gone from just a trickle of AI safety research coming out in 2023, like, literally just a couple papers a month back then, to now something like 50 to 60 papers per month as of mid twenty twenty six. And it's growing quickly as you would expect. Right? It's it's more than 10 x over the last three years, and the trend is, like, obvious. And, certainly, I think we can expect that to continue. How does that compare to the American research output? Obviously, like, just counting up the numbers of papers is not that great of a measure in the first place, but I did ask Claude and to do it just to give me kind of a rough point of comparison. They both of course said, well, it depends on exactly what you wanna count and blah blah blah. But they each basically gave me an estimate that goes from 50 to a couple 100, maybe a few 100 papers per month coming out in The US or the, let's say, the Anglosphere on AI safety. So the volume is, like, definitely higher in The US, but it's not like order of magnitude higher. It's like multiple higher, not order of magnitude higher.

[1:02:24] And I don't think that's surprising at all, but it just kinda gives a sense that the Chinese ecosystem is, like, growing quickly and it's, like, getting to roughly the same or at least approaching kind of the same scale as what we have with a thanks to our dynamic relatively permissionless nonprofit sector. We've had, like, a significant head start, and I would say even the Chinese ecosystem is catching up, least in terms of this kind of crude measure of raw volume of research papers put out. Now what are these papers? There's when there's 50 to 60 a month, it's obviously gonna be very difficult to summarize what they are. But I would say that from what I've been able to understand, they're kind of across the broad spectrum of different risks and types of research that we see in the West as well. I thought maybe the best way to give you a little sampling of it was just to pull together some titles of papers and read those, and you can judge for yourself how similar they sound. So here's one from I thought this one was notable, especially because it was 2024, which is pretty early. Paper is called Frontier AI Systems Have Surpassed the Self Replicating Red Line. Definite echoes of Palisade. I had Jeffrey Ladish on the podcast not too too long ago, and we were talking about that kind of work, right, where they were showing that an open source model, and they were using a Chinese open weights model, was able to go hack another server and copy itself and set up itself. This is a very similar line of research happening roughly contemporaneously in China. Palisades has been doing that kind of stuff for a few years now. This was 2024 out of a Chinese group. A May 2025 paper, which has been updated in 2026 as well, called Evaluation Faking, Unveiling Observer Effects in Safety Evaluation of Frontier AI Systems. So, again, extremely recognizable analog to the eval awareness work that I would say we've mostly seen coming out of anthropic, but certainly others have picked that up as well. And they're well aware of it in the Chinese context as well. September 2025, r squared AI towards resistant and resilient AI in an evolving world. This one I thought was notable because it comes from the Shanghai AI Lab. So the the supervising author, the the last author on the author list, is Zhao Boen, the same guy who coined the 45 degree line concept. And in the introduction to this paper, they cite the guaranteed safe AI paper that we've done episodes on in the past, which was led by Davedad, a recent guest, and he had a broader there's a that was a significant coalition piece that they put together, but Davitta was the lead author. And it's in the first paragraph cited as kind of motivation. What what intellectual tradition are we engaging here? Boom. There's Dalrymple right in the the opening of this paper from the such a prominent organization as the Shanghai AI Lab. Another one from October 2025 is called Deception Bench, a comprehensive benchmark for AI deception behaviors in real world scenarios. Shades of Apollo there, obviously, those guys are the kind of most focused on and leaders in the science of deception. Well, this is Deception Bench. November 2025, when alignment fails, multimodal adversarial attacks on vision language action models. This, if anything, I think is maybe a little bit more Chinese flavored because they're so focused on robotics. And we're focused on robotics too, but they're really focused on robotics. So this paper showed that there were various different ways to use data perturbations and similar techniques to cause failures in vision language action models. And nothing too remarkable about that, but again, you could see very similar work coming from all kinds of groups here. That same group later followed up six months later in April 2026 with a paper called Strong VLA, decoupled robustness learning for vision language action models under multimodal perturbations. And so this is basically the the follow-up paper where they said, hey. Last time, we showed that you can break vision language action models with these weird kind of adversarial exotic attack methods. Now here's a strategy that seems to mitigate that too. Obviously, none of these things ever work a 100%, and that's exactly the same across the Chinese and American contexts.

[1:07:23] But what they introduced there was a two step pipeline where they first did a bunch of robustness training and then came back to doing the task fine tuning, and they found that the robustness kinda held to these these various attack methods. And again, nothing perfect, but I think you could just imagine and you could probably point to all sorts of American groups who've done very similar lines of research where they're like, hey. Look. We found this way to attack and break models, and now here's a way that if you do this, you can reduce that by whatever. Usually, it's 70% to to an order of magnitude reduction. So all that stuff very similar, almost direct analogs between these two kind of I I think you you could make a direct analogy for basically every single one of these papers and find the kind of corresponding one in from an American group. Interpretability is also on the rise in China. A couple interesting recent papers just from this year. One is called mechanistic origin of moral indifference in language models. And the abstract of that paper starts with this sentence. Existing behavioral alignment techniques for large language models often neglect the discrepancy between surface compliance and internal unaligned representations, leaving LLMs vulnerable to long tail risks. That sentence really jumped out at me because, again, this is like even though it's coming from a more academic context and using more, I don't know, neutral or sort of technical sounding language than what you hear from the lesser wrong crowd in many cases. Very similar ideas. Right? We're talking about the difference between surface compliance and internal unaligned representations that could lead to long tail risks. Like, that's total less wrong cultural victory if I've ever heard one. Another one also from this year called safe seek, universal attribution of safety circuits in language models. And here's the beginning of that abstract. Mechanistic interpretability reveals that safety critical behaviors, e g alignment, jailbreaking, backdoring, in large language models, are grounded in specialized functional components. However, existing safety attribution methods struggle with generalization and reliability due to their reliance on heuristic, domain specific metrics, and search algorithms. So, again, you have this idea, and then, yep, I haven't studied this work, so I certainly can't critique it or say if it's if it's amazing or not. I'd love to have the ability to run myself in parallel in order to do more of that stuff. But, again, a very similar idea. Right? That, like, the good behavior that we see, you hear from the the AI safety voices in The US all the time. Sure. Cool. But it won't scale to superintelligence. And this is, I think, a very similar idea where they're not saying it quite that way, but they're saying existing safety attribution methods struggle with generalization and reliability due to their reliance on domain specific metrics. I think the core question is the same. Right? We've got these AIs behaving pretty well in the ways that we can anticipate that they might go badly and that we can set up test cases and that we can measure and that we can train accordingly. But are they really good under the hood? How do we know? Are they gonna generalize? Very similar questions that the mechanistic interpretability researchers in China are trying to get a handle on. I think the final one I'll give you right now, I thought this one was particularly uncanny. And this may change because I I actually have not been able to find this paper on the Internet. So I just saw it presented at an event. It was a it was an event surrounding WAIC. The line between, like, what's official WAIC and what's unofficial is kinda blurry to me, honestly. There's, like, the Expo Hall, which is, like, you're definitely at WAIC there. But then there's all these events around too where it's, like, you may go to some conference room at a nearby hotel, and they've it's like seems pretty official, but is it like official official? Anyway, I'm not even sure that matters. The point is it was at one of these surrounding events where again some research was presented. And this paper, I think, is still to be published, but it was presented to this group with the title, Toward Decoupling Capability Growth from Isolating Hazardous Capabilities in Mixture of Experts. And then the promise on the slide was harmful experts can be switched off or removed at inference. And this really jumped out at me because this is probably the technique that I've seen recently that I've been the most excited about. This came from AE Studio in collaboration with Anthropic. I've talked about it on, like, half a dozen episodes probably already.

[1:12:15] So you're probably tired of me talking about GRAM, which is a gradient routing technique that's meant to localize particular kinds of knowledge to particular identifiable experts so that you can potentially release your model open weights if you want to, but maybe hold back a small number of the experts that that contain the dangerous capabilities. So you still have the in the sort of optimistic telling of this, you still have one training run. You still can teach your model everything. It can still have these bio capabilities that you may want. Jud from AE Studio points out very simply that, like, these dual use capabilities do have a lot of good. We do want to have our biologists using AI to cure all the diseases. That's going to be pretty important. If we can't allow our biologists to use our smartest AIs to cure our diseases, we're leaving a ton of value on the table. But we also might especially as they get really powerful, we might not want to just throw that out on the Internet. So what do we do? If we can get the knowledge localized into particular experts, then we can distribute an open weights version that just has a couple experts redacted. It'll work just as well for the vast majority of people, for the vast majority of use cases, but it won't have the the dangerous capabilities that you're worried about. And those can be made available through a more structured program with know your customer and various other safeguards that will hopefully give us the right balance between freedom to use AIs and freedom to do your own research and modify them and not be always beholden to some particular company. And I think all that stuff is really important. And also not wanting to ship the ability to engineer a pandemic with the the next version of some of these models. And it does seem like given what we've seen in cybersecurity, depending on the decisions that people make, the biorest seems like it's a not too far into the future thing that is really gonna start to happen. Again, it's this is not like something that is written in stone that it must happen. It depends on people actually training the models. But the path that I understand us to be on is that we're going to get super bio capable models. I mean, they're already pretty damn capable. Right? But, like, in the next year to eighteen months, people seem to think that we'll kind of hit a similar spot to what we've now hit with cybersecurity where the models are starting to become meaningfully superhuman in some ways. Not that nobody could ever do some of the things that they are doing, but just especially when you consider the speed at which they are operating if and when that same level of capability comes to the bio domain, like, will be a really serious, serious risk. It's one thing for a model to go hack some servers and cause some cyber chaos. Like, I am fundamentally not a cyber creature. I am a biological creature. And so what happens in the biological realm matters to me in a much more profound and kind of inseparable way from what it is that I am than what happens in the cyber realm. And, anyway, I'm I'm very bullish on that. That is like a I think a very I want to have our cake and eat it too when it comes to open weights models that people can use on their own terms and preventing really crazy shit from happening, especially in the biorem. And boom. There it is in China too. Right? Decoupling capability growth from risk growth, isolating hazardous capabilities, mixture of experts, harmful experts. Harmful experts can be switched off or removed at inference. So I think you look at this research lineup, and hopefully, that gives you enough to say, okay. Of course, these are two civilizations with totally different histories and totally different languages, and they're on the opposite side of the world. Of course, there's gonna be some differences, different institutions. We've got the whole nonprofit versus academia contrast. But, like, when it comes to the ideas that are being actively worked on, there is an awful lot of overlap. I think AI safety has taken root in China, it's it's safe to say. And when you see this kind of research, I think it gives you at least some license to interpret president Xi's remarks as being, like, not narrowly scoped to social control, but they they care a lot there about expert opinion. They care a lot about they obviously have a super long tradition of scholarship and taking what scholars say seriously. They they China might be the most scholarly society in the world. We could debate that, but it's certainly a contender. And I think they are and then they've even had things like a Politburo study session on some of these topics.

[1:17:20] So at the highest level, they are there is strong evidence that they are engaged with these topics. So I think we can say pretty pretty clearly that China, whatever the the sort of the macro in the macro sense, China does care about AI safety. China is working on AI safety. China has many researchers publishing many papers across a wide range of topics that overlap in very clear and direct ways with the same topics that we are talking about. And they they're also, like, not shy about citing American sources indicating when they've taken inspiration from an American group exactly who that group is. So that's another, I think, strong reason. The fact that they're name checking these organizations. It's like they're on the same wavelength about a lot of these things. They have a little bit of a different it's come out of a different part of Chinese society, but you can find people who are thinking very seriously about just about all the same issues there as you can find here. Okay. Now let's talk about governance. So far, we've talked about actual deployed systems and how they compare. And, The US ecosystem is ahead, but really being carried by a couple leaders that are bringing up our average. And we've looked at research, and we've seen I guess, we've also looked at rhetoric briefly with Xi's speech. And I would say there, the the Chinese executive is, like, certainly much more AI safety build than anything that I could point to in the West. And then we look at the research itself, and we see that there is just a lot of overlap. I think an interesting game to play, maybe I'll even ask Claude to cut up this this this game, would be just take two AI safety paper titles, and you have to guess which one is American and which one is Chinese. I think you would find it very difficult based on the lineup that I just gave you. And all this should give us a lot of confidence that on an intellectual level, there has been a lot of exchange. There has been a lot of cross pollination. And, again, in part because of that, but also in part because of instrumental convergence and the fact that we're developing essentially the same technology with essentially the same methods, naturally, we're gonna see the same problems and naturally, we're gonna try to come to the same conclusions. Overall, I'd say with a smile, the Chinese AI safety community, they're just like us. Governance is going to be definitely a bit of a different topic, but here, I think you can make a pretty strong argument that the Chinese government is ahead of the US government, at least if your definition of ahead means they're doing more. You might think, depending on your your point of view, you might say the US government is doing better by doing less. But the Chinese government, I think it's pretty safe to say, is doing more. Now I am definitely far from an expert on the structure of the Chinese government, and it does seem that there are some overlapping jurisdictions both between the national government and more local governments. And then also even between different parts of the national government, there are at least a couple different agencies or ministries or whatever that have at least some jurisdiction over AI. With that said, it seems like the big one, the one that comes up all the time, is the Cyberspace Administration of China, which is also known as the CAC, the Cyberspace Administration of China. That seems to be the leading agency that is doing the bulk of the standard setting and that the AI companies are in touch with on an ongoing basis. So I'm gonna run down a bunch of stuff that the Chinese government writ large has done. Not all of that stuff was done specifically by the CAC. But if you wanna look at, like, the the big kind of omnipresent regulatory body in China, it's, as far as I can tell, the CAC. So what have they been doing? The last battle in terms of technology regulation that both civilizations have faced and have handled quite differently was social media. And the Chinese government has been much more active in terms of regulating social media for better or for worse. I think everybody can tell I'm, like, pretty pro AI safety in many ways. I do think it's possible, of course, to have overreach and, like, I want my self driving car and I want my diseases cured. So I'm very mindful of the cost that regulation imposes. But to me, it's pretty undeniable that we're going to need every lever at our disposal to make this thing go well, that includes government action at some point and hopefully in in wise ways.

[1:22:27] I don't actually necessarily feel the same way about social media. I think it's much less clear that regulation of social media would be good for society or at least would be good for American society than I think it is I think the case is much stronger that AI regulation will be good for American society. And I know a lot less about Chinese society. Something could be good for our society and not so good for their society and vice versa. So I'm not taking a position in terms of endorsing what the Chinese government has done in social media. I'm also not saying it's terrible for them. Like, it might work for them. I really don't know. But the point is just to going back to the top, to make the case systematically that they do care, that they are willing to take action, that sometimes those actions will even slow their companies down or make them less profitable or impose other kinds of costs, and that they are still willing to do those things when they feel that it's necessary, you can definitely see that in the social media era. They have been regulating recommendation algorithms since at least 2022. And over time, they've made, like, a bunch of different rules that, like, we just don't really have in the same way. They've made rules to protect gig workers, for example. We've had some efforts like this with, like, ballot initiatives in California that have tried to kind of reclassify gig workers as employees and things like that. But in the Chinese context, the government and I I kind of understand that this was in response to a big, like, investigative journalism expose kind of thing. I don't know the full story. But first of all, if you heard my thing last time that just delivery is insanely convenient in China, you can get your food delivered super quickly and cheaply. And I also told the story of when I had this little allergic reaction and then told a friend about it. And we were on a bus, and, like, when we got to our destination twenty minutes later, there was Clareton waiting for me. So they are really good at delivering stuff in a non demand way. As all these guys running around on electric scooters delivering all these these packages, food, and whatever else, and it kinda came to the public's attention and to the government's attention through some investigative journalism that, like, it really was getting pretty brutal on these gig workers with the platforms, like, pushing them harder, pushing them to take more orders, not giving them rest, giving them, like, impossible delivery time windows, and then penalizing them if they didn't make the delivery time window. So the government came in and put a bunch of rules in place and said like and again, this goes to their tech integration. Right? Last time I also talked about how when you've called a car in their Uber, which is called Didi, you can see when it's at a red light, you can see the timer counting down until the light turns green. So they have, like, a pretty good ability to analyze what's going on with traffic and know, like, what you can realistically do as a delivery driver. So they now have rules that prevent the platforms from giving unrealistic delivery timelines to the drivers. And now the drivers, like, are expected they were always expected to, like, follow the rules. But in the past, they were sometimes incentivized to break traffic rules because that's the only way that they were gonna make the delivery time that the platform gave them. That's just been squashed. Like, okay. Platforms, you can't do that. You have to give the drivers enough time so that they're not incentivized to break traffic laws to get there and get their on time delivery check, have to give them rest, etcetera, etcetera. They have also put in place in a different domain rules specifically designed to protect the elderly from scams, which is something I think I would love to see your Facebooks of The US take a little bit more of an interest in. Whatever. There's a a long digression into Meta platform governance that we don't have time for today. But there was a story not too long ago that there was a project at Facebook slash Meta to try to dig in on how much of their advertising ultimately at the heart of it is some kind of scam. And apparently, they found out it was, like, kind of a lot. And then if the reporting is to be believed, Zuckerberg shut that project down and was like, okay. We don't really need to look into this any further. The elderly are the victims of a lot of those scams, and I do wish that they were a little better protected by a platform like Facebook. In China, the government has made, again, specific rules for this. They also have, interestingly and I'm again, I'm not sure I favor this, but they do have restrictions on price discrimination. Platforms are not so free in China to charge different people different prices for the same thing as they are in The US. I personally think when we get bent out of shape about price discrimination, it's a bit misguided. I'm not so concerned about it as many people are, but there, they have rules about it.

[1:27:32] They also have requirements around platforms labeling AI generated content when they serve it up to users. If you're a TikTok user, you've seen that they also do that in The US. I'm not sure if they do it exactly the same way as they do in the the Douyin native Chinese app, but you do see that from TikTok in The US. And I would say, honestly, more prominently than I certainly see it on, like, Twitter or I think Facebook too. So that's interesting. And then just while I was there, another interesting set of rules went into effect around AI companionship. Last time I described a bit about the Dobao phenomenon and how in part because we've had a couple generations of one child policy, and so we've got a lot of grandparents where the there's four grandparents and there's only one grandkid. Right? The you've got when you have two generations in a row of one child, that means the one grandkid is the only grandkid for all four grandparents. There's just a lot of loneliness among parents and grandparents. People are turning to Doubou. They're, like, getting companionship from it. And this is, like, potentially okay, potentially problematic. But the Chinese government has made rules about this, and they just went into effect. I think they literally went into effect while I was there. These rules include anti addiction measures, a ban on children using them at all, various ways of popping up reminders that you're talking to an AI, etcetera. And people broadly seem to think that this was reasonable. There was a little bit of a replica kerfuffle online. I wasn't on the forums to read all the comments, but people were saying that, like, the users of Doubou and other companion apps were, like, sad that their their companion was gonna be taken away from them or what have you. Most people seem to think that this was a pretty reasonable set of moves for the Chinese government to take. And one person also said, look. This is really mostly about the big services, the ones that reach tens, hundreds of millions of users that aspire to reach a billion plus users. It's really those and I think to about was, like, a 150,000,000 users from what I recall. The the point was kinda like, they're regulating the mainstream stuff. If you really want a racier AI companion, you can go get that from some, like, other app still. It's really just about cleaning up the main ones and making sure that, like, the thing that most parents and grandparents are gonna use, the kind of equivalent of, like, all the boomers are on Facebook. Like, all the if all the boomers are gonna be on Doubout, we're gonna protect them. We're not gonna let them get confused or taken advantage of or addicted to this stuff. We're gonna try to make sure it it stays reasonably under control. Daubao is still there, and you can go get your if you wanna get romantic or whatever, those apps exist in other places, but you have to kinda seek it out. That was the the perspective that I got from one, I think, very informed resident of Shanghai who who had, I think, very interesting opinions on a lot of things. Okay. All of that to say, there's a long tradition in China of regulating technology companies in ways that we in The United States simply have not done. And that, I think, again, should give you some reason to believe that that trend might continue. And certainly that if we get serious or, like, try to approach them on some sort of deal, that it's not gonna be, a crazy idea to them that they might regulate their technology companies. The government has also put out fairly detailed taxonomies of risks when they put out these policies. They do their homework. Right? They're they're thorough. They map things out. They love a good taxonomy. And one of the taxonomies that I saw included as a risk from AI, the risk that, quote, emergence of AI self awareness and loss of human control. So, again, you see this not just in the speech, not just in the research, but you see it in the policy as an enumerated risk in a broader risk taxonomy. Now let's bring this to the AI era, the LLM era specifically. How are LLM products and services regulated in China today? This is where the CAC is really the main government entity, not the only, but the main one that the companies are working with. They have something called a registry where all the big AI services are listed. You can go to this registry online, and you can see all the services that they have reviewed and approved. And all your big companies and all their AI services are going to be on that list. Though interestingly, not every single model release is on that list. And so that's kind of a echo or probably has some relationship, I I imagine, to what I mentioned a bit earlier around how even the companies that are doing the safety disclosures with their model releases, they don't do it on every single release.

[1:32:35] It seems that the the first certainly, your first release, your first entry into the market is going to be, like, pretty carefully reviewed. And, actually, in terms of, like, Chinese regulators slowing down their AI companies, it has happened. In 2023, when ChatGPT had launched in December 2022, and then we had gbt four in '23, and, like, all of a sudden, wow, the world's, like, waking up to AIs. A bunch of Chinese companies were already working on this. They were not quite at the level where the American companies were, obviously. But, actually, one interesting take I heard on why they were behind is that they just didn't think LLMs were worth the investment. I'll talk a little bit more in the third episode about, like, the history of AI in China and how long it's been a strategic priority and, like, some of the investments that that they've made. But I think kind of in a similar way to, like, how Google sort of had the transformer, had a language model, but, like, didn't really quite know what to do with it at a time when it was, like, hallucinating all over and couldn't do basic math and was pretty useless. Like, you needed somebody with a real vision or even, like, a certain level of ideology to think, oh, we'll just scale through this and all the problems will be solved. Google didn't believe that, and the Chinese companies didn't really believe it either. They were in the game. They were very much, like, paying attention to this line of research. But the way one person put it to me was, it's an awful lot of money to spend to get an AI to write bad poetry. So the the one take at least on on why the Chinese companies were behind as of g p t four, same reason Google was behind as of g p t four. Wasn't that they didn't know what was going on. It wasn't that they couldn't make a language model. It was that they just didn't really think that, like, scaling up this particular line of work and pouring more and more resources into it was really about to give them anything all that interesting. So when that changed with ChatGPT and GPT four in that late twenty two to March 23 time frame. Chinese companies were not they were like Google. Right? They were a little bit behind, maybe caught a little bit kind of flat footed on, like, just how big of a deal this might really be, but they had they had language models. They knew very much, like, they were in touch with the technology. They had their own lines of work going there. And all of a sudden, were like, okay. I guess we better follow suit here. We better scale up, and we better launch some services. And in that moment, the Chinese government was also kind of caught flat footed, pretty understandably. And they sort of said, hold on a minute. My understanding is that in 2023, there was a period of about six months where a bunch of Chinese companies had the language model that they had, maybe are racing to create a little bit bigger and better one, and racing to get into the market in the chat GPT moment. And a lot of those deployments were for a time held up by the national government because they've said, hey. We don't really have our house in order here. Right? We don't have standards. We don't have a process. This technology is, like, clearly a big deal, but it's unwieldy. Probably did have a lot of concerns about content safety. We don't wanna talk about the three t's or whatever else. So they, as far as I understand, told the companies, no. You cannot launch yet. We will get our act together. We will have standards. You will then be expected to meet those standards, and then we can launch when we're all good and ready and we are pretty confident that we at least have a decent sense of what we're doing. So I understand that that period went on for six months, and that wasn't at a time, obviously, when, like, the stakes of who was winning the AI race were as high as they are now or as high as they're likely to be in the future. So I think people could dismiss it if they want to, but nevertheless, you did have the Chinese government slowing down their AI companies, preventing release, denying the public the utility, denying the companies whatever revenue and prestige they were gonna get because they wanted to make sure that they had the situation broadly under control. Now since then, they when they did get their act together, it seems that mostly things have gone pretty well and pretty smoothly. You, as a new entrant to the market, you're gonna have a pretty thorough review of your service. You have to give as I understand it, you have to give your local provincial authority access to your product. I think they're often doing this via just here's an API key. You can try the product. The local government will do its testing first. If they approve, then you go up to the national level and the CAC does their review. If they approve, then you're clear to launch. If if they don't approve, then you've got work to do to satisfy them before you can finally launch your product.

[1:37:41] Once that is done, it's not entirely clear, at least to me, what criteria is used to say we're gonna do that full process again? It doesn't seem like there's been a lot of delays recently. It seems like they've they're now pretty comfortable with their process, and the companies know how to do it. And from what I understand, like, releases have not been delayed nearly as much as they were in that initial period recently. Yet there are multiple entries from individual companies in the registry. So there's, like, sometimes when it seems to rise to this level where it's, like, considered a new thing and it's put in the registry as like a separate thing. But then other times, it's kind of considered to be like an incremental improvement or not such a big step up in capabilities that a somewhat lesser process is deemed sufficient. And that's that's kind of opaque to me. I can say that I there are have been similar things with, for example, Google. Think when they put out their deep research agent, I believe it was, something along those lines, or some pro mode where it was like just paralyzing essentially their best model and getting it to do whatever 10 times threads and then pick the best, something along those lines. They didn't do a totally new model card or safety report for it, and some of the AI safety community were upset about that. But their point of view was like, well, look. It's the same model. It's like, it's kinda got the same worst case as before. It's just that, like, it should be performing close to its best case more often because you're in fact doing 10 and picking the best or whatever exactly is under the hood. But they were kinda like, it's not fundamentally a totally new level of capability that requires us to do all this stuff again. I think something like that is going on in the ongoing dialogue between the AI companies and the CAC and other regulators in China. What what I kinda heard from the companies that I had a chance to talk to about this was like, we're always in touch with the government. It seemed to be at least weekly and potentially for some people at the companies, like daily. In close contact, close collaboration, I get the sense that the regulators are obviously, as we've just discussed, they're not afraid to impose costs on the company. They're not afraid to do something that will hurt their profitability. They're not afraid to do something that will cause them delays, but they want the companies to be successful is probably the understanding that I have. So it the sort of legitimacy of the regulators. Now would they tell me if they thought the regulators were illegitimate? I did have the one time I talked about last time where a professor found himself kind of caught in a catch 22 thing with respect to having a drone in Beijing. And he was pretty candid about saying that was kind of a ridiculous bit of bureaucracy that he found himself not, like, majorly inconvenienced by, but but inconvenienced by. So there's at least some willingness to say if you think that the government's made a mess of something. I'm not sure that that would extend to companies telling somebody like me that they think the AI regulators are making their life too difficult. The the sense that I got was, like, that the legitimacy of the regulators seemed to be quite well established, that the company sort of expected and accepted that they were gonna be in regular contact with the government, and that for these incremental releases, they were kind of in such an ongoing and close dialogue that not everything had to go through some cumbersome process, but occasionally, when it was a big deal, then something would. And as far as I could tell, everybody seemed to be feeling pretty good about that. And, again, k three came out right at WAIC. Model seemed to be coming out pretty fast from Chinese companies. Jipu or ZAI, I cross posted an episode from China Talk with a guy there who leads kind of their go to market partnerships, I guess, would be maybe a good way to say it. Definitely worth going back and listening to that one. It's maybe six to three months old at this point. One of the things that stood out most to me about it was just how fast they're going from models finishing the training process to release. So it does not seem like they're being dramatically delayed on a consistent basis by the regulators. It seemed like the regulators wanna have a good process, wanna have command of the situation. Their primary duty is to the national government or the party. But part of the way that they impress their bosses is, yes, like keeping a good clean sheet on safety issues, but also having the companies in their jurisdiction be successful and not be unduly delayed.

[1:42:41] I think their their incentives are actually pretty good as far as I can tell in that regard. And while they have demonstrated that they're willing to impose costs or even significant delays, it doesn't seem like that's happening on a regular basis. Other things that jumped out at me, the the Chinese government is definitely paying attention to trends. Again, the report for this a lot of this comes from the state of AI safety in China report from Concordia. But there was a in January, I think I maybe already mentioned this, a Politburo study session where she talked specifically about risks of technological loss of control. I mean, imagine our government having a study session. Look at what our cabinet meetings look like. It seems like, at least in some of the cabinet level meetings in the Chinese government, they're actually, like, studying important topics. Like, imagine that. Doesn't mean they're always gonna come to their conclusions, obviously, but, boy, we could stand to do a little bit more of our own homework here, Certainly, the executive level sometimes, I feel like. In also in January, there was a draft cybercrime law that was set to require AI companies to monitor for and report bulk generation of malicious code. This jumped out at me in preparing this because here we are in this moment where we just had OpenFace, and Anthropic has kind of reported similar things. They've got their models like hacking out of sandboxes, hacking into other people's systems. It sure seems like the monitoring on that wasn't great. I don't know that the monitoring is great on the Chinese side either, and I I don't think that this as far as I know, I don't think that this law has actually gone into effect as Dean Balt talked about in his last appearance on the podcast, there is lobbying in the Chinese system. There is when these draft laws come out, there is the opportunity for the companies to go talk to the government. And there have been instances where I think specifically around, like, the accuracy of outputs from AIs. The original draft was, like, saying, like, your AIs have to be accurate. And the companies went back and said, like, look. The nature of this technology, it's I can't really promise you that all the time. And they backed off. They they reduced their expectations because I think they made a I think they understood what the companies were telling them. And I think they made a considered kind of cost benefit analysis that, like, the upside of this technology is greater than the damage likely to be caused by hallucinations. That's different, of course, from, like, sensitive third rail topics. But just, like, hallucinations can't prevent them. Okay. We get that. They pulled back on that draft law. So we'll see if this cybercrime law goes into effect as drafted or if there's some pushback, but seems that at least in the initial form that it was put forward, it would require companies to do this sort of monitoring. The sort of monitoring that if it had been applied to the recent OpenAI and anthropic unreleased more powerful models might have prevented some of these hacking into third party systems. So I thought that was pretty interesting. And then what was the biggest trend of AI recently? What what kind of really took the agent moment mainstream, OpenClaw, of course. And it wasn't it was a pretty fast turnaround. February of this year, just weeks after the real kind of OpenClaw fever hit everybody, they had put out some warnings about, quote unquote, relatively high security risks in some versions of Open Claw. So you've got the national government in China paying attention, spotting trends like open claw, digging in, and issuing, in that case, just a warning. It's clear that they are very engaged with what is going on. Now an interesting question that I don't know the answer to is what would happen in China if there were an Open Face like incident where a company lost control of its AI, not maybe entirely obviously, but enough so that it was able to go hack third party systems for days and steal information or cause whatever havoc had caused. I don't know about that. I don't know how that would be dealt with in China. Again, this is the Cyberspace Administration of China. My sense is that as the main regulator, it would probably be their jurisdiction, their mess to clean up, unless perhaps it was deemed to be such a big deal that there was kind of loss of confidence in the CAC itself, in which case you could imagine some kind of bureaucratic reshuffling or reassignment or a change in structure of the government. But I think it would be the CAC that would be kind of in charge. I wouldn't be surprised if saw companies get more of a slap than we've seen certainly OpenAI and Anthropic have had so far. Right? I mean, if I guess the careful way to say it is, like, if a human did what OpenAI and Anthropic models have reportedly done, I believe it would be a felony.

[1:47:48] It doesn't seem like we're gonna have and by the way, Hugging Face did have law enforcement involved before they knew who it was. Right? So this was it did rise to the level of, like, authorities were called. Gonna happen? Is there gonna be any accountability? Is there gonna be any, like I don't I don't necessarily think that there should be, like, criminal charges filed against, like, individuals at OpenAI. I'm definitely not recommending that. But what is gonna happen? I don't know. It seems like maybe the most likely thing right now is nothing the governmental level or at least at the, like, law enforcement level of government. I strongly suspect something more serious would happen in the Chinese system, although, obviously, I can't prove that. But they they're not afraid to come down hard when they feel like things have got out of control. We've seen examples of that with, like, Jack Ma who dared to criticize the government's policy with respect to financial regulation in his company, and he was kind of sidelined for a while. There are some other interesting examples of that where, like, education businesses were because the Chinese system is so competitive for this national exam to get into college, and kids still study really hard over there from what I can tell. But it used to be even worse, and the government, like, said no more of this sort of extra private tutoring or at least, like, a dramatic reduction in it that came pretty suddenly. There's also one with games and social media where, like, you're only as a kid, you're only allowed to play games at a certain very limited time of the week. I think it's, like, Friday and Saturday, Sunday for a couple hours each or whatever. Now people did tell me that kids get around that by using their parents' devices and have the parents do the sign in for them. It's not like they have perfect control over there. But the way one person put it to me was the Chinese government feels that it can put the genie back in the bottle. If something does happen that that makes them that gets them spooked, right, that makes them sufficiently uncomfortable, they are able to take pretty dramatic and swift action to tamp it down. You see this in terms of censorship all the time on the Chinese Internet. There was a there was that little incident where, like, a small plane crashed into a building in Beijing or whatever, and apparently, that was totally removed from the Chinese Internet. People I was talking to knew about it. They were using it as an example of the kind of thing that is censored from the Chinese Internet. Because also, as I talked about last time, like, comfort with contradiction being sort of an interesting part of Chinese culture. Like, the people that were talking to me about it were, like, both kind of annoyed that they had to use their VPNs or go to international media to to learn about this thing, but also felt like it was probably at least defensible that the government wouldn't want everybody to hear about it because they they don't wanna create panic or whatever. Anyway, they can do these things. Right? And so the the point was like, even with an open weights model, right, they feel like even if an open weights model was released and it proved to be dangerous, they could still keep it under control. Now I don't think that's a assumption that translates to the rest of the world. I don't think that would translate to the American context. I don't think it would translate to most other governments, frankly. So I think this is one area in which, arguably the Chinese government is not thinking as much as they should about their impact on the rest of the world when they release models open weights. And I think that they look at their own situation and especially given, and again, trillions of parameters and, like, serious hardware required, they feel like even if a model is out there open weights, we can crack down if we need to. Right? We're in close contact with all these companies. We can make them do stuff. We're we're in close contact with, like, cloud companies and inference providers and whatever. Like, we can tell them, thou shall not run this particular model anymore, and they feel like they can scrub that off the Internet and make it inaccessible. And I would honestly go as far as to say, even for people trying to do it at home, if you were gonna try to assemble your own rig in your basement or whatever the case may be, if you were doing any serious volume, if you were trying to run some sort of illicit unreported inference company or something like this, maybe you could do it just for yourself. Right? If you're just trying to get your own precious few queries, you could probably slip under their ability to detect. But if you are trying to run some small but nontrivial registered inference business, I would guess that they would even be able to track that down just by your electricity usage. They first of all, they've banned crypto. Right? So they they've done work to look at, like, how do we understand the flow of electricity and what might be a problem for us? And then you gotta you would be shocked, certainly as an American who has, like, a utility company that kinda sucks and has, like honestly, my utility company is not so bad these days, but, like, traditionally terrible service, long wait times, blah blah blah blah blah.

[1:52:54] The state grid company in China was a remarkably prominent presence, including at WAIC where they had a major booth. I also visited one academic group where they were working on some some robot technology, and they had a sort of high voltage line kind of fake environment setup with, like, obviously, not the actual high voltage to it, but, like, the wires and the sort of the normal rigging that you would have for all these wires. They came and helped set this thing up in this one professor's lab space so that this professor could try to get his robots to do things that might eventually be useful for the state grid company. So I think even if you were just to imagine, like, okay. This model was released to open weights. We think in the West, like, oh, you can never take it. The Internet never forgets. The Chinese Internet does forget. They would be able to take it down. I think they would be able to track down illicit inference businesses running at any nontrivial scale that's like, wait a second. Why are you using 10 times as much electricity as your neighboring apartment? But I think all of that is a way to understand why, at least for now, the Chinese government isn't so afraid of open weights models. Because even if something like this were to happen, as long as it's not totally catastrophic and irrecoverable, they feel like they can, in fact, put the G and A back in the model. They are getting serious about labor market impact. Here in the West, have, like, the Anthropic Institute and the Open AI Foundation are, like, hiring economists and whatever to do this kind of stuff. And, of course, we've got some academics that are turning their attention to it. And, of course, politicians love to talk about it, but usually in a pretty substance free way. As far as I understand, in China, the government is setting up its own monitoring. They're creating retraining and job transition programs. And there was even this one report that said that the Chinese government has said that they won't you you will not be allowed to fire people from your company because, like, AI has made them unnecessary. That's if they really try to hold to that, that'll be pretty extreme. If my guess is it will be so costly from a, like, undermining dynamism standpoint, and they really are super dynamic in the private sector as I think is all well very well understood at this point. My guess is that that they would have a hard time holding that, but, again, maybe not. Right? Like, they did lockdown for COVID for a long time, seemingly longer than most observers thought would make sense. And they might be able to hold the line on you can't fire people because they've been made redundant by AI for a lot longer than we might think. And what would that impose in terms of costs on their AI sector? I think pretty significant. Right? I mean, if you can't if as a business, you can't realize cost savings from AI implementation because your headcount has to stay the same even if you're getting AI to do things that people used to do, well, that definitely reduces your incentive to do that transformation, and that in turn reduces the revenue that the AI companies are gonna be able to capture. I would bet that they don't hold a super firm line on that for a long time, but, like, this is kind of where the margin is right now in the Chinese context from what I can tell. They have a history of imposing delays on their LLM companies and their and their chatbot releases specifically. They are paying attention to things like OpenCLaw and releasing guidance around it. They are doing Polybros study sessions. They are doing all all the same kinds of research that we are doing, and they're even entertaining policies like you can't fire somebody because they've been made redundant by AI. To conclude the sort of government section, they're definitely doing a lot more in the government sector than the US government is doing. No question about it. Whether that's better or worse, obviously, depends on your perspective. But if you're an AI safety person or if you are if you're in some debate where somebody says, well, if we do that, we'll cede the race to China. They'll never slow down. They'll never take this stuff seriously. They don't care. I think that's hopefully at this point easily refuted. Two more things to close us out.

[1:57:45] One, a complaint that many people have with the Chinese companies is that those that signed on in Seoul at one of the AI safety summits to publishing risk frameworks and doing more kind of consistent updating of how their models are performing against these frontier risk frameworks. The complaint is that, hey. They never followed through. We never actually got those risk frameworks that we were promised. And what's up with that? Like, isn't that just another example of bad faith from the Chinese side where they commit to something and then they don't do it? Or they and you do hear this a lot, especially with the at the government level of people that have done negotiations on various topics with China over time. The the there is a certain jadedness or cynicism that has set in where the Western negotiators are like, well, they'll say it, but they may not do it. Now I think this is, like, clearly bad, and I wish these companies had done this, especially since they made the commitment. So I I'm definitely not here to excuse it. But if I was just gonna try to offer what I think the story from the Chinese side would be about, like, why this has happened or why they maybe don't feel the need to do it in the way that they, at one point, thought would make sense for them first of all, by the way, I would just say, look no further than Anthropic for a company that used to have a much tighter responsible scaling policy that had all these if then commitments that they abandoned because we can't really meet those commitments, and we can't stop racing because then we'll just be seeding the future to the bad guys, whether the bad guys are OpenAI or China. We're gonna move to a just trust us regime. We still put that forward as responsible scaling policy, but the old responsible scaling policy is kinda no more. Now it's like, as we put it memorably, now it's trust us. So we've got some of that going on too. Right? It's not a direct analogy because we had the policy and they kind of mostly rescinded it or largely rescinded it. In the Chinese case, they committed to making one. They never did. My guess is that the people at these companies feel that the government is really taking the lead on questions of AI safety. The government is the one setting the setting the standards. They are the government is doing these checks. The government has this model registry. The government is they're in communication with the government on a weekly, if not daily basis. And as a result, I think they probably, in many cases, think that it's not really their place to come up with safety standards. Like, the division of labor that they seem to have and that they seem to believe is legitimate and that seems to be working well enough for them so far is that's kind of the government's job. So the companies are not really in a position, and it might even be seen as kind of disrespectful or overstepping to put forward a safety framework when the government already has one. Again, I'm not sure that's the the full explanation, but I I that would be very consistent with everything that I observed and everything that I heard. And I think also that probably a decent kind of summary with a little bit of filling in the blanks about what their expectations will be going forward, the company's expectations, is that the safety standards are gonna get tougher. Again, the 45 degree line. The capabilities are definitely growing, And so they're all kind of on the same page society wide that like, yeah, it's just plain practical that we have got to have safety measures that get better roughly in some kind of vague conceptual sense, roughly the same pace that the capabilities themselves advance. Everybody seems to be bought in on that, and the division of labor has kind of settled into this situation where the government is taking the lead and the companies are responsible for hitting that standard. I think they expect that that standard is gonna rise over time, and I think they're, like, totally fine with that and, like, totally prepared to do what is asked to and it might be difficult, and they might face delays, and they face delays, as we know, in the past. So I I don't think that they are I don't think that those safety standards are necessarily less than the western ones. Although, as we talked about at the beginning, the actual deployed safeguards are currently less than. But as we as we project into the future and a lot of we talk about, like, what's the gap, right, between the West and the Chinese models? And it's like, whatever, nine months depending on how you wanna measure it. And it's obviously hotly debated, but I think nine months is like a pretty good center of the distribution of credible answers that you'd get. Well, a lot has happened in the last nine months. Right?

[2:02:52] And we got Claude four five Opus roughly nine months ago, and that was the first time that, like, agents really worked. Now we're getting some Chinese models where agents really work. Think that in another nine months when they're dealing with these things that OpenAI and Anthropic are currently dealing with, I think that the standards will probably rise to meet that. And I think I would expect that you would see a significant closing of the gap in terms of the deployed safeguards. Open weights models may be another thing. Again, I think the Chinese government feels like within their borders at least, they can take an open weights model offline. It's not it's not irreversible for them in the same way that it would be for us. I do think the impact that they may have on the rest of the world is really important, and I'm not sure they're taking that as seriously as I would hope they would, especially when it comes to, like, biorisk that even could come back low back on them. But that's only something that we'll need to watch. As of now, though, I think that everybody there kinda feels like we got work to do. The 45 degree line is the right mindset. We're maybe not quite where we need to be, but we also kinda know that the American companies have already explored what happened to this level of capability, and it wasn't anything too bad. And it's not like their systems are never jailbroken. And so at least for now, we're more focused on catching up, making sure we hit the standards, but it's not like our place to go out and try to opine publicly and broadly about, like, what safety standards should be. We're in the business of catching up, hitting the standards, whatever the government says they are, fully expect they're gonna get more demanding on us as we go. And that's just the that's just life in the big leagues of the AI game. My sense is that that's pretty much where the even like the smaller Chinese AI startups that are doing the least right now. I think that's probably a pretty good summary of where they are and how they're thinking about it. Okay. Final thing. And this is maybe the biggest gap or, like, difference. I went looking for an analog in the Chinese system, and I did not find it. If you have an analog, I would love to hear about it because I did look I asked a number of questions of a number of different people about this topic and never really got much of a response other than like, yeah. I don't really have anything for you. So the topic is, is there such a thing as alignment with Chinese characteristics? All if this episode as a whole is AI safety with Chinese characteristics. And we've seen differences certainly, but a lot of similarities across all these different aspects of AI safety. What about alignment? Is there such a thing as alignment with Chinese characteristics? I was inspired to ask about this in part because I was using one of the Chinese AIs at the Temple Of Confucius in Beijing and just trying to learn a little bit about Confucius. Who was he? What's the deal? When did he live? One of the facts that I was taught by the AIs was that to this day in the region that he comes from, his descendants who are now 79 generations hence 79 generations. Isn't that amazing? What the AI told me was that his descendants still to this day, 79 generations later, still identify as his descendants and perform certain rituals in his honor all this time later. And they got me thinking, boy, if we could project our values through 79 recursively self improved generations of AIs, we'd be doing really well. Right? We'd be doing, like, much better than I think we can reasonably expect to do if we just race into a recursive self improvement mediated or generated intelligence explosion. So there's something there with the Confucian tradition that has made values and sort of respect for what came before quite durable. 79 generations. So I was thinking, boy, is there a way that this could translate into a sort of AI constitution or an alignment target? We've got the Claude constitution. Could there be a Confucian constitution for AI? Unfortunately, for my enthusiasm on this topic, nobody seems to be working on that as far as I can tell. I didn't I did not get one pointer. One professor told me when I asked him about this, he said, it's an interesting idea, but we are probably the generation in all of Chinese history that is the weakest on this traditional philosophy. He said, you gotta remember, we're all engineers.

[2:08:02] The humanities, that's not where people were going. Right? I mean, the Chinese leadership at the political level is all engineers. I believe Xi himself is a chemical engineer. He said, we're all engineers. None of us really studied philosophy. None of us are schooled in the Confucian tradition in the way that previous generations were. And it's just not really in our wheelhouse to think that way. We're really practical. We look at problems as they present themselves, and we try to find solutions. And we do care about making sure that AI is good for people, but there's not really a lot of that kind of thought going on. So this, I think, is a really interesting opportunity still. But for now, the AI safety community in China is much more on the OpenAI side of the cordurability versus character debate. They're about having clear rules, having AI follow those rules, trying to make that as reliable and consistent as possible. I did not find and I would love to hear about it, if you know of any, but I I did not find much in the way of imagining what it would look like for an AI to grow into a wisdom tradition that they have that that an AI might be able to to embody or realize or bring to its its highest potential form in the way that the Claude constitution seems to imagine Claude growing into the greatest virtue ethicist of all time. Just couldn't find anything like that in China. So I would be really interested to know if you have any pointers. And I think maybe that's something that people should work on even in the West or across obviously, there's there's many Chinese and Chinese American people here that would have a a much more credible angle on it than I do. But I do feel like there's something missing there that could be a pretty interesting opportunity. If we imagine that a future is that a good future might be made up of, like, multiple powerful AIs that are in some sort of ecological style balance with one another, then having different wisdom traditions to base them on, to me, seems like quite a good idea. And right now, as far as I can tell, that is pretty much greenfield. Okay. That does it for today. I would welcome your feedback, your commentary, your critical commentary, your pointers to anything that I have missed. But if nothing else, hopefully, this serves to give you all the ammo you need to push back whenever people say, China doesn't care. China will never slow things down. China's not gonna stop their companies. I think that the truth is quite the opposite, that they do clearly care, that their research community is very much engaged, that their product companies still have some work to do, but that their government is pretty well on the ball. I think they're trending in the right direction even though and we can say the same for ourselves. There's a lot of work left to do. So part three will come as soon as I can get it ready for you. That's gonna be focused on The US China relationship and what, if anything, we might do to make it better and to start to work together on some of these AI issues. Probably allow myself to be this one is a little bit more just kind of fact based reporting. That one, I'll probably allow myself to be a little bit more of a idealist dreamer. So stay tuned for that coming soon. But for now, this has been AI safety with Chinese characteristics, and I thank you for being part of the cognitive revolution.

Outro

[2:15:36] If you're finding value in the show, we'd appreciate it if you'd take a moment to share it with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts, which is now part of A16Z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast.ing. And thank you to everyone who listens for being part of the Cognitive Revolution.


Great! You’ve successfully signed up.

Welcome back! You've successfully signed in.

You've successfully subscribed to The Cognitive Revolution.

Success! Check your email for magic link to sign-in.

Success! Your billing info has been updated.

Your billing was not updated.