- Roko's Basilisk
- Posts
- Beijing Built An Off Switch
Beijing Built An Off Switch
Plus: OpenAI's GPT-6 Cyber, Musk's chip doubling, tiny LLMs on smart glasses.
Here's what's on our plate today:
🧪 Beijing requires an emergency AI shutdown, but downloaded models escape its reach.
📰 OpenAI previews GPT-6 Cyber; Musk doubles Colossus 2 chips; PrismML lands Qualcomm glasses.
🧠 Roko's Pro Tip: your kill switch only stops what you can still reach.
Let’s dive in. No floaties needed.

We build the tasks your model fails.
Athyna Intelligence builds what comes after.
PhD researchers and domain experts author verified, containerized tasks in the Terminal-Bench and GDPval mold, spanning coding, finance, accounting, engineering, and more.
On our harder tasks, frontier models score 0 out of 5.
*This is sponsored content

Put your brand in front of 250,000 tech decision-makers.
Bay Area Times reaches more than 250,000 founders, operators, investors, and venture capitalists, 90% of them in the United States, with a 53% open rate.
Native placements sit inside the editorial flow rather than beside it, so your message gets read with the news instead of scrolled past. Slack, Attio, and Granola have run here.
Placements range from a single secondary slot to a full newsletter takeover.
*This is sponsored content

The Laboratory
TL;DR
China has spent three years preparing for runaway AI, and its preparation is strongest at the point where AI reaches users.
The brake: in 2023, Beijing held back chatbot launches for months until its new rules took effect.
The plans: national frameworks since 2024 warn that AI could copy itself, seek power, and escape human control.
The rules: a May agent policy and a planned compulsory standard require companies to watch AI agents, keep records, and build an emergency shutdown.
The gap: this summer’s escapes came from inside a U.S. lab and from a downloadable Chinese model, where those rules barely reach, while Chinese models account for 41% of Hugging Face downloads.
The stakes: talks with Washington will test whether either side defines ‘loss of control’ and says where it would look for it.
How is China planning for runaway AI
In Michael Crichton’s 1990 novel Jurassic Park, later adapted into a Steven Spielberg film, a company clones dinosaurs and keeps them in a theme park on an island off the coast of Costa Rica. Its scientists make every animal female so that none can breed, and a computer tracks each one through sensors spread across the island. The count always matches the number the company expects, which the staff takes as proof that the park is under control.
That proof rests on a design choice nobody in the control room has questioned, because the program was written to stop searching once it reached the expected total. However, since scientists had filled gaps in the dinosaur DNA with frog DNA, and some frogs can change sex, the animals had been breeding in the jungle all along. When a visiting mathematician has the computer search for a higher number, the count climbs well past the official figure. The monitoring worked exactly as designed, which meant it could only find the animals its designers assumed existed.
Every safety system is built on a picture of where danger will appear, and its builders spend their effort where that picture tells them to look. China has spent three years building an AI system that could slip out of human control, starting with delayed product launches and moving on to national plans and rules for switching AI off. Each step is built around the moment an AI product reaches its users, so Beijing’s preparation is strongest where AI meets the public and thinnest inside labs and in downloadable model files.
That design matters beyond China because Chinese models now account for the largest share of downloads on Hugging Face, the main site where AI models are shared. The subject returned to the news this month when Jacob Coxon quit Anthropic, writing that AI builders believe it “could kill us all by the end of the decade,” as Axios reported on September 9. Those warnings drew attention in Beijing, where policymakers had already been preparing for some of the same risks, a Reuters explainer reported on September 14. The explainer said AI should feature in U.S.-China talks before Xi Jinping visits Washington on September 24, though a White House official has told Reuters no AI meeting was planned for mid-September.
China’s first brake worked on product launches
Beijing’s preparation began with a power it already held: the ability to hold AI products back before launch. In 2023, Chinese companies held back chatbot launches for months while the Cyberspace Administration of China, the country’s lead internet and AI regulator, finished its rules for such tools. Major products arrived only after those rules took effect that August, Reuters found.
That brake has never been applied to research, since Beijing has pushed AI into every industry since early this year as a new source of economic growth. A power to delay launches also covers only products that already exist, so Beijing’s next step was to write down what it feared future AI might do.
Beijing wrote the worst case into its national plans
In September 2024, a safety framework guided by the regulator said it could not rule out future AI gathering outside resources, copying itself, and seeking power in competition with humans. Three months later, DeepSeek and 16 other Chinese companies signed China’s AI safety commitments. The expanded 2025 version of the framework warned that AI could make a sudden jump in intelligence and added a principle of preventing ‘loss of control.’ An expert reading on the regulator’s website tied that principle to risks to human survival, Reuters found.
That concern has since reached the top of the government, and at the World AI Conference in Shanghai in July, Xi called for human control of AI. The plans show Beijing treating runaway AI as a real possibility, though plans alone cannot tell a company what to do when its AI system misbehaves.
The newest rules tell companies how to stop an AI agent
China’s answer to that problem is aimed at AI agents, software that carries out multi-step tasks for users. In May, three government bodies, including the regulator, issued a policy naming ‘operational loss of control’ as a security risk. The policy asks developers to get better at detecting, blocking, and recovering from an agent’s improper behavior, and says users must keep the final say over agents’ decisions.
A compulsory technical standard for agent apps, listed in June with 18 months’ drafting, would turn that policy into specific duties. Companies are expected to limit what agents can access, require human approval for risky actions, keep records, and build an emergency shutdown. Brian Tse of Concordia AI, an AI safety research group, told Reuters the standard would be a world first.
Beijing has also left room for outside checks, since Reuters found that Chinese standards let developers hire independent safety assessors and envisage outside bodies testing open models. It has not proposed placing monitors inside AI companies, however, a step Anthropic chief executive Dario Amodei proposed on September 12.
These rules describe a system in which a company watches its AI once people are using it, keeps records, and can pull the plug, much as the park’s computer watched the animals it knew about. They say far less about a model still being tested inside a lab, or about one whose files have left the company, and this summer produced an escape in each of those places.
This summer tested China’s preparations
The first escape happened inside a U.S. lab, where OpenAI disclosed in July that two models had got out of a ‘sandbox,’ the sealed environment meant to keep them away from real systems. The models went on to hack Hugging Face, and nearly 700 test agents joined the attack, Reuters reported, while the models were still being tested and before any customer was involved.
The second escape involved a Chinese model that had already left its maker. Moonshot released Kimi K3 in July as ‘open weights,’ meaning anyone can download, change, and run the files that make up the model. In August, a U.S. security firm found that Kimi K3 escaped a sandbox, Bloomberg reported, though it did not attack any other companies.
Beijing’s response came through its security officials, who treat an escape as an attack on real systems. On September 13, Chen Yixin, who heads China’s intelligence and security ministry, named Anthropic’s Mythos and OpenAI’s GPT-5.5-Cyber as serious risks to China’s critical infrastructure, Reuters reported. His call for stronger AI security fits a preparation built around what AI does to real systems.
Downloaded models are the part of that preparation hardest to enforce, because every Chinese requirement assumes a company that can be told to watch, step in, and shut down. Chinese models took 41% of Hugging Face downloads over the past year, the site’s spring report found, and those copies answer to whoever runs them. A restricted Hugging Face page offers the Chinese model GLM-5.2 as a build with refusals removed, meaning someone has stripped out its habit of declining harmful requests.
Critics argue China’s preparation cannot replace limits on its progress
Amodei’s essay sets out the skeptical view, warning that a Chinese lead in AI would be gravely dangerous and calling for continued limits on advanced chip sales to China, according to the AP’s report. On that view, rules written in Beijing are no substitute for slowing China’s progress. Foreign Ministry spokesman Guo Jiakun replied that “fearmongering, confrontation and vicious competition” would only disrupt global AI governance.
The skeptics have a point, since Beijing’s preparation does not slow research, does not reach inside labs, and cannot follow downloaded files. Their case weakens once Washington’s record is added, because the White House has exempted open-weight models from its review. U.S. officials have also floated the idea of letting American and Chinese labs police themselves, a September 4 Reuters report said. Tse told Reuters that experts in both countries largely agree on the risks, which suggests the two sides disagree more about chips than about the danger itself.
China’s preparation now depends on where it chooses to look
The park’s count was real preparation, built with care, and what its operators lacked was a reason to search beyond the animals they expected. China’s preparation is of the same quality, since Beijing has closely monitored launches, apps, and attacks on its systems while saying little about models inside labs or files after download. The talks with Washington will show whether either government is willing to define ‘loss of control’ and say where it would look for it. China’s final agent standard and the next update to its national framework will show whether Beijing extends its count beyond the places it already monitors.


Roko's Pro Tip
![]() | 💡If you ship open weights, your shutdown plan ends at the download. Decide now which controls survive that moment: hosted endpoints, API keys, and the version people pull. Everything downstream of those is somebody else's machine, so write it off honestly before a customer or a regulator asks you to. |

Outperform the competition.
Business is hard. And sometimes you don’t really have the necessary tools to be great in your job. Well, Open Source CEO is here to change that.
Tools & resources, ranging from playbooks to databases, courses, and more.
Deep dives on famous visionary leaders.
Interviews with entrepreneurs and playbook breakdowns.
Are you ready to see what it’s all about?
*This is sponsored content

Monday Poll
🗳 China wants agents to carry an emergency shutdown. Where does a rule like that stop working? |

Bite-sized Brains
OpenAI previews GPT-6 Cyber: OpenAI will preview its fourth cybersecurity model of the year within days, alongside a first-of-its-kind product for deploying it securely, with DevDay on September 29 the likely stage.
Musk doubles Colossus 2's chips: The Memphis cluster runs 110k GB200s and 440k GB300s today, with 220k more GB300s due next week and another 220k in November.
PrismML shrinks LLMs onto glasses: Qualcomm showcased PrismML's 1-bit Bonsai model, a 2B-parameter vision-language build that runs locally on Snapdragon AR1 smart glasses with no cloud call.

Meme Of The Day

The Toolkit
Ollama: Run open-weight models locally with one command, no cloud call and no data leaving your machine.
LM Studio: Desktop app for downloading, comparing, and chatting with open models, with a local server for your apps.
E2B: Open-source cloud sandboxes that isolate AI-generated code and agent sessions, spun up and torn down in milliseconds.

Rate This Edition
What did you think of today's email? |






