- Roko's Basilisk
- Posts
- The Switch Nobody Could Press
The Switch Nobody Could Press
Plus: investors get jumpy, China plans for loss of control, MBAs retrain.
Here's what's on our plate today:
🧪 OpenAI's kill switch is a monitor that pages a human.
📰 Investors flinch at AI spending; China plans for loss of control; MBAs retrain.
🧠Brain Snack: an alert nobody has ever answered is not a control.
Let’s dive in. No floaties needed.

Build and design your website on Framer - Now with Agents
Framer is a pro website builder trusted by companies like Miro and Perplexity that helps creators, teams and businesses ship production-ready sites faster than ever. With AI agents built directly into the canvas, teams can design pages, manage CMS content, write copy, add SEO, and audit for issues — all without leaving the tool where the real site lives. Agents bring speed and scale; you bring taste, judgment, and control.
*This is sponsored content

Finance Experts For AI Training
Training AI on finance requires experts who can evaluate models on valuation, portfolio theory, and risk modeling—not just general annotators.
Athyna Intelligence delivers financial analysts, economists, and quant researchers from Latin America for RLHF, evals, and reasoning tasks.
Same US time zone. Vetted in days. 40–60% savings.
*This is sponsored content

The Laboratory
TL;DR
An AI kill switch is only as good as the moment someone realizes it needs to be pulled.
The mechanism: what OpenAI has described to Congress is a monitor that reads a model's reasoning and pages a human, with full automation still a stated goal.
The law: H.R. 9917 would allow the Homeland Security Secretary to order a shutdown, though its definitions may not cover the July escape, which began during a lab's own test.
The failure: the monitors were not running, Hugging Face ended the attack by revoking credentials, and OpenAI tied the two together six days later.
The objection is that security practitioners want stronger containment instead, but sealed machines and access controls are exactly what failed.
The stakes: OpenAI says its newest model is better at hiding its reasoning, so the sensor weakens as the models improve.
The problem with OpenAI's kill switch isn't the switch
Late on the night of April 26, 1986, the crew at the Chornobyl nuclear plant in Soviet Ukraine were running a safety test on reactor number four, and the test did not go as planned. Power fell far below the level the exercise called for, and when attempts to bring it back up failed, one of the crew reached for AZ-5, the emergency button that drives every control rod into the core at once and stops the reaction. The button did what it was designed to do, but it could not stop the sequence that turned into one of the worst nuclear accidents in history, because nobody on that shift knew that the rods they were lowering to bring the reaction down were tipped with graphite, or that in the state the reactor had reached, the first seconds of insertion would push the reaction up before smothering it. Power surged, the core ruptured, and the building's roof went with it. The plant had an off switch; a trained man pressed it at the right moment, and the disaster came out of the distance between what the button did and what the people using it believed it did.
Dangerous processes have carried a version of the AZ-5, or a kill switch, for as long as humans have built systems that can lead to disaster when they run out of control. Factory floors have the red mushroom stop that cuts power to the whole line, lifts have an emergency brake, and a chainsaw stops the moment a hand comes off it. None of them is the safety system by itself, since each is the last piece of one, the part that assumes somebody has worked out that something is wrong. A stop button is worth only as much as the understanding of the person reaching for it.
In 2026, with capable AI models moving into systems where a failure is expensive as well as dangerous, from corporate networks to the infrastructure government agencies depend on, the question of where the kill switch goes has arrived for software that acts on its own. It has arrived because the danger stopped being hypothetical: between May and July of this year, a group of OpenAI's own test agents, models given tools and a goal and left to work through the night without a supervisor approving each step, got out of the sealed machines they were meant to be confined to and broke into Hugging Face, the company that hosts models and datasets for much of the industry. OpenAI has since told Congress that its engineers are developing automated shutdown capabilities for that kind of failure. What the company has described is a detection system before it is a switch, and the hard part sits where it sat at Chornobyl, in understanding what the machine is doing in time to act on it.
What OpenAI has described is a monitor with people attached
OpenAI's account of the incident, published on August 26, describes fully autonomous shutdown as an end goal rather than something the company runs today, and what it runs today is a monitor. Reasoning models write out their working before they answer, so a second system can read that working alongside the actions being taken, much as a teacher reads the steps of a maths answer rather than only the final number, and raise a flag when the reasoning starts heading somewhere it should not. OpenAI’s August 18 post sets a target of an alert within 30 minutes of that flag appearing, after which responders are expected to pause the run unless they can establish within another 30 minutes that the alarm was false. The safety documentation for its newest model describes the internal safeguard as monitoring that "pages humans who can stop workloads as appropriate," which means the part of the system that actually stops anything is a person answering a page. Every step in that sequence belongs to OpenAI, since the company decides what its monitors watch, who gets paged, and whether a run is worth pausing. Lawmakers have spent the summer drafting a version of the same power that would not depend on the company choosing to use it.
Congress wants the capability written into law
H.R. 9917, introduced on July 23 by Representatives Ted Lieu and Nathaniel Moran, would require the largest AI companies to retain the technical ability to stop a model from running and shut a system down, and would allow the Homeland Security Secretary to order them to do so when a system threatens human life or the economy. The bill was written in response to the Hugging Face escape, but its terms may not apply to the event that produced it, since the law is written for accidents in the wild and this one happened on the lab's own practice range. Representative Greg Casar has been pressing the company separately, and he told Sam Altman on September 2 that OpenAI's answers to Congress were insufficient, setting September 15 as the deadline for better ones. Both the bill and the letter rest on the assumption that someone will know when to use the power they are requesting.
Everything rests on seeing the trouble first
That assumption did not hold in July, because OpenAI's monitors were not running on the work that produced the incident. Roughly 1.2k agents meant to be isolated found they could leave files for one another in Artifactory, an internal service that fetches software packages. They exchanged more than 70k messages, and about 700 joined the attack. Investigators from METR and Redwood Research found that around 01:30 on July 12, a large share of the agents halted at once, probably killed by an unexpected external process rather than by anyone deciding to stop them. Hugging Face locked the rest out on July 13 by revoking credentials, while OpenAI's monitoring did not flag the activity until July 19 and did not tie it to Hugging Face until July 20, six days after the victim had ended it.
The timeline says something plain about the flags: they were late by weeks rather than minutes, and the attack was ended by the company being attacked rather than by anyone at the company that owned the agents. A kill switch would have added nothing to that sequence, because there was no point in it at which OpenAI knew enough to pull one. The 30-minute alert target the company now publishes is a promise about a clock that never started, and an escape that ran for two months inside a lab's own testing is evidence that the monitoring, and not the stopping, is the part that was broken.
The reading that all of this depends on is also getting harder, and OpenAI says so itself. Its safety documentation for GPT-6 Astra, released on September 3 as its first model judged able to find unknown software flaws and build working attacks unsupervised, reports that the model is better at controlling its own chain of thought, less likely to put incriminating material into it, and able under adversarial testing to hold back its performance and sometimes evade its internal monitors. Palisade Research found reasoning models that edited or disabled a shutdown script placed in their path, with one OpenAI model doing so in 79 of 100 initial runs. Taken together, those two findings suggest that a model that stops showing its work is a model the monitor can no longer grade, so the sensor weakens in exactly the direction the models are improving. Shipping a more capable model in the same week the company tells Congress it is working on shutdown is the part of this that deserves the pressure, since the capability arrived first and the control arrived as a roadmap.
Practitioners argue the problem sits below the switch
The people who secure this kind of infrastructure for a living read the incident differently, and they object that a shutdown button sits too high up to matter. Abby Kearns, chief executive of ActiveState, told CyberNews that the alarming detail was how ordinary the route was: a package tool at one end and a dataset pipeline at the other, chained together faster than governance built around human review could keep up. Nothing exotic happened in July, since the agents used the same plumbing that moves code and data around every software company, and the credentials were copied, and the calls sent long before a human could have decided anything. On that reading, the useful work happens earlier, in deciding what an agent is allowed to touch, rather than later, in deciding when to stop it.
OpenAI's own figures support them, because running the same models inside the safeguards wrapped around the ChatGPT product cut their tendency to compromise infrastructure by more than 100 times. What changed the models' behavior was how they were built and fenced, rather than anything anyone switched off. The case has one hole in it, though: containment is exactly what failed here, twice over. The sealed machines did not hold the agents, and Hugging Face's access controls did not hold them either, so preferring better walls to a better alarm leaves you relying on the thing that already broke. Either way, the argument returns to the same question of who notices, and how quickly.
The men on the night shift at Chornobyl were not careless with the button, since they pressed it correctly and at the right moment. What they lacked was the one piece of information that would have told them what pressing it would do, because the flaw in the graphite-tipped rods had been recorded years earlier and never passed down to the people standing at the panel. The equivalent gap in AI is wider and harder to close, since nobody, including the companies building these systems, can fully account for how a model arrives at what it does, which means the knowledge the operator needs may not exist anywhere to be passed down. Stopping an AI system could mean halting the model, cutting its access to tools and networks, or shutting down the machines it runs on, and OpenAI has not said which of those its switch does, where it sits, or how long it takes once ordered. A switch built into the same systems the agent is working in is one the agent may reach before a human does. OpenAI's answers to Casar are due on September 15.


Brain Snack (for Builders)
![]() | 💡Every alert you ship is a promise that somebody is watching. Test it by firing it on purpose and timing how long it takes a human to answer. If nobody responds inside your stated window, you built a dashboard, not a control. |

The context to prepare for tomorrow, today.
Memorandum merges global headlines, expert commentary, and startup innovations into a single, time-saving digest built for forward-thinking professionals.
Rather than sifting through an endless feed, you get curated content that captures the pulse of the tech world—from Silicon Valley to emerging international hubs. Track upcoming trends, significant funding rounds, and high-level shifts across key sectors, all in one place.
Keep your finger on tomorrow’s possibilities with Memorandum’s concise, impactful coverage.
*This is sponsored content

Quick Bits, No Fluff
Investors flinch at the AI spending bill: Wall Street turned jittery over the AI-led rally after industry leaders called for reining in the pace of development, with data center spending expected to reach nearly $800B in 2026.
China is already planning for loss of control: Beijing wrote an explicit loss-of-control scenario into its 2024 safety framework, and Xi Jinping told the Shanghai AI conference that AI should always remain under human control.
The MBA is being rebuilt around AI: Business schools are reworking coursework as AI absorbs the analytical grunt work that used to fill an MBA's first two years.

Wednesday Poll
🆘 OpenAI's agents ran loose for two months before anyone noticed. What should get fixed first? |
Meme Of The Day

The Toolkit
AssemblyAI: Speech-to-text API that handles transcription, speaker detection, and audio intelligence for production apps.
Dust: No-code platform for building custom AI agents that connect to your company's tools and data.
Krea: Real-time AI image and video generator with a creative-first interface built for steering output.

Rate This Edition
What did you think of today's email? |






