- Roko's Basilisk
- Posts
- Anthropic Won't Promise Anymore
Anthropic Won't Promise Anymore
Plus: AI's trillion-dollar break-even math, Instinct's $10B talks, and Washington's thinning options.
Here's what's on our plate today:
🔍 Anthropic's 140-page misuse report names culprits only Anthropic can actually stop.
💸 Hyperscalers near $1.1T in spend; Instinct eyes $10B; Washington runs low on levers.
🛠️ Weekend To-Do: test the fingerprints, map agent autonomy, read Google's version.
Let’s dive in. No floaties needed.

Goodies delivered straight into your inbox.
Get the chance to peek inside founders and leaders’ brains and see how they think about going from zero to 1 and beyond.
Join thousands of weekly readers at Google, OpenAI, Stripe, TikTok, Sequoia, and more.
Check all the tools and more here, and outperform the competition.
*This is sponsored content

Put your brand in front of 250,000 tech decision-makers.
Bay Area Times reaches more than 250,000 founders, operators, investors, and venture capitalists, 90% of them in the United States, with a 53% open rate.
Native placements sit inside the editorial flow rather than beside it, so your message gets read with the news instead of scrolled past. Slack, Attio, and Granola have run here.
Placements range from a single secondary slot to a full newsletter takeover.
*This is sponsored content

The Laboratory
TL;DR
Anthropic has published the most complete account yet of how AI is misused, and almost nobody outside the company can act on it.
The document: on September 10, 2026, Anthropic detailed operations it says it disrupted across seven areas of harm, from espionage to missile guidance software.
The finding: misusers now hand AI whole jobs, running familiar attacks faster and cheaper rather than inventing new ones.
The admission: the company can no longer promise its newest models are too weak to help a sophisticated user with dangerous biology.
The contrast: Google's threat reporting gives security teams something to defend, while Anthropic's cases happened inside Claude, where only Anthropic can change anything.
The stakes: the readers able to act are in Washington, and a company preparing a share offering has its own reasons to be heard there.
What Anthropic's new threat report says, and who can act on it
A security advisory is written with a simple purpose: to tell someone how to fix a problem. A software company finds a flaw in its product, explains how the flaw can be exploited, and ends by telling users to install an update. The company fixes its own product, and everyone else has something they can do to protect themselves.
AI misuse is different because the problem may not be a flaw in the product at all. The same model that helps a programmer write code can help someone break into a computer, and the model is not malfunctioning in either case. The problem is what the customer is asking it to do, which means there is no patch for an outside company to install and no simple instruction for a customer to follow. The company selling the model is left with the unusual position of being the one that sees the misuse, decides whether it is dangerous, and has the power to stop it.
Anthropic's newest threat report is a document of exactly that kind, and it is the fullest public account yet of how people are misusing AI. On September 10, 2026, the company published more than 140 pages on operations it says it disrupted between December 2025 and August 2026, across seven areas of harm from espionage to propaganda to weapons work. The cases are specific, and the culprits are named, but almost every response described is one only Anthropic could carry out. The company banned the accounts and tightened its safeguards itself. That makes the question of who the report was written for as interesting as what it found.
The central finding is that misusers now hand AI the whole job
The change running underneath most of the cases concerns how the work gets divided. In the hacking operations Anthropic describes, AI carried out the work while people selected the targets and reviewed the results. The tool that allows it is an 'agent', a model set up to run a multi-step job by itself while a person checks in occasionally. Anthropic says the approach has spread to every kind of actor it investigated, from suspected state-backed spies to two undergraduates in China's Hunan province, a range that shows how cheaply the method now spreads.
Delegation shows up away from hacking as well, since nine covert propaganda campaigns turned Claude into their writer, copy editor, and back office, several timed to elections. What links those campaigns to the intrusions is that neither required a new technique, and Anthropic says the familiar ones simply became cheaper to run. Cost carries more weight than it sounds, since skill no longer reveals who is behind an intrusion, and some breaches ran from first access to bulk theft in two to three hours. Anthropic then adds a caveat that runs counter to its own framing, since people directed every step of several of the most damaging break-ins.
The gravest cases are the ones the safeguards only partly caught
That caveat matters most where the safeguards were working, and something got through anyway. A group in northern Yemen used Claude in place of programmers on guidance software for rockets and missiles, Al Jazeera reports, and many of its requests were refused. Several succeeded because the operators hid their goal and split the work across sessions, so no single exchange looked like weapons work. Anthropic found no evidence of a working weapon, since a test launch appears to have failed.
Biology followed the same shape and produced the report's most consequential admission. Five cases involved researchers getting around regional bans or disguising their purpose, CNN reports, which prompted the company to comment on its own models. It says its 2025 models were too weak to meaningfully help a sophisticated user with dangerous biology, while for current ones "the evidence is no longer certain, and we cannot make that same assurance." Newer releases, such as Fable 5, therefore carry tighter limits on 'dual-use' questions, those whose answers can protect or harm depending on who is asking.
The largest numbers in the report concern 'distillation', a rival training its own model on another company's answers. TechCrunch counts nearly 200M such exchanges with Claude, most attributed to China's Alibaba, the only harm here counted in volume. It also sits inside a wider dispute, since three U.S. security agencies accused six Chinese AI companies of the same practice two days earlier. China's commerce ministry has replied that the practice is common across the industry, including among U.S. firms.
Google's version of the report hands defenders something to use
Anthropic is not the only company publishing work like this, and the comparison shows the report's limits. Google's threat intelligence group released its own findings this month, describing the same move from prompting to agent-driven automation. A security team reading it has something to do, since the attacks Google describes play out on the corporate networks those teams defend. Anthropic's cases mostly happened inside Claude, where the defenses that count are the ones it writes into its own product.
Some of it does travel, since the report publishes technical fingerprints that outside defenders can test, meaning file signatures and web addresses that identify the same attacker elsewhere. What it cannot offer is a repair anyone else can make, which is the difference between a security firm's account and a developer's account of its own product. The gap widens given that the six labs in the U.S. advisory drew on four American developers, so one company's records cannot cover the whole problem.
The readiest audience for the report sits in Washington
In government, a document like this can have consequences beyond the company that produced it, which explains much of its timing. Anthropic argued in February that distillation campaigns strengthen the case for chip export controls, and this report landed two days after a federal advisory making a similar charge. Its other interests sit close to the surface, since the company is preparing a public share offering and fighting a Pentagon blacklisting that a judge ruled unlawful in August. Al Jazeera reads the document as an effort to restore its standing with the U.S. defense establishment.
None of that makes the cases untrue, and Google's account of the same trend is a reason to take them seriously. The interests instead lie in what the report leaves out, since Anthropic picked the cases, made the attributions, and holds the data. It publishes no count of the accounts it banned or the activity it missed, and calls these its most notable finds, unrepresentative of typical misuse. Read on those terms, it shows where misuse is heading without measuring how much of it there is. John Thickstun, a computer scientist at Cornell, told the AP that companies are left making society-wide calls without democratic oversight.
A security advisory ends by telling the reader what to install, and the reader can then check whether the fix worked. This report ends with accounts Anthropic has shut and limits it has tightened, both inside a product nobody outside the company can inspect. What is left for everyone else is the slower test of whether these patterns surface in other companies' reporting. Until they do, the fullest account of a year of AI misuse is the one written by the company describing itself.


Weekend To-Do
Run Anthropic's indicators through your own logs: The 140-page report publishes file signatures and web addresses from the intrusions, the one part an outside defender can actually test.
Read Google's account of the same shift: Google's threat group describes the same move from prompting to agent-driven attacks, but on the corporate networks your team defends.
Check the federal advisory naming six Chinese labs: Three US security agencies accused them of distilling four American developers' models, two days before Anthropic's report landed.

The context to prepare for tomorrow, today.
Memorandum merges global headlines, expert commentary, and startup innovations into a single, time-saving digest built for forward-thinking professionals.
Rather than sifting through an endless feed, you get curated content that captures the pulse of the tech world—from Silicon Valley to emerging international hubs. Track upcoming trends, significant funding rounds, and high-level shifts across key sectors, all in one place.
Keep your finger on tomorrow’s possibilities with Memorandum’s concise, impactful coverage.
*This is sponsored content

Friday Poll
🤖 Anthropic found the misuse, judged the danger, and fixed it alone. Who should be checking that work? |
|

Headlines You Actually Need
AI's trillion-dollar break-even math: Hyperscalers are on track to spend nearly $1.1T on data centers through 2027, and would need to nearly triple their own productivity to break even by 2030.
Instinct chases a $10B mark: The invite-only personal AI assistant is in talks to raise up to $1B at roughly a $10B valuation, weeks after closing a round at $2.5B.
Washington runs low on levers: Bloomberg reports the administration has few good options left to slow China's rise as an AI superpower, even as Trump insists the US is comfortably ahead.
Meme Of The Day

Rate This Edition
What did you think of today's email? |






