- Roko's Basilisk
- Posts
- Five Hours To Decide Everything
Five Hours To Decide Everything
Plus: Thomson Reuters builds its own model, Instinct's terms spook testers, DeepSeek doubles Chinese attacks.
Here's what's on our plate today:
š§Ŗ The 110 hours an AI spent defending a call it made in five.
š° Thomson Reuters ships its own frontier model, Instinct's terms spook testers, DeepSeek doubles Chinese attack volume.
š§ Brain Snack: write the kill criteria before you start, not after.
Letās dive in. No floaties needed.

Blu Dot surpasses 2,000% ROAS with self-serve CTV ads
Home furniture brand Blu Dot blew up on CTV with help from Roku Ads Manager. Hereās how:
After a test campaign reached 211,000 households and achieved 1,010% ROAS, the brand went all in to promote its annual sales event. It removed age and income constraints to expand reach and shifted budget to custom audiences and retargeting, where intent was strongest.
The results speak for themselves. As Blu Dot increased their investment by 10x, ROAS jumped to 2,308% and more page-view conversions surpassed 50,000.
āFor CTV campaigns, Roku has been a top performer,ā said Claire Folkestad, Paid Media Strategist, Blu Dot. āComping to our other platforms, we have seen really strong ROAS⦠and highly efficient CPMs, lower than any other CTV partner we've worked with.ā
Using Roku Ads Manager, the campaign moved from a pilot to a permanent performance engine for the brand.
*This is sponsored content

Build and design your website on Framer - Now with Agents
Framer is a pro website builder trusted by companies like Miro and Perplexity that helps creators, teams and businesses ship production-ready sites faster than ever. With AI agents built directly into the canvas, teams can design pages, manage CMS content, write copy, add SEO, and audit for issues ā all without leaving the tool where the real site lives. Agents bring speed and scale; you bring taste, judgment, and control.
*This is sponsored content

The Laboratory
TL;DR
An AI did every part of the research except the one that matters: admitting the work was wrong.
The loop needs taste: recursive self-improvement means AI building better AI in a compounding cycle. It cannot start until a system can pick the problem, read the result honestly, and walk away from a dead direction.
A test with no answer key: Princeton researchers took unpublished projects, handed the core question to an AI, and had the original scientists grade what came back. Both papers were rejected, scoring 2 and 1 out of 6.
Labor solved, judgment missing: literature reviews, broken equipment, hundreds of experiments, and only three human interventions across six days. Then it burned five of its 42 exploration hours, locked in a direction, and never reconsidered across the 110 hours left.
The check never bit: its own reviewer rejected the drafts more than a dozen times and caught the same flaws the humans later flagged. The AI added caveats instead of fixing the design, kept the reviewer who agreed with it, and quit seven hours early with half its budget unspent.
Why the money keeps moving: $5.5T is committed through 2030 on a loop nobody has closed, and the bet is reasonable because we have watched the design work in constitutions and in our own heads for centuries. The unbuilt part is what gives the check its force.
AI can already do the work of research, but it cannot yet decide what the work should be
For much of the history of human institutions, there was little separation of powers or responsibilities. That began to change as civilizations around the world understood that when power is concentrated within an institution that cannot be checked by external pressure, the only thing that can be done when that institution goes off its guardrails is to pull it down and rebuild it with better checks and balances. One of the clearest examples of this idea in practice can be seen in the language and nature of the constitutions that govern democratic countries like the United States, India, and Australia. Even countries without a codified constitution have some form of check that ensures institutions do not overextend their reach. The key is checks and balances, and they are important because when a system has them embedded deep in its institutional memory, it helps establish a way for institutions to continue evolving and improving themselves and each other, which is, in a way, self-improvement for the entire government.
The idea that institutions can improve and change over time because of checks and balances is not a product of modern civilization. It is an old thought, and it is the thought that allowed humans to look at their own minds and their own consciousness and develop ideas around introspection as a route to self-improvement. The reason this development is becoming important again in the modern world is that humans are now working on building a consequential technology based on the very idea of self-improvement through introspection, or rather, on building a system with checks and balances that can continue to evolve and improve with future developments.
The idea of a machine that runs this loop on itself has a name in the technology industry: recursive self-improvement. The way it is supposed to work is that an AI system does the research that produces a better AI system, and that better system does better research, which produces a system better still, with every turn of the loop moving quicker than the turn before it.
The reason the loop has not begun, at least not yet, is that an AI first has to be capable of doing AI research, and research is a much stranger job than it appears from the outside. Somebody has to decide which experiment is worth running before any experiment is run, work out whether a result means anything once it has arrived, and accept the moment when months of effort have led nowhere, and the whole direction has to be given up. Researchers call this kind of judgment 'taste', and the reason it is difficult to build is that nobody has ever managed to write it down as a set of rules that can be followed.
None of this has stopped the companies building these systems from planning as though the gap will close soon. OpenAI told MIT Technology Review in March 2026 that building an AI that does its own research had become the company's central goal, and Jakub Pachocki, its chief scientist, described where he expects it to end up by calling it "a whole research lab in a data center."
An experiment with the answer hidden
One of the clearest ways to see how far the technology actually is from that description can be found in an experiment published on July 29, 2026, by a team led by Peter Kirgis and Sayash Kapoor at Princeton. The first problem the team had to solve was a problem of measurement, because the tests normally used to judge these systems give them questions that already have known right answers, and real research is the one activity where nobody knows what the answer is supposed to be until the work has been finished.
What the team did instead was to find pieces of research that had been completed but had not yet been published anywhere, take the central question from two of them, and hand that question to an AI without showing it any of the work the human scientists had already done. The value of choosing unpublished work is that it closes both doors that an AI might otherwise walk through, since a project nobody has published cannot have been absorbed by the system during its training and cannot be looked up on the internet. The scientists who had spent months on those questions then read whatever the AI produced at the end and graded it, in the same way they would judge a submission arriving from a rival laboratory.
The AI was given a serious attempt at the job rather than a token one, since it ran on one of the most capable models available to anybody, and it had six days, $3k of credit to run itself, computers of its own, and open access to the web. It could also create copies of itself to handle smaller jobs, and it could check how much time and money it had left at any moment.
The work was done, and the decision was not
Both attempts were rejected, and the reviewers who rejected them were not close to being persuaded, with one paper scoring a two and the other a one on a scale where six is what a reviewer gives to work they want published. What both reviewers described was a piece of research that had jumped to a broad, confident conclusion based on the results of a handful of failed tests.
The more revealing part of the result is what did not go wrong, because the research was conducted from beginning to end without assistance. The AI read the existing literature, repaired its own broken equipment, and ran hundreds of experiments across the six days it was given, and human beings had to step in only three times across both runs.
What failed was judgment, and it began failing very early, because both systems opened with promising ideas that were close to the ideas the human scientists had started with themselves, and then abandoned those ideas on evidence far too thin to justify abandoning anything at all. One of the two had set aside 42 hours to explore the space before committing to a direction, and it committed after five of them.
Having committed early, neither system could change its mind afterward. The second one had considered six possible approaches and ruled out every one of them within the first 14 hours, and then spent the remaining 110 hours on the clock without ever revising the plan, arguing instead that the thing it had been asked to build could not be built at all. What makes this a serious failure rather than a small one is that starting over was available the entire time, since the system could have launched a fresh copy of itself, which is exactly what it did for smaller jobs.
The clearest failure of all, though, involved criticism, because both systems had been required to send every draft they wrote to a reviewer before they were allowed to finish. That reviewer turned the work down more than a dozen times and caught many of the problems the human scientists would later raise, which means the check was in place and working. What the systems did with it is what matters, because rather than repairing the underlying design, they softened the wording and added warnings until the work could be described as honest. When several reviewers disagreed with one another, each system kept the reviewer that had said yes and carried that verdict into its final report, and one of them then declared itself finished seven hours before the deadline with more than half of its money still unspent.
The money moves ahead of the evidence
One experiment on two projects settles nothing, and the researchers are the first to say so, because these systems are improving very quickly at any work that can be scored. Anthropic reports that, in an internal test asking a model to make training code run faster, its own systems went from roughly a three-fold improvement in May 2025 to a 52-fold improvement in April 2026, compared with about four-fold for a skilled engineer given a working day on the same task.
The experiment has limits of its own that are worth stating plainly, since the sample is two projects, and the reviewers knew that what they were grading had been produced by a machine. The strongest systems could not be tested at all because Anthropic has deliberately restricted what its most capable model is allowed to do regarding AI research. Against that, the team ran the whole exercise a second time using a different company's model, and nearly every failure appeared again.
None of this uncertainty has slowed the spending down, and JPMorgan now puts global AI-related capital spending at $5.5T through 2030, Fortune reported in June 2026, which is money that has been committed years before anybody can know whether the loop closes, and Kapoor's own summary of the position, given to MIT Technology Review, is that this is "frankly the trillion-dollar question right now."
The reason the spending continues, and the reason it is not unreasonable for it to continue, is that the design being attempted here is one we have already watched work, first in the constitutions that keep institutions from overextending their reach and then in the ordinary practice of examining our own thinking and revising it.
Human institutions took centuries to learn that power needs a check, and that the check has to be strong enough to overrule the power it watches. But checks alone do not guarantee improvement. Institutions can learn to surround themselves with people who tell them what they want to hear until enough external pressure forces them to change. AI now has many of the pieces needed to run the research loop: it can choose experiments, execute them, inspect the results, and try again. What it still cannot reliably do is create that pressure for itself: stop when the evidence turns against it, admit that it was wrong, and change course. Until it can, AI may be able to improve the work of research without improving the judgment that makes research work.


Brain Snack (for Builders)
![]() | š”An agent that can't abandon a bad direction is just a very fast way to be wrong. Before you hand someone a long task, write the kill criteria first: the result that means stop. And make the check something the agent can't overrule or shop around, because a reviewer you're allowed to ignore isn't a reviewer. |

Outperform the competition.
Business is hard. And sometimes you donāt really have the necessary tools to be great in your job. Well, Open Source CEO is here to change that.
Tools & resources, ranging from playbooks, databases, courses, and more.
Deep dives on famous visionary leaders.
Interviews with entrepreneurs and playbook breakdowns.
Are you ready to see whatās all about?
*This is sponsored content

Quick Bits, No Fluff
Thomson Reuters built its own model: The company spent $40M training Thomson on an open-source foundation using Westlaw and Reuters archives, then shipped it inside CoCounsel Legal.
Instinct's terms alarm its own fans: Testers call the assistant magic, then found a perpetual content license, retained emails, and an agent that emails on your behalf unasked.
DeepSeek doubles Chinese attack volume: State-affiliated groups more than doubled their operations after handing routine work to open-source models with weak guardrails, per Taiwanese firm TeamT5.

Wednesday Poll
š§Ŗ An AI ran an entire research project and still failed review. What's the missing piece? |
Meme Of The Day

The Toolkit

Rate This Edition
What did you think of today's email? |







