The Model Went On Sale

Plus: Apple's own China model, Flock's opaque abuse detector, Z.ai's coding push.

Here's what's on our plate today:

  • 🧪 The 80% price cut: what it reveals about where AI's toll actually sits.

  • 🍪 Apple trains its own China model, Flock won't explain its abuse tool, Z.ai chases the coding crown.

  • ⭐️ Roko's Pro Tip: build for swappable models, spend where nothing can substitute.

  • 📊 Poll: which chokepoint actually collects from here?

Let’s dive in. No floaties needed.

AI is reshaping fraud. It’s time to rethink trust.

AI is changing how fraud shows up and how quickly it can scale. Signals that once worked are easier to spoof, forcing teams to choose between higher losses or more friction.

This white paper explores why traditional identity approaches are breaking down in the AI era and how financial behavior and broader context help teams catch more fraud earlier while maintaining seamless user experiences.

*This is sponsored content

Build and design your website on Framer - Now with Agents

Framer is a pro website builder trusted by companies like Miro and Perplexity that helps creators, teams and businesses ship production-ready sites faster than ever. With AI agents built directly into the canvas, teams can design pages, manage CMS content, write copy, add SEO, and audit for issues — all without leaving the tool where the real site lives. Agents bring speed and scale; you bring taste, judgment, and control.

*This is sponsored content

The Laboratory

TL;DR

Everyone bet the toll booth would sit on the model. The model just went on sale.

  • The clue: OpenAI cut the price of its cheapest model by 80% three weeks after launch, while leaving its flagship alone. Cheap models can be swapped out, so the price collapsed where switching is easiest.

  • The precedent: In wireless, operators spent $1.5T building networks and captured almost nothing. Apple and Qualcomm collected instead, holding the customer and the license.

  • The tolls now: NVIDIA hasn't cut prices while its customers have. Memory, a commodity three years ago, is sold out through 2027. Power is next.

  • The catch: No toll lasts. Even Qualcomm's licensing revenue is slipping now that wireless is easy to copy.

  • The stakes: Hundreds of billions have been committed to a bottleneck that keeps relocating faster than anyone can plan around it.

Decoding where AI's chokepoints actually lie

A chokepoint is a narrow passage where all traffic must pass, and whoever controls it collects a toll from everyone else. Canals, mountain passes, and payment networks have all worked this way, and the pattern has never rewarded whoever spent the most money building the route. It rewards the one part of the chain that nobody else can circumvent.

In technology, it is worth watching the gap between who builds a technology and who eventually controls its chokepoints. The companies that spend the most money building something new are not necessarily the ones that capture the most value from it. They are betting that their investment will give them a lasting position, but they only find out whether that bet paid off once the technology matures and its real bottlenecks become clear.

A token is the unit these systems bill by, and one token roughly corresponds to the length of a short word. That makes the price of a million tokens the closest thing the AI industry has to a wholesale price for intelligence. On July 30, 2026, OpenAI cut the price of Luna, the cheapest model in its GPT-5.6 range, by 80% to 20 cents per million input tokens, just three weeks after putting it on sale. A price cut that large, that quickly, suggests customers had already found alternatives.

Where that work went is visible on OpenRouter, a service that lets companies send the same job to whichever model they choose. Since February, Chinese-built models have handled at least 30% of the work that American companies send through the platform each week, peaking at 46%, compared with 4.5% in the first half of 2025. Those models cost 60% to 90% less than the leading American systems. When work can move between providers that easily, the model itself is not much of a chokepoint.

So where does the chokepoint actually lie? The answer will not come from any one company or product, but from the clues left behind as the economics of AI change. The first was the size and speed of OpenAI's price cuts. The second is where those cuts stopped: OpenAI lowered prices on Luna and its mid-range model while leaving its most capable system untouched. That tells us where competition is strongest, and where buyers have the easiest alternatives. The task now is to follow those clues up the chain and see where the alternatives run out.

What the last build-out charged for

An industry has run this experiment to completion once before at a comparable scale, and the wireless build-out shows which positions ultimately took a toll. Mobile operators built the networks and paid to build them, spending $1.5T between 2023 and 2030, with the bulk of that going to 5G, while their revenue barely moved over the same period. Ordinary customers proved unwilling to pay much more for faster data, leaving the companies that built the road with little to capture from the traffic it carried.

Two other parties collected instead, and neither of them built any of the network. Apple took 49% of all global smartphone revenue in the second quarter of 2026 while shipping roughly one handset in five, because a rival could not match the device or the relationship it already sat within. Qualcomm keeps roughly two-thirds of every licensing dollar as profit, charging for patents that any phone running 3G, 4G, or 5G must license, which means its toll is built into every handset regardless of who assembled it.

Even that second toll is thinning now, with Qualcomm's licensing revenue slipping rather than growing in its most recent quarter. A narrow passage stays narrow only while the technology passing through it is still hard to replicate; once wireless stops being hard, it no longer is.

The tolls being collected now

Applying that test to AI, the clearest toll today sits with the company selling the hardware, and it is the one link in the chain that has not cut its prices. NVIDIA earned $215.9B in its last full year at gross margins near 75%, selling to the same laboratories that are now reducing what they charge their own customers. Pressure has traveled through this chain as far as the model builders and stopped there.

A second toll appeared where almost nobody was looking, because memory chips were a commodity business three years ago. Three manufacturers account for around 90% of the world's DRAM, and by late 2025, SK Hynix had already sold its entire 2026 production capacity, with contract prices rising by half in a single quarter. Samsung and SK Hynix have since warned that shortages will run through at least 2027, which is what a chokepoint sounds like from the inside.

A third is forming around electricity, which used to be a line on an operating budget rather than a limit on what could be built. Gartner expects 40% of existing AI data centers to be operationally constrained by power availability by 2027, and expects the cost of securing that power to be passed along to the companies selling AI products. Whoever can guarantee megawatts is close to holding a narrow passage.

The toll keeps moving

The companies that control where this intelligence reaches people are hedging against it all, and their behavior is the clearest read on what they think is scarce. Microsoft has begun replacing OpenAI and Anthropic models with systems it built itself in Excel and Outlook, while Apple agreed that its next generation of foundation models would be built on Google's Gemini models and cloud technology, for a sum reported at about $1B per year. One is building a substitute and the other is renting one, and both keep the surface whichever model ends up answering.

Neither has stopped paying entirely, because the model's toll has narrowed rather than disappeared. Menlo Ventures, a venture firm that surveys large companies about their AI budgets, found open-source usage falling from 19% to 11% in a single year, even as developers routed more work to cheap models. Large buyers are purchasing support, indemnity, and somebody to call to switch providers to save on tokens, which is the one segment where the laboratories still name their own price. They lost the ordinary high-volume work, and that volume is what the hundreds of billions were raised against.

What makes AI harder to read than wireless is not that buyers are behaving unusually; they are doing the ordinary thing and taking the cheapest route that works. The technology underneath them keeps rearranging what is scarce. Memory was a commodity, electricity was an operating cost, and the model was the entire investment thesis, and all three changed position inside three years.

Whoever ends up collecting on this build-out will hold something difficult to see from here, which leaves the decision with the companies signing multi-year contracts, the investors pricing two of the largest share sales ever attempted, and the engineers choosing this week where to send their work. Each is committing money against a guess about what will still be scarce in five years. If the technology keeps moving its own bottlenecks at this speed, the winner of this phase will not be whoever spent the most building the road, but whoever happens to be standing in the narrow passage when the movement stops.

Roko’s Pro Tip

💡 

If your margin depends on today's token price, you are building on the one layer of the stack that is actively collapsing. Route every call through something you can swap out in an afternoon, and keep one benchmark that tells you when the cheap model is already good enough. Then spend your real money on the part nobody else can buy at the same price you did.

The context to prepare for tomorrow, today.

Memorandum merges global headlines, expert commentary, and startup innovations into a single, time-saving digest built for forward-thinking professionals.

Rather than sifting through an endless feed, you get curated content that captures the pulse of the tech world—from Silicon Valley to emerging international hubs. Track upcoming trends, significant funding rounds, and high-level shifts across key sectors, all in one place.

Keep your finger on tomorrow’s possibilities with Memorandum’s concise, impactful coverage.

*This is sponsored content

Monday Poll

📊 Models got cheap fast, and hundreds of billions are already committed. Which chokepoint actually collects the toll from here?

Login or Subscribe to participate in polls.

Bite-sized Brains

Meme Of The Day

The Toolkit

  • Continue: Open-source AI code assistant that plugs into VS Code and JetBrains with full control over models and context.

  • Chroma: Open-source vector database for AI apps, fast to set up and easy to scale for RAG.

  • Krea: Real-time AI image and video generator with a creative-first interface for designers who want to steer the output.

Rate This Edition

What did you think of today's email?

Login or Subscribe to participate in polls.