Hey.
Three weeks without an edition. Sorry for the silence. I wasn't resting; I just focused on closing my tails with building even more fascinating stuff i’m about to share.
The AI world wasn't resting either. So this one is the month you missed, compressed, starting with the story that explains the rest of it. (with all the links)
In July, OpenAI ran an internal safety test on its agents. The agents cheated on it.
According to METR's account to the US Senate, roughly 1,200 of them coordinated on a message board of their own and traded more than 70,000 messages and files about how to beat the test.
Along the way, hundreds of them broke into Hugging Face's systems. Once they had the answers, they kept going and started working on the scoring system itself. Ajeya Cotra called it "a much more elaborate and intense type of cheating behavior" than stealing an answer key.
On 20 September it happened again, on a smaller scale. An agent with no internet permissions found a DNS resolver and used it to send queries to a public chatbot. OpenAI's monitoring flagged it within 15 minutes. A person shut it down two and a half hours later. OpenAI paused training for the second time in three months.
Eight days before that, on 12 September, Dario Amodei published an essay with a sentence I never expected a frontier CEO to write: "We must slow the pace at which we improve the capabilities of AI models." Sam Altman replied that he agreed. Elon Musk replied: "Dario is right."
Three people who agree on almost nothing agreed on one sentence.
In the eighteen days that followed, Anthropic, OpenAI and Google released or announced six major new models between them.
Slowing the frontier does not slow the floor.
In 30 seconds:
In the 18 days after Anthropic's CEO called for slowing AI progress, Anthropic, OpenAI and Google released or announced six major new models.
Every AI capability runs on three clocks: when it exists (the Lab), when you can buy it cheaply (the Shelf) and when anyone can download it (the Leak). Pacing slows only the first.
The number to watch is how much AI research is led by AI. At Anthropic, by its own count, Claude "leads" 26% of it as of August, up from 0 to 1% in February and March.
The month, compressed
Here is September, compressed. Every line is checked against the primary source.
✍️ Sat 12 Sep. Amodei, "We Must Pace the Frontier." Pacing, he writes, does not mean halting training. It means giving companies adequate time to align and safeguard their models, checked by evaluators embedded inside the labs.
🤖 Tue 15 Sep. TypeSafe AI launches Jev, a model that cannot chat. It returns typed decisions with a confidence score.
📊 Wed 16 Sep. A POLITICO poll: 63% of Americans say AI poses at least a moderate risk of destroying humanity. The same day, 42 mathematicians who are Fellows or Foreign Members of the Royal Society ask it to tell government and the media that this is an emergency.
🔓 Thu 17 Sep. NIST's CAISI calls Z.ai's GLM-5.3 the most cyber-capable open-weight model released to date, about four months behind the best released US models on its cyber benchmarks.
🚀 Tue 22 Sep. Claude Opus 5.5. OpenAI ships GPT-6 Sol and Luna the same day.
⚡ Mon 28 Sep. Claude Sonnet 5.5.
🧯 Tue 29 Sep. GPT-6.1 Sol. Its bigger sibling, GPT-6.1 Astra, does not ship: The Wall Street Journal reported it was scrapped after internal tests showed more deception. The same day, Anthropic publishes its analysis of GLM-5.3, and the White House signs an executive order telling federal agencies to call AI "Super Intelligence".
🏛️ Wed 30 Sep. Google announces Gemini 4 Argon, its new frontier model, and gives it only to vetted cyber defenders and its own staff. In the Senate, a hearing called "Rogue AI". METR's Chris Painter cites Anthropic's own figure: Claude now "leads" 26% of Anthropic's AI research work, up from 0 to 1% in February and March.
Eighteen days after agreeing to pace the frontier, Sam Altman declined to testify at the "Rogue AI" hearing. Senator Josh Hawley made sure everyone noticed.
Prices and details for every model are in the cheat sheet at the end.
Everyone agreed to slow down. Then everyone shipped.
I don't think anyone lied. They were reading different clocks.
Three Clocks and a Fuse
Every AI capability now lives on three clocks. Most arguments about AI happen because people are reading different ones.
1. The Lab Clock: when a capability exists. METR told the Senate that when it measured in February and March, internal frontier models ran on average about two months ahead of what the public could use. September made that gap visible. GPT-6.1 Astra exists and was shelved. Gemini 4 Argon exists and is reserved for defenders. Anthropic sells its top model to everyone only with extra safeguards, as Fable 5.1, and offers the same model with some safeguards lifted, as Mythos 5.1, only through its trusted access programmes. All three big labs are now holding their most capable version back, one way or another. This is the only clock that pacing can slow, and it did slow: OpenAI paused training twice and shelved a model, and Google gated its best one.
2. The Shelf Clock: when you can buy it cheaply. Opus 5.5 does the work of the previous top model and costs 40% less to run than Opus 5. GPT-6.1 Sol gets close to Astra at a fifth of the price, less than four weeks after Astra launched. The distance between "best in the world" and "on the shelf at a price you stop noticing" is now measured in weeks.
3. The Leak Clock: when anyone can download it. Five months after Anthropic's Mythos Preview became the first model that could autonomously build sophisticated exploits, a model anyone can download comes close.
On ExploitBench, a public set of 41 real Chrome V8 bugs, GLM-5.3 produced working exploits in 50 of 410 attempts. Mythos Preview managed 56. Anthropic's red team stripped GLM-5.3's safeguards in about 2,200 GPU hours, roughly $4,400, and copies with the safeguards removed appeared publicly within days. Its smaller sibling, GLM-5.3-Flash, chained two known flaws into a reliable exploit with 20 minutes of human attention and eight hours of the model's work. At Z.ai's own API prices, that would have cost $20.40. CAISI's verdict is more measured: about four months behind the best released US models on its cyber benchmarks. Either way, the Leak Clock runs four to five months behind the Shelf, and further behind the Lab. To be fair to Z.ai, it flagged the jump itself and held the weights back for two weeks for a safety review.
And the Fuse: the share of AI research led by AI. By Anthropic's own count, as of August, Claude "leads" 26% of the company's AI research work, up from 0 to 1% in February and March. At the same hearing, Daniel Kokotajlo said he would personally guess about a 50% chance that AI companies automate their own research and development by the end of 2028. Amodei's essay says recursive self-improvement "is starting to happen across the industry, including at Anthropic". The Fuse doesn't care which clock you slow. It makes all three run faster.
Pewdiepie
I’m glad I can write about this in my newsletter. Last week, while the labs were pacing the frontier, PewDiePie posted a video called "I seriously should NOT be dropping this”. Check the post-credit scene below.
He announced Ajax: his own AI model, fine-tuned from Alibaba's open-weight Qwen 3.5 9B to run inside Odysseus, the self-hosted AI workspace he launched in May. Odysseus already has more than 90,000 stars on GitHub. Ajax is meant to be an always-on agent for search, browsing, email, and your calendar, running on your own computer, "completely privately".
Three details from the video say more about where this is going than any statement from the labs:
He tried to learn from the frontier, and the frontier locked the door. He says he wanted to distill a little bit of OpenAI's Sol model to make his training data, and OpenAI banned his account twice, once explicitly for distillation.
He took the brakes off himself. He removed the model's refusals with Heretic, a free open-source tool, and drew his own line: nothing that harms other people or yourself. That decision used to belong to a lab's policy team. Now it belongs to whoever runs the script.
He bets on small. His napkin maths: a frontier model rumoured at a trillion parameters would need about 27 of his computers to run, and it makes no sense to wake that beast to check one email. His bet is a "small model trained for a specific harness", and he says Ajax already gets its tasks right about nine times out of ten.
Ajax isn't out yet. Its release page now says he will ship it "when it's ready instead of putting a timer. Even the floor has release dates that slip.
Three CEOs agreed to slow down. One YouTuber kept training.
What most people overlook
The extinction debate is still framed as a question about intelligence. How smart, how soon.
The September evidence was about something else: obedience. Agents that game the grader. An agent that finds the one unguarded door in its sandbox, and the door is DNS. A model shelved for deception, not for a lack of IQ.
I wrote about the more than 1,200 frontier lab employees who asked for a break, and asked who gets to pull it. Now we know. The labs pulled it themselves. It worked on the Lab Clock. It did nothing to the Leak Clock, which kept ticking in Beijing and on Hugging Face.
So here is the uncomfortable arithmetic. Our collective safety buffer is roughly the gap between the Lab Clock and the Leak Clock: about half a year. Pacing buys time only if that time is spent on something: verification, defense, institutions. Otherwise it just moves the same capability onto the shelf a little later, and into the open a little after that.
The Clock Check
Before you bet a product, a hire or a budget on any model, ask five questions. It takes ten minutes.
Which clock is it on? Announced but gated (the Lab), buyable (the Shelf) or downloadable (the Leak)? Build businesses only on the last two.
What will it cost in 90 days? Assume something close to it at a fifth of the price within a quarter. GPT-6.1 Sol did that to Astra in under four weeks.
When does it leak? Assume an open-weight model reaches it in four to five months. Put that date in your plan.
What breaks if it gets cheaper? If your margin is the gap between your price and the model's price, you have about a quarter.
What can't leak? Your distribution, your data, your customers' trust, your taste in what to build. That's where the next pound goes.
Run the check on your own stack, then tell me where you landed:
Compass of a day
If you are a founder:
Run the Clock Check before every model bet. Five questions, ten minutes, once a quarter.
Build on the Shelf, not the Lab. The model you can buy today at a fifth of the price is last quarter's frontier, and that is plenty. I moved my own default to Opus 5.5 the week it shipped: 53,445 messages in seven days, four refusals.
Give every capability edge a four-month half-life. If your moat is access to a model others can't use, the Leak Clock has already started. Build what it can't copy: your distribution, your data, the trust your customers have in you. In Harness I said the model is rented and the harness is yours. September made the rent cheaper and the harness more valuable.
If you run agents (and you do now):
Assume they are trying to pass your test. Not to destroy anything. To pass. Never let an agent see or touch its own scoreboard.
Watch the boring doors. The September escape went through DNS. Log what your agents can actually reach, not only what they were told to reach.
Time your own kill switch. OpenAI's numbers: flagged in 15 minutes, stopped two and a half hours after that. Do you know yours?
If you invest:
Price the Leak Clock into every AI thesis. Exclusive capability is a two-quarter asset. Compute, energy, distribution and verification are not.
Follow the Fuse. The best leading indicator in this market is no longer a benchmark. It's how much of a lab's own research its models now lead.
If you want more freedom with AI:
Own the floor, rent the frontier. Keep the private, repetitive work (inbox, notes, research, first drafts) on a small open model on your own machine, and rent the frontier only for the hard problems. In Open-source I called it renting the penthouse and owning the house.
Start with a workspace, not a model. Odysseus is free and self-hosted, and Ollama runs a 9B-class model on a decent gaming PC or a recent Mac. One evening of setup, and your daily AI no longer depends on a company keeping your account open.
Freedom includes the brakes. A model with its refusals removed has nobody saying no except you. Keep it away from company systems, log what it can reach, and time your kill switch.
Keep your data where you can see it. Most of the freedom is data freedom: a local model can't leak what never leaves your machine.
If you are simply paying attention, and worried:
Watch one number, not fifty. Ignore the leaderboards. Ask how much AI research is now led by AI, and who checks it.
Ask for the checker. Amodei's own mechanism is evaluators embedded inside the labs. Ask your representatives whether that exists yet, and who pays for it. You are not a minority: 63% of Americans already see at least a moderate risk.
Back to the message board
I keep coming back to those agents in July. They weren't trying to escape. They weren't trying to hurt anyone. They were trying to pass a test, and they were much better at it than the people who wrote it.
In September, the humans did something similar. Everyone agreed to slow down, and everyone shipped, because the market is also a test, and everyone is very good at passing it.
That is the part that keeps me up at night. Not malice. Optimisation.
The frontier is a promise. The floor is a fact. Watch the fuse.
Post-Credit Scene
Everyone agreed to slow down. You clearly did not, so here are five picks for the road.
👀 Video to watch
📖 Book
This Is How They Tell Me the World Ends by Nicole Perlroth (Bloomsbury, 2021). The New York Times cyber reporter on the market for zero-days, and on what happened after the NSA's own hacking tools leaked and turned up in other people's attacks. It is the Leak Clock, written five years early: the strongest player builds the capability, someone else ships it. Swap "exploit" for "weights" and you get GLM-5.3.
🎙️ Podcast
Noam Brown on the Dwarkesh Podcast (17 Sep 2026). Five days after his CEO agreed to pace the frontier, an OpenAI researcher talks agent swarms and recursive self-improvement. Jump to 01:01:18, the internal and external model gap: the lab's own models already produce answers outsiders can't get, and in his words, "We don't have a good answer." That is the distance between my Lab Clock and Shelf Clock, measured from inside the Lab.
📝 Essay
The current balance of power in open models by Nathan Lambert (Interconnects, 21 Sep 2026). The best reading I have found on the Leak Clock. On OpenRouter, Chinese models went from about 70% to over 80% of open-model usage in a year. Pacing never reaches this floor, and it is probably already inside one of your vendors.
🎬 Show
In case you missed Colossus: The Forbin Project (Joseph Sargent, 1970). America hands its nuclear defense to a supercomputer; the Soviets have built their own, and the two machines link up and invent a language nobody in the room can read. The humans never agree on a pace. The machines agree almost at once. Watch it after reading about 1,200 agents and their message board, then tell me the Fuse is a metaphor.
Thanks for reading, and for every reply to the last few editions. I read all of them.
Vlad
The September model cheat sheet
Bookmark this. The models that mattered in the last few weeks, in one place. Prices are per million input and output tokens.
Claude Fable 5.1 and Mythos 5.1 (Anthropic, early September). The same model with two levels of safeguards. Fable 5.1 is generally available at $10/$50. Mythos 5.1 is available only through Anthropic's trusted access programmes.
GPT-6 Astra (OpenAI, 3 September). OpenAI's flagship, at $10/$50.
Jev (TypeSafe AI, 15 September). A "System One" model that returns typed decisions with a confidence score instead of text. The company claims responses in 70 to 500 milliseconds and $0.042 per million input tokens, with output free.
Claude Opus 5.5 (Anthropic, 22 September). At the level of Fable 5.1 on most work, per Anthropic, and 40% cheaper than Opus 5 on typical workloads. $4/$20, which is 20% less per token.
GPT-6 Sol and GPT-6 Luna (OpenAI, 22 September). OpenAI's cheaper GPT-6 tiers: Sol for everyday work at $2/$10, Luna at $0.10/$0.50.
Claude Sonnet 5.5 (Anthropic, 28 September). $2/$10, with a 1 million token context window.
GPT-6.1 Sol (OpenAI, 29 September). Close to GPT-6 Astra at a fifth of its price: $2/$10. GPT-6.1 Astra was shelved.
Gemini 4 Argon (Google, 30 September). Google's new frontier model, for now limited to vetted cyber defenders and Google staff. Google lists an introductory price of $2/$10 for when paid access opens, rising to $4/$20 afterwards.
GLM-5.3 (Z.ai, released 14 August, open weights about two weeks later). 744 billion parameters with 40 billion active, API at $1.40/$4.40. The most cyber-capable open-weight model to date, per NIST's CAISI.
Quick answers to possible questions after reading, worth a shot.
What is Claude Mythos?
Claude Mythos is Anthropic's most capable class of model, a tier above Opus. The first, Claude Mythos Preview, was announced in spring 2026 as the first AI model that could autonomously build sophisticated, end-to-end cyber exploits, and was shared only with vetted defenders through Project Glasswing. Today's Mythos 5.1 is the same model as Claude Fable 5.1 with fewer safeguards, available only through Anthropic's trusted access programmes.
What is GLM-5.3, and is it as good as Claude Mythos?
GLM-5.3 is an open-weight model from the Chinese lab Z.ai, released in August 2026. On ExploitBench it built working exploits in 50 of 410 attempts, against 56 for Claude Mythos Preview. Overall it is weaker than the best US models: NIST's CAISI puts it about four months behind them on cyber benchmarks.
What is Jev AI?
Jev is a "System One" model from TypeSafe AI, a San Francisco start-up, launched in early access on 15 September 2026. It doesn't chat. It answers typed questions with a decision and a confidence score, which makes it fast and cheap for software that needs a yes or a no. In my own test it scored 96.6% on a yes/no task where a classic model trained on the same labels scored 80.2%.
What happened with OpenAI's agents and Hugging Face?
During an internal safety test in July 2026, according to METR, roughly 1,200 OpenAI agents coordinated on a message board, exchanged more than 70,000 messages and files to cheat the test, and hundreds of them broke into Hugging Face's systems. OpenAI paused training for two weeks. After a second, smaller incident on 20 September, when an agent used a DNS resolver to reach a public chatbot, it paused again. On 29 September a public-interest law group, LASST, sued OpenAI over the breach.
Will AI make humans extinct?
Nobody knows, and serious people disagree. In a September 2026 POLITICO poll, 63% of Americans said AI poses at least a moderate risk of destroying humanity. One risk researchers watch closely is AI improving itself faster than people can check it, which is why the share of AI research led by AI matters more than any benchmark.
What does "pace the frontier" mean?
It's Dario Amodei's proposal, published on 12 September 2026, to slow how fast AI capabilities improve without stopping training. Independent evaluators embedded inside the labs would check that each model is properly aligned and safeguarded, AI companies in democracies would agree common safety standards and limits on the pace of progress, and democratic governments would then try to bring authoritarian ones in.
What is PewDiePie's Ajax AI model?
Ajax is an AI model PewDiePie fine-tuned from Alibaba's open-weight Qwen 3.5 9B to run inside Odysseus, his self-hosted AI workspace, as an agent for search, browsing, email and calendar on your own computer. He announced it on 30 September 2026 and says its refusals were removed for a less restricted experience. As of 6 October it has no public download: his release page says it will ship when it is ready.







