The real reason AI releases slowed down
If you’ve been following AI, you’re probably used to seeing massive leaps forward every couple of weeks, but it's been a while since the last set of big model releases.
Get ready: the dam is about to break. Almost every major frontier lab is sitting on massive upgrades that are right around the corner.
First up is Anthropic. As I was writing this, they released Fable 5.1. From my early testing, this is not a small bump in capability. It's a big leap forward, and (I've heard from sources I trust) it can handle work that takes multiple days to do, completely autonomously.
OpenAI is in a similar position with its upcoming model, Astra. The details and research coming out around Astra suggest a massive leap in its ability to handle projects that take days or weeks, solve math problems that have stumped people for decades, and (most importantly for you!) use a computer like a human.
At the same time, Elon Musk mentioned on X that Grok is about to get dramatically better with the release of Grok 4.7. Considering the massive jump Grok 4.6 made in coding and real-world task benchmarks, 4.7 should firmly cement itself among the very best models available (and make tools like Grok Bot ridiculously capable).
So why did things feel quiet for a minute?
A lot of people assumed we hit a wall on progress, but the bottleneck hasn’t been capability. It’s actually been getting these models cleared to release safely.
These models are finally reaching autonomy and cybersecurity thresholds where they can cause real-world damage if used in a malicious way. If a lab wants to make one of their most capable frontier models public, it's no longer just running internal benchmarks, testing with a few outside folks, and hitting deploy. It now involves extremely extensive testing, external evaluations, and government coordination before anything goes public.
The models are built, the safety work is likely wrapping up, and we’re about to see just how good the new models have gotten.
Tokens as currency
But there’s another bottleneck I’m starting to feel personally: having enough tokens to put these models to work.
Yesterday, I was using Fable 5.1 to build a new version of my Gauntlet Loop for a project I’ll hopefully share soon. I expected to have a first version within a few hours, but instead, I hit my usage limits. So I bought another $200 Claude subscription so I could keep going. With Claude running so many agents at once, it burned through that account’s available token allowance in about 20 minutes. I bought another, then had to dramatically reduce how much work it was doing in parallel (and scale back my ambition) to keep going.
That’s a really strange feeling... the model could do more of the work, but I didn’t have nearly enough capacity to let it. As these systems get better, the amount of AI you can afford starts to determine how many things you can try, how quickly you can build, and how ambitious you can be. It starts to look a lot like being able to hire a bigger team.
I think we’re in the early innings of a pretty significant divide between people who can afford to keep these models working and people who can’t. And the advantage compounds: more access means more hands-on time with the models, which means getting better at using them. I’m incredibly excited about what’s coming, but yesterday was the first time in a while that I had to make a project smaller because I couldn’t keep the AI working on it.
Clear enough for my dad. Sharp enough for the frontier. If a week is boring, you don’t hear from me.