This week, the world’s biggest open-weight model became too popular for its own creator to serve. That alone says a lot about where AI is heading.
Models are getting strong enough. Now the pressure is moving elsewhere: to the GPUs needed to run them, the CPUs keeping agents moving between model calls, the organizations struggling to turn saved time into actual value, and the governments deciding which models people may use at all.
From the biggest open-weight models and AI’s new limitations — GPUs, CPUs, and more — to questions of governance and society, this issue is about what happens when model capability stops being the only constraint.
Open weights are changing their purpose. And open source is getting to big to run
The hottest topic of the week was Kimi K3, the biggest ever open-weight model built by Chinese Moonshot AI. But these wasn’t only Kimi K3.
We looked at two very different open-weight strategies:
Moonshot’s Kimi K3 is pushing the capability frontier
Thinking Machines’ Inkling bets that customization may matter more than topping every benchmark.
Together, they show that open weights are no longer just about cheaper alternatives. They’re becoming infrastructure for shared innovation and adaptation.
Then, just 3 days later, the story took another turn. Kimi K3 sold out after overwhelming Moonshot’s GPU capacity. Demand was enormous. At the same time:
Xi Jinping publicly backed open-source AI
Alibaba teased another massive open-weight model, and
Reports emerged that the US may ban Chinese AI models.
The bottleneck shifted again: from model capability, to compute, and perhaps next — to permission.
Loop engineering lasted about six weeks before the industry said it was dead! And moved on to graphs. But the terminology is moving faster than the technology.
That’s why we decided to separate the useful engineering idea from the hype. In the weekly digest, we explain when an agent loop genuinely grows into a larger graph of models, tools, evaluators, code, and human approvals — and why “graph” can also mean a control flow, knowledge graph, execution trace, or improvement system. It also fact-checks the viral claim that Microsoft, Stanford, and Anthropic had all adopted graph engineering with dramatic accuracy and cost gains.
Employees are saving time with AI, yet many companies still can’t find the impact in revenue, costs, or productivity metrics. The problem is that faster tasks do not automatically create business value.
We explain the four-stage Capacity-to-Outcome Chain: task-level gain → released capacity → organizational absorption → business outcome. It shows where AI returns disappear, why scattered minutes are often impossible to reuse, and how workflow redesign, management decisions, and clear ownership determine whether saved time becomes greater volume, faster delivery, better quality, lower risk, or simply more meetings.
What happens if AI makes intelligence, services, and even some forms of labor dramatically cheaper, but society is not ready for the consequences?
Here we look at a Post-Necessity Institute: a place to test public AI access, new civic roles, education, cash support, and other systems that could help people build meaningful lives beyond survival-driven work.
The main idea is to start with small, measurable experiments now, before abundance turns into an institutional emergency.
And more about the bottleneck →
We’ve spent years measuring AI by how fast models generate tokens. But once agents start writing code, launching sandboxes, running tests, and using tools, the bottleneck shifts beyond the GPU.
In this video, we unpack NVIDIA’s new Vera CPU and the bigger idea behind it: AI performance is becoming a system problem, not only a model problem. We look at why CPUs matter for agent loops, what NVIDIA’s benchmarks really show, where the claims deserve skepticism, and whether Vera represents a genuine shift in AI infrastructure — or simply a clever way to sell the entire data center.
And since we are speaking a lot about hardware like GPUs and CPUs, we thought that you definitely need this full practical guide on AI chips →


