An exploration of DeepSeek founder Liang Wenfeng's strategy of deliberate restraint, continual learning, and open weights in the race toward AGI.
Notes on restraint, world models, and why the most interesting story in AI right now is one of deliberate subtraction.
In 2025, at Spotlight in Amsterdam I put up a slide that read:
2037. AGI. China probably.
It earned a few laughs and a few raised eyebrows, which is roughly what you want from a prediction slide. Now, I want to revisit that line, because a leaked recording out of Hangzhou has quietly turned it into my base case.
The recording is a four-hour investor meeting with DeepSeek's founder, Liang Wenfeng, held on the 20th of May. Fragments have been circulating in Chinese AI and investor circles for a month, and a fuller transcript ran this week, north of a hundred separate answers. Before anyone treats it as gospel, it is unconfirmed by Liang himself, it has been curated down from the full session, and it was delivered in a room full of people deciding whether to wire him billions of yuan.
It makes sense though. Let's get into it.
Liang says no a great deal. No genius myth. No profit maximization. No closed weights. No chasing user numbers. No video generation, no 3D, no world models, no bid to build the next super-app. His justification fits in a sentence: restraint is a strategy, and giving things up raises the probability of reaching AGI.
The roadmap underneath the refusals is clean. Chain of thought was last year's rung. Agents are this year's. After agents comes continual learning, then self-iteration, then AI that helps humans build better AI, and only at the far end, embodiment. He calls the whole climb a gradual singularity, which is the most useful three-word compression of that worldview I have come across.
The rung that should make every researcher put their coffee down is continual learning, and Liang makes it the entry requirement for anything calling itself next-generation. His argument is the clearest version I have seen of a problem the whole field is stuck on. A human picks things up on the job and keeps them. Today's models start every task cold, needing a person to reload all the relevant context by hand, which is close to impossible at any real scale. Until that changes, a model cannot function as labor. And labor, Liang points out, is what people actually want from this technology. Nobody was asking for another computer interface.
The single no that puts Liang on a collision course with the rest of the field is world models. He waves them off directly: world models have little bearing on the ceiling of intelligence, and multimodality, however much it matters to products and consumers, remains one component of a system, sitting some distance below intelligence itself.
Set that against two of the most serious people working today.
Yann LeCun has bet his second act on the opposite view. He now runs AMI Labs out of Paris, which raised north of a billion dollars in March to chase exactly this, and his group has spent 2026 turning the JEPA thesis into something with mathematical teeth. Their May paper, When Does LeJEPA Learn a World Model?, proves that the architecture can recover the true hidden variables behind raw observations, a property called linear identifiability. The proof comes with a sharp condition attached: it holds when those latent variables are Gaussian and evolve under stable, predictable dynamics, and the Gaussian case is the only one where the guarantee survives. A companion benchmark then showed current systems buckling under minor visual shifts. So the world-model camp now has both a target and a clear measure of the distance still to cover.

Demis Hassabis sits at the other pole, all in on world models and forever pointing at Veo as evidence. This is where my skepticism kicks in. A system that earns its understanding of the world by predicting pixels is paying rent on an enormous amount of detail that most tasks never ask about.
Here is the third position, and the one I stood on that stage to argue. Evolution already ran this experiment, billions of times. A fly has a world model. So does a mouse, a cat, an ape, a person. Each of those is enough to dodge a predator, find food, contest a patch of ground, win a mate. These are not trivial feats. And not one of those models is a full simulation of physics. Complete world understanding was never on the menu, because nothing in the environment ever paid for it. The environment paid for cheap, quantized, compressed models with narrow attention and quick reflexes. The intelligence scales all the way from insect to human, and at every rung the world model stays light enough to run in real time on a few watts of wet tissue.
Two properties made that possible, and both are missing from the machines. Biological minds are always on, and they fine-tune continuously against a stream that never stops. This is where Liang's continual-learning rung and my light-world-model argument turn out to be the same insight approached from two sides. A model light enough to run continuously is a model that can keep learning while it runs, and a model that keeps learning while it runs has no need to hold a bloated simulation of everything in its head. The lightness and the always-on are the same coin. Veo-style generative world models are heavy. Light world models fly, in every sense of the word.
Pull the camera back from Hangzhou and the individual lab stops being the story. The ecosystem is.
China at NeurIPS 2025:

By May, Chinese open-weight models were carrying roughly 61% of all tokens routed through OpenRouter, the largest neutral model router going. Alibaba's Qwen family crossed a billion cumulative downloads on Hugging Face in March, the fastest any model family has reached that mark, and for stretches accounted for more than half of all open-source model downloads on the planet. DeepSeek's V4-Pro, a 1.6-trillion-parameter mixture-of-experts model with around 49 billion active, has been trading blows with the closed frontier on agentic coding. Moonshot's Kimi K2.6 has demonstrated autonomous coding runs lasting half a day with thousands of tool calls. Zhipu's GLM line ships under MIT with a million-token context, and one of its recent flagships was trained end to end on a cluster of roughly 100,000 Huawei Ascend chips, with no Nvidia in the loop at all.
China at ICLR 2026:

That last detail is the one to sit with. Cost discipline plus hardware independence is a different kind of moat from the one American labs have been digging. Liang is explicit that when two models are equal in quality, the contest is decided first on cost, then on time to ship, and only then on user experience, which he declines to treat as the deepest defense. It is a commoditization thesis, and it happens to crown the axis where his own house is strongest. The tell does not make it wrong. His releases have repriced this market more than once.
The open-weight strategy sits inside the same logic. Liang frames a mid-size lab as occupying a rare sweet spot: too small to be a giant tripping over its own org chart, too serious to be a startup with no frontier ambition. Give the weights away, keep nothing stronger hidden inside, and you erode your competitors' margins at close to zero cost to a lab that shrugs at API revenue in the first place. The idealism and the ruthlessness point the same direction, which is exactly why the strategy holds.
So, the slide. Why did I write China probably, and why do I still believe it a year on.
The structural case is that cost discipline compounds, an open ecosystem compounds, the talent was never the bottleneck, state capital is now flowing in, and, above all, a willingness to subtract is rarer and more valuable than a willingness to add. Everyone can add. Almost nobody with a billion-dollar valuation can look at video, 3D, world models, and a consumer super-app and say we will do none of these, on purpose.
The counter-case is real and I hold it in the same hand. The compute gap has not closed. And the money is now arriving in volume: a first round near a 52-billion-dollar valuation, then early talk of a second at around 71 billion, with state money holding voting rights that most other investors gave up. The interesting question for the next two years is whether the restraint survives the capital. Liang left the sharpest possible marker for himself on that. His closing line to the room was that the moment your vision becomes taking more, you have already lost. That is a beautiful thing to say to your own future cap table. It is also about to be tested by it.
Here is the part that keeps the 2037 bet grounded, and keeps it fun. Nobody has the primitive yet. Spend an afternoon in the deeper corners of the field, the long cognitive-science conversations, the first-principles arguments about compression and attention and what a world model even is, and the mood is unmistakable. Everyone is toying with the pieces. Nobody has clicked them together. The light-world-model, always-on, continually-learning breakthrough is sitting there unclaimed.
Unclaimed things get claimed from unexpected places. It could come from DeepSeek's relaxed, no-KPI research floor. It could come from LeCun's Paris theorem factory. It could just as easily come from a kid on a dev box who picks an architecture because it fits on the machine and lands inside the edge of their patience, which is roughly how Alec Radford settled the shape of the first GPT. The disruption that resets all of this may already be training overnight on a single rack in a bedroom somewhere.

I said 2037 in Amsterdam and I am not moving the date. The direction of travel is what I would watch, and right now the most disciplined version of it is being run, out loud and mostly in the open, from China.