# Recursive self-improvement, under the hood

Published: 2026-09-18

Google DeepMind published "Dream-RSI: Recursive Self-Improvement through Evolving Worlds" — and the coding model inside it never changes. Under the hood: what the phrase meant for 60 years, what actually loops in Dream-RSI (a Python exploration policy that "dreams" over recorded search trees), and why 317 agent calls instead of 550 is a discount, not a takeoff. Verdict: NEEDS REVIEW.

Canonical: https://thedailydiff.dev/video/2026-09-18-self-improving-ai/

## What this video covers

- 1965 → 2008: I. J. Good's "intelligence explosion", Yudkowsky's seed AI — a machine that rewrites its own weights
- 2025: AlphaEvolve — 48-multiplication 4×4 matrix multiply, 0.7% of Google's compute recovered, Gemini kernels 23% faster
- 2026: Dream-RSI — the policy improves, the coder, judge and rewriter are all frozen; the prompt says "Do not solve the scientific task"
- X: "Google just cracked recursive self-improvement" (819 likes) vs HN: "calling this RSI seems misleading"
- Numbers used: §4.1 Lasso 550 → 317 agent calls, 3587 → 2931 ms; Table 1 circle packing 2.635983 = AlphaEvolveV2; §4.3 KernelBench 2.43× / 1.79× fewer generations, up to 2.09× faster; Appendix B.2 prompt (all from the PDF above)

## Transcript

0:00 Recursive self-improvement is the idea that an AI makes itself smarter, then uses the smarter self to do it again, and this week Google published a paper with those two words in the title, in which the model's weights never move. Three numbers. Anthropic's CEO says this loop has been running since the summer, which is his reason to slow the whole industry down. The tweet announcing that Google just cracked it collected eight hundred likes in a day.

0:24 And the paper's own headline result is 317 model calls instead of 550, which is a discount, not a takeoff. In three minutes: where the phrase came from, what actually loops inside Dream-RSI, and how to read the next paper with RSI on the cover. This is The Daily Diff, under the hood. 1965. I. J. Good, a Bletchley Park statistician who worked next to Turing, writes that an ultraintelligent machine could design even better machines,

0:52 calls it an intelligence explosion and the last invention man need ever make, provided the machine is docile enough to tell us how to control it, which is the provided-that people stop reading after the comma. For forty years the phrase meant exactly that: a system that edits its own source or weights and comes out smarter every pass. Eliezer Yudkowsky's 2008 essay made it the load-bearing word of AI safety, and nobody had a machine to test it on, which is the ideal condition for a definition. Then in May 2025 DeepMind shipped AlphaEvolve,

1:23 a Gemini agent that proposes code, gets it scored, keeps the winners and mutates them again. It multiplied four-by-four complex matrices in 48 steps, beating a record Strassen set in 1969, and it recovers about zero point seven percent of Google's worldwide compute, which at Google's scale is a data centre found under the sofa. Gemini was now helping train Gemini, and the phrase quietly moved from safety blogs into press releases. Dream-RSI bolts a controller on top of that loop.

1:52 A policy, an actual Python program, decides which branches of the search tree to expand, how many in parallel, and when to give up. Gemini writes the candidate code, and a fixed evaluator scores it. Every run leaves behind a tree of what was tried and what it scored. That tree becomes a replay simulator: a second, equally frozen Gemini rewrites the policy, and the new policy is tested against the recorded outcomes instead of on real GPUs.

2:16 The paper calls this dreaming, and one real run pays for thousands of dreamed ones at zero execution cost. The winner goes back online, the pool of recorded worlds grows, repeat. And the prompt in Appendix B.2 tells the self-improving AI, in writing: edit only this one file, and do not solve the scientific task. Measured. On the Lasso solver task, fixed exploration spent 550 Gemini calls to reach three point six seconds, Dream-RSI spent 317 to reach two point nine, and both beat scikit-learn.

2:46 On KernelBench it reached the same kernel speed with about two point four times fewer generations. And on circle packing it scored two point six three five nine eight three, which is precisely, to six decimals, what AlphaEvolve version two scored last year. So X read the title: Google just cracked recursive self-improvement. Hacker News read the PDF: calling this RSI seems misleading. And the same week a competitor's CEO said the loop had been running since summer and everyone should slow down.

3:13 Three sentences, the same two words, three different machines. So, Monday, how to read the next RSI paper. One: ask what object changes, the weights, the code, or the search settings; Dream-RSI changes the third. Two: ask who is frozen; here it's the coder, the judge and the rewriter, all three. Three: ask what the headline number is a unit of. Compute saved is optimisation; a benchmark the model couldn't reach before is the takeoff, and nobody has published that one yet.

3:40 Verdict, under the hood: needs review. The loop is real, the savings are measured, and the two words on the cover describe a machine that isn't in the paper. If you'd rather read this than hear me say it, the diff lands in your inbox every morning, free at the daily diff dot dev, link below. Tell me what to open up next in the comments. And that's the diff for today. I'm Niko from Axrisi.

4:00 Merge responsibly.

## Sources

- [Paper: Dream-RSI (Zheng et al., Google / Google DeepMind / UMD / UVA, 14 Sep 2026)](https://arxiv.org/abs/2609.14858) — arxiv.org
- [Code (293 stars, no license) — https://github.com/zhengkid/Dream-RSI · project page](https://dream-rsi.com/) — dream-rsi.com
- [Hacker News thread (169 points)](https://news.ycombinator.com/item?id=49726955) — news.ycombinator.com
- [Hugging Face papers (278 upvotes)](https://huggingface.co/papers/2609.14858) — huggingface.co
- [I. J. Good, 1965, "Speculations Concerning the First Ultraintelligent Machine"](https://en.wikipedia.org/wiki/I._J._Good) — en.wikipedia.org
- [Eliezer Yudkowsky, "Recursive Self-Improvement", LessWrong, 1 Dec 2008](https://www.lesswrong.com/posts/JBadX7rwdcRFzGuju/recursive-self-improvement) — www.lesswrong.com
- [Google DeepMind, AlphaEvolve (14 May 2025): 0.7% compute, 48 multiplications, 23% kernel speed-up](https://deepmind.google/discover/blog/alphaevolve-a-gemini-powered-coding-agent-for-designing-advanced-algorithms/) — deepmind.google
- [Dario Amodei, "We Must Pace the Frontier" (12 Sep 2026)](https://darioamodei.com/post/we-must-pace-the-frontier) — darioamodei.com
- [Tweet](https://x.com/thesupermannx/status/2100153430849499395) — x.com
