AI agents can now test thousands of new strategies by replaying their own past search results, without paying for new model runs. The method, called Dream-RSI and developed by researchers at Google and DeepMind, lets an agent use its recorded history to figure out which paths would have paid off, then carry the best strategy into its next live search. In tests on program synthesis, math optimization, and GPU kernel writing, the approach reached equal or better results while cutting the number of attempts by up to 2.43 times.
What Dream-RSI actually changes
Dream-RSI does not modify the underlying model. It changes how the agent searches. Self-improving agents typically propose a solution, score it, learn from the result, and try again. The hard part is exploration: deciding which branches to follow, which to run in parallel, and which to abandon.
Most existing approaches fall into one of two camps. A fixed strategy cannot learn from experience, so it keeps hitting the same dead ends. Adapting the strategy during a live run works better and avoids that rigidity, but every new idea has to be tested with a fresh, expensive run from the model and the evaluator. That cost limits how many alternatives an agent can realistically try.
Dream-RSI’s contribution is a cheap way to test alternatives. The agent saves its attempts and their outcomes as it searches, building a recorded search tree. New strategies can then be run against that stored data instead of a live system. Because all the results already exist, thousands of options can be checked without calling the evaluator again.
How the dreaming loop works
The team describes the idea using an analogy. On a first visit to an unfamiliar area, you hit dead ends, double back, and struggle to find a route. Once you have a mental map, you can plan another route without walking every spot again.
Dream-RSI applies that map idea to recorded search histories. Rather than testing a new strategy in a live run, the agent replays it against stored results. The system does not invent entirely new solutions during replay; it tests different decisions inside the recorded search tree. The researchers call this process “dreaming.” The agent plays through thousands of variations and picks the best one before putting it into a live search.
The cycle then repeats. After each live round, the agent uses the recorded results to test better strategies, then applies the improved version to its next run. Throughout, only the search strategy changes. The model that actually generates solutions stays untouched.
What the experiments showed
The team tested Dream-RSI with Gemini 3.1 Pro and Gemini 3.7 Flash on eight tasks spanning three areas. Each comparison used a baseline with the same starting conditions but a fixed search strategy.
One task asked the system to write the fastest possible program for a statistical calculation commonly used in genomics and finance. Dream-RSI’s program ran faster than the established libraries sklearn and glmnet on all six test datasets. With Gemini 3.1 Pro, average runtime fell from 3,587 to 2,931 milliseconds, and the number of attempts dropped from 550 to 317. Dream-RSI also outperformed a competing system called SimpleTES, which needed 51,200 runs to Dream-RSI’s 317 attempts.
The same pattern held for math optimization tasks and for writing efficient GPU kernels, with comparable or better results at much lower computational cost. On two GPU tasks, Dream-RSI matched performance while cutting the number of runs by a factor of up to 2.43. On two others, it delivered up to 2.09 times the performance within the same budget.
A second pattern emerged in how the learned strategy behaved over time. As performance improved, it initially reduced the number of attempts. When progress stalled, it increased the search effort again, which coincided with further gains. The strategy tightened exploration early to save compute, then loosened it when more searching paid off.
Where explicit instructions can backfire
In a follow-up analysis, the researchers tested a different way to use search histories. Instead of replaying them to test strategies, they condensed them into instructions telling the agent where to search. On one GPU task, the version with these instructions performed worse than the version without them.
The team suggests that overly specific directions can narrow the search space too much, keeping it from exploring a broader range of approaches. The finding matters for other systems that turn past failures and successes into reusable instructions, since those instructions can restrict exploration on open-ended search tasks.
How it fits with other self-improvement work
Recursive self-improvement has drawn growing attention in AI research. Google DeepMind introduced AlphaEvolve in 2025, using the same broad principle: Gemini Flash generates code proposals, Gemini Pro analyzes them, and an evolutionary algorithm selects the best versions. Dream-RSI works one level above that process by optimizing the search strategy itself, rather than the candidate solutions.
AutoTTS takes a related approach, using a coding agent to search for algorithms in a simulated environment. Those algorithms decide when a language model should start, expand, or abandon reasoning paths, and the resulting methods beat manually designed methods while using less compute. Meta’s Hyperagents push further, letting agents rewrite the mechanism that controls how they improve.
The researchers have shared code and more details on GitHub for teams that want to study the replay loop or apply it to their own search tasks.
FAQ
What is Dream-RSI?
Dream-RSI is a method from Google and DeepMind that lets AI search agents reuse their recorded past attempts to test new strategies cheaply. It changes only the search strategy, not the underlying model.
How does Dream-RSI cut compute costs?
It stores results from a completed search as a search tree. New strategies are tested against that stored data rather than through fresh, expensive runs, so thousands of alternatives can be checked without calling the model or evaluator again.
What results did Dream-RSI achieve in testing?
Across eight tasks in program synthesis, math optimization, and GPU kernel writing, Dream-RSI reached equal or better performance than fixed-strategy baselines. On a statistical-calculation task it cut average runtime from 3,587 to 2,931 milliseconds with Gemini 3.1 Pro, and on two GPU tasks it matched performance with up to 2.43 times fewer runs.
This article summarizes reporting from the-decoder.com.
