AllSpark releases Iris-mini and Iris-pro, the strongest open-weight search agents in their class

Iris-mini and Iris-pro search agent benchmark results

Written by

in

Two new open-weight search agents from Chinese lab AllSpark, Iris-mini and Iris-pro, deliver the strongest results in their respective size classes on four established web research benchmarks, according to the team’s published paper. Both models were trained on questions reverse-engineered from the link structure of web pages, with the training data and models also improving performance on tasks they were never trained for, including general tool use and office work.

Search agents built on language models research the web on their own. They need to understand the question, decide what to search for, interpret the results, and judge when they have gathered enough evidence for an answer. On established benchmarks, leading AI systems of this kind mostly use the web to confirm knowledge they already picked up during training, and how much of the work the model itself is doing, versus the scaffolding around it, remains contested.

What AllSpark released

Iris-mini has 35 billion parameters and Iris-pro has 397 billion. Both build on Qwen-series models, specifically Qwen3.6-35B-A3B and Qwen3.5-397B-A17B, and work with a 256,000-token context window. The team published the model weights on Hugging Face and the code on GitHub. The initial release includes the Iris Harness with the agent loop, tools, context management strategies, and the four benchmarks with evaluation. The harness runs against any OpenAI-compatible endpoint. The data construction and training pipelines are planned for later release.

How the training data is built

The training pipeline constructs tasks backward from the link structure of web pages. Starting from a seed page and its outgoing links, it builds a graph of terms and relationships, then generates a multi-step question whose answer requires chaining several connected steps. Every term except the final answer is replaced with a paraphrase so no clue can be resolved through a simple text search. The agent has to reason, not just look things up.

Only questions that a reference model cannot solve without tools but can solve with the right sources make it into the dataset, which keeps the tasks both hard and clearly verifiable.

Two-stage filtering weeds out bad training data

A stronger teacher model generates solution paths made up of reasoning, search queries, and results. These paths go through two rounds of filtering. The first checks the full path for correctness, repetition loops, and search depth. The second is a step-by-step review by a judge model whose criteria were derived from the data itself rather than set by hand, according to the paper. After that, the model is improved through reinforcement learning against a live web search. The judge model and result summaries run inside the training cluster, powered by the team’s own large Qwen model, so training does not depend on external services.

Supervised fine-tuning and reinforcement learning alternate in a process the authors call “SFT-RL climbing.” The hardest solved tasks and the most efficient solution paths from each round feed back into the next training cycle.

Why context management may matter more than model differences

The team argues that runtime context management on common benchmarks often makes a bigger difference than the reported gaps between systems. During long research sessions, the context can fill up before the agent has resolved all sub-questions. Tricks like discarding the conversation history extend the research artificially but say little about the model’s actual quality.

To isolate the effect, the team tests every benchmark with and without context management while keeping tools, context limits, and the judge model constant. Results reported only with management turned on cannot be cleanly split into what comes from the model and what comes from the scaffolding around it. The Iris scores also come from a single agent, with no helper agents and no extra verification steps at the end.

Results across four benchmarks

Testing covered BrowseComp, which tests the ability to find rare facts from indirect clues, its Chinese counterpart BrowseComp-ZH, DeepSearchQA, which evaluates the completeness of retrieved evidence, and Humanity’s Last Exam, which poses academic questions at expert level. With context management turned on, Iris-mini scores 82.2, 84.8, 86.9, and 52.3 on the four benchmarks respectively. Iris-pro reaches 88.6, 85.1, 92.9, and 56.4.

In the smaller class, Iris-mini leads on three of four benchmarks and beats the next-best model, XYZ-Aquila-mini, on BrowseComp by 3.4 points, though it trails on DeepSearchQA. Iris-pro leads or ties in the larger class and sometimes approaches systems that need far more compute, according to the authors.

Context management has a much bigger effect on the smaller model, boosting BrowseComp scores by up to 21.2 points. The reason is not a smaller token budget but faster consumption, according to the paper. Iris-mini needs more steps for the same tasks and hits the context limit more often. On Humanity’s Last Exam, the gains are smaller because the benchmark leans more on domain knowledge and academic reasoning, where web search plays a supporting role.

The best scores come from combining history discarding with a second attempt. If the first try fails, the system condenses it into a short note that records what was already checked and ruled out. That note gets appended to the task for the next run.

When the ground truth is wrong

In the paper’s appendix, the team describes a case where its agent was marked wrong even though the answer was backed by the source material. A question in BrowseComp-ZH targeted the series “Game of Thrones.” The agent answered “Bolton,” but the ground truth said “Lannister.” The character in question, Sansa Stark, actually marries Ramsay Bolton in her second marriage. The agent’s answer was correct. The team says contradictions like these between ground truth and source material motivate them to build better benchmarks.

Search as a foundational skill

Beyond search, the authors report an unexpected side effect. Both the generated training data and the specialized models improved performance on tasks they were never trained for, including general tool use and office work. The team suggests that search may function more as a foundational skill than a narrow specialty, since the learned behavior helps wherever an agent has to work with incomplete information.

FAQ

What are Iris-mini and Iris-pro?

Iris-mini and Iris-pro are open-weight search agents released by Chinese lab AllSpark. Iris-mini has 35 billion parameters built on Qwen3.6-35B-A3B, and Iris-pro has 397 billion parameters built on Qwen3.5-397B-A17B. Both use a 256,000-token context window and, according to the team’s paper, lead their respective size classes on four web research benchmarks.

How were the Iris models trained?

The training pipeline builds multi-step questions backward from the link structure of web pages, paraphrases every clue except the final answer, and filters the resulting solution paths through two rounds of review, a full-path check and a step-by-step judge model. Supervised fine-tuning and reinforcement learning against a live web search alternate in a process the team calls SFT-RL climbing, with the hardest solved tasks and most efficient paths fed back into the next cycle.

Where can the model weights and code be downloaded?

The model weights for Iris-mini and Iris-pro are available in a collection on Hugging Face, and the code is on GitHub. The initial release includes the Iris Harness with the agent loop, tools, context management strategies, and the four benchmarks with evaluation. The data construction and training pipelines are planned for later release.


This article summarizes reporting from the-decoder.com.