AllSpark Research has released two search agents, Iris-mini and Iris-pro, that repeatedly traverse the web to find answers. They are open-weight models in the 35B and 397B classes, and the research team reports that each achieved the strongest overall results in its size class across four search benchmarks. But the trained weights are not the only thing driving the scores. Context management, which discards the history that bloats during long searches, produced a larger gap on the same test than scaling up the model did. What Iris raises is a question: can a search agent's score be read as the capability of the model alone?

AD

What "strongest in its class" means, and what is not yet public

Iris-mini is built on Qwen3.6-35B-A3B, with 35B total parameters, of which 3B are active at inference. Iris-pro is built on Qwen3.5-397B-A17B, with 397B total and 17B active parameters. Both are Mixture of Experts (MoE) models with a 256K-token context length. The weights are distributed on Hugging Face under the Apache 2.0 license.

The representative figures in the preprint dated September 3, 2026 are single-attempt-per-question results with "discard-all," which wipes the long history, enabled. Iris-mini scored 82.2 on BrowseComp, 84.8 on BrowseComp-ZH, 86.9 on DeepSearchQA, and 52.3 on Humanity's Last Exam (HLE). Iris-pro scored 88.6, 85.1, 92.9, and 56.4, respectively. DeepSearchQA is measured by F1 and the other three by accuracy, and the HLE figure covers only the 2,158 text-based questions.

"Strongest" has limits. Mini beat its 35B-class peers on three metrics (BrowseComp, BrowseComp-ZH, and HLE) but fell short of XYZ-Aquila-mini's 89.5 on DeepSearchQA. Pro led the roughly 400B class on three metrics and tied XYZ-Aquila-pro at 85.1 on BrowseComp-ZH. If one includes the high-compute configurations listed separately in the paper and the closed frontier models, Iris is not first on every metric.

What is public is also split in two. GitHub hosts Iris-Harness, which bundles the agent loop, the two tools for search and page retrieval, history management, the four benchmarks, and the graders. With enough compute to run the weights, the evaluation setup can be replicated. However, the README lists the data construction and training pipeline as "coming soon." What is in hand now is a trained climber and a measured course, not the blueprint of the training facility that produced the climber.

Iris's training approach runs in the opposite direction from the usual one, starting with how questions are collected. It first selects a seed page describing the object that should become the answer, then follows the links leading out of that page to build a local web graph. Rather than use full page text as is, it compresses the pages into an entity graph of objects and relations, and assembles questions that reach the answer only by following multiple relations in sequence.

That alone would let through questions that can be solved by pasting a proper noun into a search box. So the names and aliases of entities other than the answer are rewritten into descriptions that uniquely identify them. For example, instead of keeping a person's name, the question asks the agent to identify that person from attributes and relationships found on another page. This removes string-matching shortcuts and turns the act of linking evidence across multiple pages into the thing being trained.

Generated questions are filtered on two conditions: a reference model cannot answer them without tools, and it can answer correctly when given the entity graph as evidence. The former guarantees difficulty, and the latter confirms that the answer is unique and solvable. Questions that are merely hard and broken, as well as those solvable from memory without searching, are dropped.

Next, a strong teacher model produces solution trajectories in ReAct format, alternating reasoning, search, page retrieval, and observation. Iris does not learn from correct trajectories unconditionally. Beyond the correctness of the final answer, it checks for loops that repeat the same sentence or search, endless thinking, consecutive calls with identical arguments, and shallow searching. A judge also labels each turn KEEP or MASK, excluding locally bad turns from the training loss. MASK is capped at 10% of a trajectory, and masked turns are not removed from the surrounding history. The design thins out only the bad moves mixed into successful examples.

AD

Alternating SFT and RL, and resuming long searches midway

The pipeline is not a single round of supervised fine-tuning followed by one round of reinforcement learning. Iris alternates supervised fine-tuning (SFT) with reinforcement learning (RL) on live search. The research team calls this procedure SFT–RL climbing. RL discovers successful trajectories that are rare for the current model, and SFT imprints those trajectories directly into the next model.

The problems sent back to the next SFT round are not the ones that are easily solved on every attempt. They are problems where, across a group of RL rollouts, the accuracy is above 0 but still at or below 1/2. Among successful trajectories containing at least a certain number of tool calls, only the shortest is kept, which avoids lucky one-shot answers and needlessly long searches while raising difficulty to match where the model currently stands. Rather than cycling through fixed data many times, the model carries the climbing routes it only barely found into its next lesson.

RL with long searches has another problem that leaves compute idle. Even after most rollouts finish, if a few long ones remain, the entire synchronized training step has to wait. Iris halts the remaining requests midway, saves the completed prefix, and resumes from it in the next step. Because the prefix and the continuation are produced under different weight generations, truncated importance sampling corrects the mismatch. The method keeps unfinished long trajectories rather than discarding them, and allocates compute with roughly 2x oversampling as headroom.

Reward judging and summarization of retrieved pages use Qwen3.5-397B-A17B running inside the training cluster. This avoids dependence on external APIs, but it does not mean training closes around the search model alone. It works only when a large judge and summarizer, live search, and the RL infrastructure are operated together as one training system. The figure of 3B active parameters for Iris-mini alone does not allow one to estimate training cost or ease of reproduction.

History discarding moved scores more than model size did

Discard-all at evaluation time works as follows: when the running context reaches a set threshold, the accumulated tool history is erased and the search continues from the original question. For retries after failure, what was investigated and ruled out is compressed into a short memo and passed to the next search. Neither changes the model's weights. What changes is the effective length of search available and the number of attempts.

Table 2 of the paper reports values within Iris with the search and retrieval tools, the 256K-token cap, the judge, and the maximum number of turns held constant, varying only history management. Iris-mini's BrowseComp score was 64.7 without management and 82.2 with discard-all, a difference of 17.5 points. By comparison, moving from mini to pro without management raised the score from 64.7 to 72.6, a difference of 7.9 points.

On BrowseComp, the 17.5-point increase from adding discard-all to Iris-mini was about 2.2 times the 7.9-point difference from scaling from mini to pro without management.

The figure of about 2.2 is 17.5 divided by 7.9, rounded to one decimal place. It uses three cells from the same BrowseComp, single-attempt, same-tools, same-judging setup. However, the 7.9 points cannot be called a causal effect of parameter count alone. Mini and pro differ in base model as well as scale, and the score gap cannot be converted into inference cost or practical research quality. Still, it is clear that the surrounding machinery is too large to label 82.2 purely as the capability of a 35B model.

Combining discard-all with retries lifts mini to 85.9 and pro to 90.3. But retries require additional full searches and increase compute. The research team itself positions this as a configuration to probe the attainable ceiling, and adopted discard-all without retries for the representative table. Being able to choose a high number is not the same as choosing a number suited to comparison.

AD

Don't treat the four numbers as one yardstick

The effect of history management varies greatly by task. For Iris-mini, adding discard-all and retries improved BrowseComp by up to 21.2 points. BrowseComp-ZH, HLE, and DeepSearchQA also improved, but by less than BrowseComp. The relationship is not simply that lower-scoring tasks benefit more from history management: HLE, the lowest-scoring task without management, gained less than half as much as BrowseComp.

The difference lies in what each benchmark measures. BrowseComp asks for a short answer by locating a long-tail target from indirect, mutually constraining clues. Searches tend to run long, and once history uses up the 256K tokens, context length becomes the limit on action. DeepSearchQA measures, via F1, the ability to comprehensively collect the required items, and HLE strongly demands specialized knowledge and reasoning that search alone cannot replace. Discarding history to extend search time does not make up for gaps in knowledge or reasoning.

BrowseComp-ZH shows a different kind of ceiling. Of 289 questions, mini with discard-all plus retries, pro with discard-all, and pro with discard-all plus retries all converged on 246 correct, or 85.1. The paper notes that this may indicate a constraint other than model capacity, and in an appendix examines one question where the official answer was "Lannister" and the agent answered "Bolton." The claim is that tracing the character's second marriage in the story supports the latter. But this is a single example, and the label error rate across all 289 questions is unknown.

Evaluation used a single ReAct agent, no auxiliary agents, no additional test-time verification, and one attempt per question. According to the paper, benchmark distribution pages were excluded from search results and refused in page retrieval, and cases where a known URL was produced directly were stopped by post-hoc checks. Transparency is high. Even so, the comparison figures for other models include numbers reported by different teams under different history management. Comparisons between conditions within Iris are easy to make, but a side-by-side ranking table is not a pure contest of model weights.

What can be reproduced is the evaluation; what cannot yet be reproduced is the training

What a third party can verify today is limited to connecting the public weights to Iris-Harness and running them with the same two tools, history management, benchmarks, and graders. Even then, the live web keeps changing. Unless the search index, retrieval time, summarizer and judge, and inference budget are recorded, the same configuration file will not necessarily reach the same evidence. Replication needs two kinds of test: one comparing weights and harness on a fixed web snapshot, and one comparing them on the live web, including real-world variability.

Many pieces are also missing for reproducing the training results. The total number of training questions, the source web corpus, the size of each SFT and RL round, total compute, the specific prompts for the reward judge and summarizer, and the data generation and training code have not been published. There are also reports that the synthetic data and specialized models for search positively affected general tool use and office work, but the preprint describes this only as an initial observation and provides no detailed figures. The interpretation that search is a general foundational capability remains a hypothesis.

Iris has widened the unit of evaluation for search agents from weights to systems. The next test is whether, once the data construction and training pipeline are released, another team can reproduce the same training procedure and confirm its effect on practical tasks, including inference cost and citation integrity. If those conditions are met, what Iris showed can be judged to be not a one-time high score but a process for rebuilding search capability.