AllSpark releases open-weight search agents Iris-mini (35B) and Iris-pro (397B) with training recipe and benchmark results
Chinese lab AllSpark released two open-weight search agent models, Iris-mini (35B parameters, built on Qwen3.6-35B-A3B) and Iris-pro (397B parameters, built on Qwen3.5-397B-A17B), both with 256,000-token context windows, along with a paper describing the training recipe, an evaluation harness ("Iris Harness") with agent loop, tools, and context management strategies, and results on four benchmarks (BrowseComp, BrowseComp-ZH, DeepSearchQA, Humanity's Last Exam). Model weights are on Hugging Face and code is on GitHub; the data construction and training pipelines are not yet released (planned for later).
Entities: AllSpark, Iris-mini, Iris-pro, Qwen3.6-35B-A3B, Qwen3.5-397B-A17B, The Decoder
0 primary
What happened
Chinese lab AllSpark released two open-weight search agent models: Iris-mini (35B parameters, built on Qwen3.6-35B-A3B) and Iris-pro (397B parameters, built on Qwen3.5-397B-A17B), both with 256,000-token context windows. Alongside the weights on Hugging Face and code on GitHub, AllSpark published a paper describing its training recipe and an evaluation harness, with results reported on four benchmarks: BrowseComp, BrowseComp-ZH, DeepSearchQA and Humanity's Last Exam. The data-construction and training pipelines themselves are not yet released; AllSpark says those are coming later.
Why it matters
Developers and researchers building search or research agents now have two more open-weight base models to test, with a documented harness for agent loops, tools and context management they can inspect or copy. That is genuinely useful and immediately actionable, since the weights and code are real and downloadable today. But both models are fine-tunes of existing Qwen bases rather than new foundation models, so the impact is incremental (better tooling and technique on top of known architectures) rather than a shift in who controls the underlying compute or model layer.
What is noise
The headline and the "strongest in their class" claim are AllSpark's own self-reported superlative, relayed by The Decoder without independent verification or benchmark links in the coverage. The claim that search skills "generalise unexpectedly" to general tool use and office work, and that search may be a "foundational" capability, is speculative framing from the lab's paper, not a demonstrated or falsifiable result. Missing from the coverage: any comparison methodology, third-party evaluation, or explanation of what happens once the promised training pipeline is (or isn't) released.
Watch next
- 01Whether AllSpark releases the promised data-construction and training pipelines, and whether independent developers can reproduce the claimed benchmark scores using them
- 02Third-party or independent benchmark runs of Iris-mini and Iris-pro on BrowseComp, BrowseComp-ZH, DeepSearchQA and Humanity's Last Exam, since current results are lab-reported only
- 03Uptake signals on Hugging Face and GitHub (downloads, forks, derivative fine-tunes) over the next 4-8 weeks to see if this becomes a reference model or is quickly superseded
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- AI systems outperform expert humans in persuasive communication22 Jun 202681
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680