Researchers publish taxonomy and benchmark for LLM tool hallucination, showing 322-476 hallucinated tool calls across ten hosted models and MCP servers
Researchers released a paper defining a five-class taxonomy (H1-H5) of LLM tool hallucination (agents calling nonexistent tools or passing undeclared arguments), proposed a training-free closed-world 'Resolution Rung' resolver (registry membership plus signature check), extended the taxonomy to Model Context Protocol (MCP) merged-namespace failures (M1-M5), measured 322 genuine hallucinations across ten hosted models under two invocation surfaces and 154 hallucinations on a live MCP surface, and released a versioned benchmark (Hallucinated-Tools Benchmark, HTB) for comparing resolvers.
Entities: Resolution Rung, Hallucinated-Tools Benchmark (HTB), Model Context Protocol (MCP)
0 primary
What happened
Researchers published an arXiv preprint proposing a five-class taxonomy (H1-H5) for LLM "tool hallucination" (agents inventing nonexistent tools or passing undeclared arguments), extended it to five more failure modes specific to Model Context Protocol (MCP) merged-namespace setups (M1-M5), and released a versioned benchmark (HTB) for testing fixes. They measured 322 hallucinated tool calls across ten unnamed hosted models under two invocation surfaces, plus 154 on a live MCP setup, and propose a training-free "Resolution Rung" resolver based on registry membership and signature checking.
Why it matters
This targets a real gap: current tool-selection and tool-gating defenses assume a call refers to a real tool, so none catch outright hallucinated calls before any safety gate runs. The finding that a 675B model hallucinates at roughly the same rate as a 7-8B model is notable because it means scaling alone will not fix this, which matters for anyone building or deploying MCP-based agents at enterprise scale. Practical impact is limited today: there is no vendor adoption, no named affected products, and the proposed resolver overlaps with checks some careful agent runtimes already implement.
What is noise
The "structural blind spot" framing overstates novelty since disciplined engineering teams already validate tool calls against a registry; this paper formalises and benchmarks that practice rather than inventing it. This is an unreviewed preprint with no named models, no evidence links captured, and no independent replication, so treat the specific hallucination counts as provisional rather than settled fact.
Watch next
- 01Whether the HTB benchmark gets adopted or cited by agent framework maintainers (LangChain, MCP spec authors, etc.) within the next few months
- 02Publication status of the paper (peer review, conference acceptance) and whether the 322/154 hallucination figures hold up under independent replication with named models
- 03Any MCP specification or major model provider (Anthropic, OpenAI, Google) response addressing merged-namespace hallucination risk directly
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- AI systems outperform expert humans in persuasive communication22 Jun 202681
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680