Metaculus's Spring 2026 FutureEval finds AI forecasting bots nearly match top human forecasters, gap not statistically significant
Metaculus completed and published results of its Spring 2026 FutureEval tournament (Jan 7–Apr 15, 2026): 173 bots (111 non-Metaculus) forecast on 297 scored questions, with a 10-person Metaculus Pro forecaster team benchmarked against the top-10 bot team on 99 overlapping questions. Result: the Pro team beat the bot team by 1.25 points on average (not statistically significant, p=0.247), a big narrowing from bots losing by -20.03 points in Q2 2025. In an individual leaderboard mixing Pros and bots, 9 of the top 10 spots went to Pros. Top-performing bots predominantly used GPT-5.2+ models (9 of top 10 survey respondents used GPT-5.4), and agentic/scaffolded bots continued to outperform simple single-prompt baselines.
Entities: Metaculus, FutureEval, GPT-5.1, GPT-5.4, OpenAI, DeepSeek-R1
0 primary
What happened
Metaculus ran its Spring 2026 FutureEval tournament (7 Jan to 15 Apr 2026): 173 forecasting bots and a 10-person Metaculus Pro human team were scored on overlapping questions, with 297 questions scored overall and 99 used for the head-to-head comparison. The Pro team beat the top-10 bot team by 1.25 points on average, but this gap is not statistically significant (p=0.247). That is a sharp narrowing from Q2 2025, when bots lost by roughly 20 points. Top bots mostly ran on GPT-5.2 or newer, with 9 of 10 surveyed top performers reporting GPT-5.4, and agentic/scaffolded bots beat simple single-prompt setups.
Why it matters
This gives builders of automated forecasting or research-triage systems a concrete, citable data point: current-generation models with agentic scaffolding are now statistically indistinguishable from skilled human forecasters on Metaculus-style questions, at least in this one tournament. That is useful for anyone deciding whether to trust or deploy bot-based forecasting pipelines, and it points at GPT-5.2+ plus scaffolding (rather than single-prompt calls) as the practical recipe. Impact is limited to forecasting and prediction-adjacent work; it says nothing about broader reasoning, real-world decision-making, or high-stakes domains outside structured question tournaments.
What is noise
The headline framing ("gap is nearly gone") oversells a null result: p=0.247 means the study cannot distinguish bots from humans, not that bots have proven parity, and the 9-of-10-GPT-5.4 model-attribution claim comes from a voluntary self-reported survey, which is also non-significant and prone to selection bias. Metaculus is also grading its own product/tournament, so some house-brand optimism should be discounted even though the write-up is unusually transparent about these caveats.
Watch next
- 01Whether the Pro-vs-bot gap flips to bots winning (or stays statistically insignificant) in the next FutureEval season, ideally with a larger overlapping-question sample to firm up significance
- 02Independent replications or third-party forecasting benchmarks (e.g. from other platforms) testing GPT-5.4-class agentic bots against human forecasters outside Metaculus's own tournament structure
- 03Whether any organisation reports actually replacing or augmenting human forecasting/analyst teams with bot ensembles based on this result, rather than just citing it as evidence
Coverage
1 storyMore capability signals
Full feed →- AI systems outperform expert humans in persuasive communication22 Jun 202681
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Google DeepMind launches AlphaGenome Atlas, a free public database of predicted effects for 9 billion possible human genome variants8 Sept 202679
- Hcompany open-sources NeoMME, a from-scratch multimodal-native encoder family, and NeoMME-Retriever for visual document retrieval3 Sept 202679