UK AISI publishes verified evaluation results and configuration data via EvalEval's Evaluation Cards, accompanying new inference-compute study
The UK AI Security Institute (AISI) published verified evaluation results, context, and configuration data through EvalEval's Evaluation Cards platform, using the Every Eval Ever (EEE) reporting schema. The release covers five benchmarks (HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, Terminal-Bench 2.0) across six frontier models (Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, GPT-5.4), plus two cyber evaluations (Cyber CTFs, The Last Ones) on a different, partially overlapping model set. The data accompanies AISI's paper "How Inference Compute Shapes Frontier LLM Evaluation."
Entities: UK AI Security Institute, EvalEval Coalition, Every Eval Ever, Evaluation Cards, Hugging Face, HealthBench
1 primary
What happened
The UK AI Security Institute, via EvalEval's Evaluation Cards platform on Hugging Face, published verified evaluation results and full configuration data for five benchmarks (HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, Terminal-Bench 2.0) run against six frontier models (Claude Opus 4, 4.5, 4.6 and GPT-5, 5.2, 5.4), plus two cyber evaluations on an overlapping model set. The release uses the "Every Eval Ever" (EEE) reporting schema and accompanies a separate AISI paper on how inference compute affects evaluation scores. This is a data and tooling release, not a new benchmark or a new model.
Why it matters
This targets a real and specific problem: evaluation results across the industry are typically reported in inconsistent formats without enough detail for anyone else to reproduce them, and re-running frontier evaluations is expensive. Researchers, model developers and anyone trying to sanity-check benchmark claims (including regulators and enterprises doing procurement due diligence) get a documented reference point and a shared schema to compare against. The impact is real but narrow: it improves the plumbing of evaluation science, it does not change what any model can do, what it costs, or how it is deployed.
What is noise
The framing that this "supports broader, more reliable meta-research across the evaluation ecosystem" is a forward-looking claim that depends entirely on other evaluators actually adopting the EEE schema, which has not happened yet at any scale. Treat "adoption" language as aspiration, not evidence of a trend. The release also only covers five benchmarks and a handful of models, so calling it a fix for evaluation reproducibility industry-wide overstates the current scope.
Watch next
- 01Whether any other lab or evaluator (e.g. Epoch AI, METR, a major AI company) publishes results in the EEE schema within the next 3-6 months, as opposed to just AISI
- 02Whether the accompanying AISI paper's findings on inference compute get cited or challenged by other evaluation researchers, which would validate or undercut the methodology
- 03Whether Evaluation Cards usage expands beyond this initial five-benchmark, six-model release, e.g. a public count of cards submitted by external teams
Coverage
1 storyMore capability signals
Full feed →- Deepseek releases V4.1-Flash, an open-source model that sharply cuts KV cache memory and input-processing compute for AI agents10 Sept 202682
- Anthropic threat report: Claude abused for malware, drone/missile software, mass surveillance, and industrial-scale distillation by Chinese AI labs11 Sept 202680
- WIRED investigation: Flock Safety's AI person-search tools let police run broad description-based surveillance, with weak guardrails against misuse3 Sept 202680
- Anthropic launches Claude Opus 5.5, matching Fable 5.1 performance at ~40% lower operating cost with faster output and less "Claudish" writing22 Sept 202679