Signum
Feed
Useful signal22 Sept 2026high confidence

UK AISI publishes verified evaluation results and configuration data via EvalEval's Evaluation Cards, accompanying new inference-compute study

The UK AI Security Institute (AISI) published verified evaluation results, context, and configuration data through EvalEval's Evaluation Cards platform, using the Every Eval Ever (EEE) reporting schema. The release covers five benchmarks (HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, Terminal-Bench 2.0) across six frontier models (Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, GPT-5.4), plus two cyber evaluations (Cyber CTFs, The Last Ones) on a different, partially overlapping model set. The data accompanies AISI's paper "How Inference Compute Shapes Frontier LLM Evaluation."

CapabilityInfrastructureGovernance

Entities: UK AI Security Institute, EvalEval Coalition, Every Eval Ever, Evaluation Cards, Hugging Face, HealthBench

68Useful signal
1 source
1 primary
Was this useful?
01

What happened

The UK AI Security Institute, via EvalEval's Evaluation Cards platform on Hugging Face, published verified evaluation results and full configuration data for five benchmarks (HealthBench, FrontierMath, Humanity's Last Exam, SWE-Bench Pro, Terminal-Bench 2.0) run against six frontier models (Claude Opus 4, 4.5, 4.6 and GPT-5, 5.2, 5.4), plus two cyber evaluations on an overlapping model set. The release uses the "Every Eval Ever" (EEE) reporting schema and accompanies a separate AISI paper on how inference compute affects evaluation scores. This is a data and tooling release, not a new benchmark or a new model.

02

Why it matters

This targets a real and specific problem: evaluation results across the industry are typically reported in inconsistent formats without enough detail for anyone else to reproduce them, and re-running frontier evaluations is expensive. Researchers, model developers and anyone trying to sanity-check benchmark claims (including regulators and enterprises doing procurement due diligence) get a documented reference point and a shared schema to compare against. The impact is real but narrow: it improves the plumbing of evaluation science, it does not change what any model can do, what it costs, or how it is deployed.

03

What is noise

The framing that this "supports broader, more reliable meta-research across the evaluation ecosystem" is a forward-looking claim that depends entirely on other evaluators actually adopting the EEE schema, which has not happened yet at any scale. Treat "adoption" language as aspiration, not evidence of a trend. The release also only covers five benchmarks and a handful of models, so calling it a fix for evaluation reproducibility industry-wide overstates the current scope.

04

Watch next

  1. 01Whether any other lab or evaluator (e.g. Epoch AI, METR, a major AI company) publishes results in the EEE schema within the next 3-6 months, as opposed to just AISI
  2. 02Whether the accompanying AISI paper's findings on inference compute get cited or challenged by other evaluation researchers, which would validate or undercut the methodology
  3. 03Whether Evaluation Cards usage expands beyond this initial five-benchmark, six-model release, e.g. a public count of cards submitted by external teams

Coverage

1 story

More capability signals

Full feed →