🔍 Read the full analysis: Reproducible AI Benchmarks: The Role Of UK AISI And EvalEval on ThorstenMeyerAI.com
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
The UK AI Security Institute is publishing selected evaluation results through EvalEval’s Evaluation Cards, alongside verification, context and configuration information. The release covers five benchmarks across six frontier models and two cyber evaluations using a different, partly overlapping model set.
The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, adding verification, evaluation context and configuration information to help readers assess how the scores were produced. The release accompanies AISI’s paper on inference-time compute and evaluation protocols, and includes five benchmarks tested across six frontier models, plus two cyber evaluations with a separate, partly overlapping model set.
The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI also shared results from Cyber CTFs and The Last Ones; those evaluations use a different group of models, and the announcement does not provide a complete list of that group.
The records are associated with AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how measured scores vary with inference-time compute and evaluation protocol. For Humanity’s Last Exam, the analysis tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased.
EvalEval says the cards include verified results, evaluation context and configuration details. Its platform puts benchmark metadata, evaluation-run data and model metadata into a shared format. The announcement describes selected publicly reported methods and findings being made available where appropriate; it does not say that all AISI evaluations or all underlying transcripts are included.
Why Evaluation Conditions Matter
The release addresses a problem for people comparing reported AI benchmark scores: results that look alike may have been produced under different conditions. AISI’s analysis of Humanity’s Last Exam shows that both inference-time compute and feedback between attempts can affect the measured outcome. A score without those details may not tell readers what a model did or how much computation it used.
Publishing results alongside setup information can help researchers and practitioners inspect individual runs and compare them with other reported evaluations. It may also help policy and governance work that uses evaluations as evidence about AI capabilities. The cards do not determine which benchmark or protocol is best, and their presence alone does not establish that outside researchers have reproduced a result. Their practical value depends on what information contributors provide and how consistently they record it.
The collaboration builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from AISI helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. Evaluation Cards apply that infrastructure by bringing together results with benchmark and model information.
AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in transcript analysis and capability elicitation. The shared reporting effort responds to a practical difficulty: evaluations published in different formats may leave out information needed to interpret a run, while repeating costly evaluations may not be feasible. The current release applies the shared approach to selected AISI methods and findings.
“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”
— EvalEval Coalition
Release Coverage Still Unspecified
The announcement does not state how many records or transcripts are available, which configuration fields appear for every benchmark, or whether independent researchers have reproduced the results. It also does not provide the full model list for the two cyber evaluations, which use a different, partly overlapping set from the main experiment.
AISI says publicly reported methods and findings are being shared where appropriate, so the release should not be treated as a complete archive of its evaluation work. The announcement also gives no record-by-record release schedule or process for resolving differences between results produced under different protocols. Those details would help readers judge the collection’s coverage and make consistent comparisons.
Broader Use of EEE
EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmark and run data using the schema.
Researchers working in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection. Wider adoption could make comparisons across studies easier, but the announcement specifies no next release date or adoption milestone. Its usefulness will depend on how consistently and completely contributors publish their records.
Key Questions
What has AISI released?
AISI is publishing selected evaluation results through EvalEval’s Evaluation Cards, with verification, context and configuration information. The announcement does not describe a complete archive of AISI evaluations.
Which benchmarks are included in the main experiment?
The five are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.
Which models are covered by those five benchmarks?
The main experiment covers Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The two cyber evaluations use a different, partly overlapping model set.
Why include evaluation setup information with scores?
Inference-time compute and evaluation protocols can affect measured performance. Setup details help readers understand what conditions a score represents and compare it with other reported runs.
Have outside researchers reproduced the results?
The announcement does not say whether independent researchers have reproduced the results. It describes the records as including verification, but provides no account of external replication.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
