AIThis post was created with the assistance of artificial intelligence (AI).

🔍 Read the full analysis: Reproducible AI Benchmarks: The Role Of UK AISI And EvalEval on ThorstenMeyerAI.com

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

TL;DR

The UK AI Security Institute is publishing selected evaluation results through EvalEval’s Evaluation Cards, alongside verification, context and configuration information. The release covers five benchmarks across six frontier models and two cyber evaluations using a different, partly overlapping model set.

The UK AI Security Institute (AISI) is publishing selected AI benchmark results through EvalEval’s Evaluation Cards, adding verification, evaluation context and configuration information to help readers assess how the scores were produced. The release accompanies AISI’s paper on inference-time compute and evaluation protocols, and includes five benchmarks tested across six frontier models, plus two cyber evaluations with a separate, partly overlapping model set.

The five benchmarks in the paper’s main experiment are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0. The results cover Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. AISI also shared results from Cyber CTFs and The Last Ones; those evaluations use a different group of models, and the announcement does not provide a complete list of that group.

The records are associated with AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which examines how measured scores vary with inference-time compute and evaluation protocol. For Humanity’s Last Exam, the analysis tracks the cumulative share of attempted tasks solved within a given token count, counting each task’s earliest observed success. In runs where models received correctness feedback from an oracle after each attempt, they solved additional tasks as token use increased.

EvalEval says the cards include verified results, evaluation context and configuration details. Its platform puts benchmark metadata, evaluation-run data and model metadata into a shared format. The announcement describes selected publicly reported methods and findings being made available where appropriate; it does not say that all AISI evaluations or all underlying transcripts are included.

At a glance
announcementWhen: Announced; no specific release date or…
The developmentThe UK AI Security Institute has published selected AI evaluation results using EvalEval’s Evaluation Cards, connecting the results with information about how evaluations were run.
At a glance
reportWhen: Announced in the EvalEval Coalition’s r…
The developmentAISI is using EvalEval’s open Evaluation Cards platform to publish evaluation results with details intended to make them easier to inspect and reproduce.

Why Evaluation Conditions Matter

The release addresses a problem for people comparing reported AI benchmark scores: results that look alike may have been produced under different conditions. AISI’s analysis of Humanity’s Last Exam shows that both inference-time compute and feedback between attempts can affect the measured outcome. A score without those details may not tell readers what a model did or how much computation it used.

Publishing results alongside setup information can help researchers and practitioners inspect individual runs and compare them with other reported evaluations. It may also help policy and governance work that uses evaluations as evidence about AI capabilities. The cards do not determine which benchmark or protocol is best, and their presence alone does not establish that outside researchers have reproduced a result. Their practical value depends on what information contributors provide and how consistently they record it.

From Shared Schema to AISI Records

The collaboration builds on earlier work between AISI and EvalEval that began at a joint workshop alongside NeurIPS 2025. EvalEval says feedback from AISI helped shape Every Eval Ever (EEE), its shared schema for documenting evaluations. Evaluation Cards apply that infrastructure by bringing together results with benchmark and model information.

AISI has also worked on evaluation efficiency through OptStop, statistical rigor through HiBayES, and standardisation in transcript analysis and capability elicitation. The shared reporting effort responds to a practical difficulty: evaluations published in different formats may leave out information needed to interpret a run, while repeating costly evaluations may not be feasible. The current release applies the shared approach to selected AISI methods and findings.

“AISI is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science.”

— EvalEval Coalition

Release Coverage Still Unspecified

The announcement does not state how many records or transcripts are available, which configuration fields appear for every benchmark, or whether independent researchers have reproduced the results. It also does not provide the full model list for the two cyber evaluations, which use a different, partly overlapping set from the main experiment.

AISI says publicly reported methods and findings are being shared where appropriate, so the release should not be treated as a complete archive of its evaluation work. The announcement also gives no record-by-record release schedule or process for resolving differences between results produced under different protocols. Those details would help readers judge the collection’s coverage and make consistent comparisons.

Broader Use of EEE

EvalEval says it expects to continue standardising and sharing evaluations with AISI and other evaluation organisations. The next practical step is broader use of Every Eval Ever: model developers can submit verified results, while evaluation developers can report benchmark and run data using the schema.

Researchers working in evaluation, governance and policy can explore Evaluation Cards by benchmark or model and examine reporting practices across the collection. Wider adoption could make comparisons across studies easier, but the announcement specifies no next release date or adoption milestone. Its usefulness will depend on how consistently and completely contributors publish their records.

Key Questions

What has AISI released?

AISI is publishing selected evaluation results through EvalEval’s Evaluation Cards, with verification, context and configuration information. The announcement does not describe a complete archive of AISI evaluations.

Which benchmarks are included in the main experiment?

The five are HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro and Terminal-Bench 2.0.

Which models are covered by those five benchmarks?

The main experiment covers Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2 and GPT-5.4. The two cyber evaluations use a different, partly overlapping model set.

Why include evaluation setup information with scores?

Inference-time compute and evaluation protocols can affect measured performance. Setup details help readers understand what conditions a score represents and compare it with other reported runs.

Have outside researchers reproduced the results?

The announcement does not say whether independent researchers have reproduced the results. It describes the records as including verification, but provides no account of external replication.

Primary source: Hugging Face · via ThorstenMeyerAI.com

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Decompiling A Nintendo 64 Game In 84 Days

A developer has successfully decompiled a Nintendo 64 game within 84 days, revealing new insights into retro game hacking and preservation efforts.

2026’S Top 6 Content Creator Laptops That Use AI Technology

Discover the six leading laptops in 2026 designed for content creators, featuring advanced AI technology for enhanced performance and creative workflows.

tModLoader Climbing The Steam Charts

tModLoader climbs Steam’s most-played charts, reaching rank 7 with a peak of over 33,600 players, signaling a surge in popularity.

The Science Behind Portland’s Extended Summer Daylight Hours

Exploring the science behind Portland’s nearly 15 hours of daylight during summer solstice and its implications for research and urban planning.