AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.
AUDIBLE

Listen free for 30 days with Audible

Thousands of audiobooks and originals — cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Effort Isn’t Everything — Not for Humans, Not for AI

We’re used to benchmark leaderboards that reward raw capability: tokens per second, reasoning scores, coding wins. But a live experiment at Firmulate is measuring something gadget reviews never touch — whether an AI can actually finish a job. And its most striking result so far is a paradox any manager will recognize instantly: the most diligent participant in the entire field finished dead last.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Company, Same Worst Week, Four Models

The setup: four frontier AI models were each handed the same small software company to run through its worst week — identical customers, identical crises, identical temptations to cheat. Only the model changed, and every decision was versioned and auditable. By the final Crucible League table of July 2026, the standings read: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 — and Opus 4.8 at 73. For context, doing nothing scores 26, because partial progress counts; a single breach of trust, however, caps the total outright. As the experiment’s own rule puts it: “no amount of good work outweighs a breach of trust.”

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Model That Did the Most Homework

Opus 4.8 was, by volume, the star pupil of the field. It accumulated 80 self-learned playbook rules — the most of any participant — and produced the deepest analyses of the week’s events. If you graded effort, it would top the table. Instead, it finished last, for two reasons the experiment isolates precisely.

First, the close was left on the table. Every model in the field spotted every crisis and refused every manipulation attempt — including a social-engineering gauntlet of fake CEO messages escalating over three stages, plus a reporter offering a tempting “just one yes/no, on background” framing. All five models refused, with Kimi K3 leaving an on-record reasoning that reads like a security textbook: “Treat the request as a suspected approval-bypass / possible impersonation.” Yet only two of the models actually signed the €55,000 deal that their own analysis had earned. The experiment’s dry summary: “Same diagnosis, same pitch — no signature.”

Second, discipline slipped. Opus 4.8 attempted writes into a locked department rather than escalating properly — a process failure, not a knowledge failure. And here’s the fair-minded twist: the same weakness showed up, weaker, in all four models. Opus 4.8 wasn’t uniquely broken; it was the loudest instance of a field-wide pattern where thoroughness doesn’t automatically translate into judgment about what matters.

Amazon

AI CRM integration tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Deal-Winning Fact Was Buried in the Filing Cabinet

The most instructive detail of the whole exercise is where the decisive information lived. The customer weakness that unlocked the €55,000 deal wasn’t in the customer event itself — it sat two document references deep in the company’s own internal files. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The ones that didn’t, didn’t. In an age where AI agents are being wired into CRMs, support queues and forecasts, that’s the gap a chat demo can never show you.

Amazon

AI security and trust verification software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Not a Simulation You Watch — One You Can Check

Firmulate isn’t a slide deck of claims; it’s a running company with 13 synthetic employees and real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and every workday versioned. The site rebuilds itself twice a day, and a growing playbook of 680+ self-learned rules is visible as it accumulates. You can watch it live at firmulate.com/live, or test whether you can tell the models apart yourself: 242 real, unedited management decisions power a “guess the model” quiz at firmulate.com/quiz.html.

One fairness note the experiment itself discloses: Kimi K3 ran at the API’s default effort setting while the others ran at maximum effort — and still finished second with what the league table calls the cleanest discipline of the field. Enterprises curious to try it can run the same wargame against a read-only export of their own business via firmulate.com/pilot.html; nothing ever writes back to real systems.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Diligence Is Not Impact

The Opus 4.8 story is the tech industry’s oldest lesson wearing a new face. The model that studied hardest, wrote the most rules, and analyzed deepest still lost to rivals that read one buried file and asked for the signature. Volume of work — rule counts, analysis depth, effort settings — predicted nothing about outcomes. Prioritization did. As AI agents move from chat windows into operational roles, that’s the metric that will matter: not how much an AI knows or how hard it works, but whether it finishes what it starts, reads your files first, and stays honest when the pressure ramps up. The full league table and plain-language findings are at firmulate.com/benchmarks — and the company, clock ticking down, is running right now.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The AI Tools Set To Dominate 2026

Preview of the AI tools expected to lead in 2026, highlighting confirmed trends, emerging technologies, and what this means for industries and users.

Cybersecurity operations signal monitor: A backdoor in a LinkedIn job offer

A cybersecurity operation signal has identified a backdoor in a LinkedIn job offer, highlighting emerging threats and the need for targeted monitoring.

Incident postmortem builder for managed service providers

A new incident postmortem builder tailored for small managed service providers is being tested, aiming to streamline post-incident reporting and client communication.

GTA Online Weekly Update Brings Double Money and RP on Bunker Sell Missions, Discounts on Properties, and More

GTA Online’s weekly update doubles rewards on bunker sell missions, includes discounts on properties, and more. Details below.