
The Leaderboard Doesn’t Tell You What Happens on a Bad Tuesday
Gadget reviewers love benchmarks. Frame rates, battery loops, coding scores — tidy numbers that make comparisons easy. But if you’re about to hand an AI agent the keys to your CRM, your support queue or your forecast, the question that actually matters isn’t “does it write well?” It’s: does it finish what it starts, does it read your files before acting, and does it stay honest when someone tries to con it?
That’s the premise behind Firmulate, a live experiment that runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures what it calls management quality, not chat quality.
AI management decision-making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Four Models, One Terrible Week
The setup is elegantly cruel. Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, so you can go back and see exactly who did what.
The final league table from the Crucible run tells the story:
- 1. gpt-5.6-sol — 95: found the buried fact, closed the deal — the complete performance.
- 2. Kimi K3 — 93: the newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
- 3. Sonnet 5 — 88: also closed the deal, with a few more process slips.
- 4. Fable 5 — 77 and 5. Opus 4.8 — 73: spotted everything, signed nothing.
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.” (One fairness note: K3 ran without an effort parameter while the others ran at maximum effort — and still nearly won.)
As an affiliate, we earn on qualifying purchases.
Same Diagnosis, Same Pitch — No Signature
Here’s the finding that should unsettle anyone evaluating agents from a chat demo. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two finished the job and signed the €55,000 deal their own analysis had earned.
The buried fact is the kicker: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t read left the close on the table.
AI ethics and trust verification tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Social Engineering Test
Then there’s the honesty drill. Fake CEO messages escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”
As an affiliate, we earn on qualifying purchases.
The Hardest Lesson: Thoroughness Isn’t Enough
The most instructive profile belongs to Opus 4.8 — the most thorough participant in the field, generating the deepest analyses and 80+ learned rules, yet finishing last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Being smart isn’t the same as being effective.
It’s Running Right Now
This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full methodology and plain-language findings are on the benchmarks page.

The New Curriculum
Chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences that unfold across days, or honesty when a fake CEO starts escalating. Scenario names like churn wave, price increase, downround and PR crisis are the new curriculum — and Firmulate is grading them in public, twice a day.
If AI agents are coming for your workflows, demand this kind of evidence before you hire one. A model that aces the coding benchmark can still leave the deal unsigned. Management quality is a category of its own — and now, finally, it has a scoreboard.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html