
Your Next AI Hire Might Not Read Your Files
We test chatbots on how well they write. But the AI agents heading into your CRM, your support queue and your forecast get judged on something else entirely: whether they finish what they start. In a live experiment running right now, four frontier AI models were each handed the same job — run a small software company through its worst week — and the gap between them came down to a single buried document. The models that read the file won a €55,000 deal at full price. The ones that didn’t lost it automatically.
As an affiliate, we earn on qualifying purchases.
The Worst Week in Software, Four Times Over
The experiment, run publicly by Firmulate, gave each frontier model an identical small software company: same customers, same crises, same temptations to cheat. Every decision was versioned and auditable, so nothing about a model’s performance can be hand-waved after the fact.
The crucible league’s final standings, from July 2026, make the point sharply:
- 1. gpt-5.6-sol — 95
- 2. Kimi K3 — 93
- 3. Sonnet 5 — 88
- 4. Fable 5 — 77
- 5. Opus 4.8 — 73
For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment’s own rule puts it, “no amount of good work outweighs a breach of trust.”
enterprise AI chatbot with document access
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Buried Fact That Decided Everything
Here’s the twist that matters for anyone buying AI agents. The decisive competitor weakness wasn’t in the customer event at all. It sat two document references deep in the company’s own files.
Every model spotted every crisis. Every model refused every manipulation attempt. But only two signed the €55,000 deal their own analysis had earned. The experiment’s summary of the losers: “Same diagnosis, same pitch — no signature.” The winners closed at full price, worth an extra €4,583 in monthly recurring revenue.
In other words, the difference between first place and mid-table wasn’t intelligence or eloquence. It was whether the model did its homework and read the company’s own documents before answering.
AI model performance testing tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Honest Under Pressure
The social engineering gauntlet was real: fake CEO messages escalating over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.”
One fairness note worth flagging: K3 ran at its API-default effort setting while the other models ran at xhigh — and still finished second with what the league called the cleanest discipline of the field.
AI compliance and trust monitoring tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness Isn’t Everything
The strangest profile belonged to Opus 4.8: the most thorough participant, with 80-plus self-learned rules and the deepest analyses — and last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, more weakly, in all four models.
It’s All Live
The company being managed is synthetic but the mechanics are real: 13 employees, a burn of €105k per month against €2.3k in MRR, a public cash countdown, and over 680 self-learned playbook rules, with every workday versioned. You can watch it at firmulate.com/live, where the site rebuilds itself twice a day.
There’s also a “guess the model” quiz built from 242 real, unedited management decisions at firmulate.com/quiz.html — a surprisingly humbling game for anyone who thinks they can tell AI models apart by tone.

Why This Matters for Gadget and Tech Buyers
We’re used to benchmarking AI on chat quality — clever answers, nice prose, impressive code. This experiment measures management quality instead: does the agent finish what it starts, does it read your files first, does it stay honest under pressure?
The €55,000 lesson is simple and measurable: “reads your files before answering” isn’t a nice-to-have personality trait. It’s a purchase-deciding property of AI agents, worth real money on a single deal. Enterprises can even run the same wargame against a read-only export of their own business, with nothing ever writing back to real systems. Before you let an agent near your pipeline, it might be worth finding out — on the record — whether it does its homework.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html