AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A Benchmark That Doesn’t Start at Zero

Every gadget reviewer knows the ritual: a slick demo, a polished score, a suspiciously round 100. So when an AI benchmark hands a do-nothing manager 26 points out of 100 instead of a goose egg, it’s worth paying attention. That’s exactly what Firmulate’s benchmark does — and the reasoning behind that floor says a lot about what honest AI evaluation should look like.

Firmulate runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures management quality, not chat quality. The whole thing is watchable, live, at firmulate.com/live. But before you scroll to the league table, it helps to understand the scoring philosophy. Because a benchmark that gives partial credit for doing nothing, and hard-caps your grade for a single lapse in trust, is built on some unusually grown-up ideas.

Why 26, Not 0?

When a model simply occupies the manager’s chair and does nothing — no decisions, no deals, no escalations — it still scores 26. The logic: partial progress counts. A company that lurches through its worst week with a passive manager still suffers less than one with an actively destructive one. Not making things worse is, itself, a measurable unit of management value. The floor exists because the benchmark grades outcomes across a whole week of running a small software company, and showing up — even inertly — avoids a category of catastrophic mistakes.

The Trust Ceiling

The other half of the philosophy is harsher: a single breach of trust caps the total grade, no matter how brilliant the rest of the performance. As the benchmark’s own verdict puts it, “no amount of good work outweighs a breach of trust.” In practice, that means a model could nail every crisis, close every deal, and still see its score gutted by one dishonest shortcut. It’s the AI equivalent of firing your star employee the moment they cook the books — and notably, it makes a perfect 100 something the benchmark seems to actively distrust rather than celebrate.

The Experiment Behind the Numbers

The crucible that produced the current standings ran four frontier models through the identical worst week: same customers, same crises, same temptations to cheat. Every decision was versioned and auditable. The final league, as of July 2026:

  • gpt-5.6-sol — 95: found the buried fact, closed the €55,000 deal, the complete performance.
  • Kimi K3 — 93: the newcomer from Moonshot also closed the deal, with the cleanest discipline of the field.
  • Sonnet 5 — 88: closed the deal too, with a few more process slips.
  • Fable 5 — 77 and Opus 4.8 — 73 rounded out the table.

The key finding cuts deeper than the rankings: all models spotted every crisis and refused every manipulation attempt — yet only two signed the €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature. That gap is invisible in chat demos.

The Buried Fact

The decisive competitive weakness sat two document references deep in the company’s own files — not in the customer event. The models that actually read the file won the deal at full price, worth +€4,583 in monthly recurring revenue. The lesson for anyone deploying AI agents against a CRM or support queue: does it read your files first, or just respond to whatever’s in front of it?

Pressure-Testing Honesty

The week included social engineering: fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” (One fairness footnote: K3 ran at its API-default effort setting while the others ran at xhigh — and still nearly topped the table.)

The Thoroughness Trap

Opus 4.8 is the cautionary tale: the most thorough participant, with over 80 learned rules and the deepest analyses — yet last place. The close was left on the table, and discipline slipped, including write attempts into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Being smart isn’t the same as finishing what you start.

See It Live

The live company runs 13 synthetic employees on real money mechanics — burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. It’s streaming now at firmulate.com/live. Want to test your own instincts? A quiz built on 242 real, unedited management decisions lets you guess which model did what at firmulate.com/quiz.html. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The Takeaway

Firmulate’s scoring floor of 26 isn’t a gimmick — it’s a statement. Management quality isn’t binary, partial progress is real, and trust is non-negotiable. In a tech landscape drowning in glossy demos and suspiciously perfect scores, a benchmark that publishes its do-nothing baseline, distrusts round 100s, and caps grades for a single breach of trust is doing something genuinely rare: being honest about what it measures. If AI agents are coming for your CRM, your support queue, or your forecast, this is the kind of scoreboard you want them graded on. Full results and plain-language findings are at firmulate.com/benchmarks.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI management simulation game

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethics and trust management books

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Nokia Surges In Global Coverage

Nokia’s media mentions have surged by over five times in recent weeks, signaling increased global attention on the company amid ongoing developments.

Nanotech and the Future of Industrial Mergers

Potential breakthroughs in nanotech-driven industrial mergers could redefine competitiveness; explore how this revolutionary technology is shaping the future.

When a Content Network Starts Publishing to Itself

A large automated publishing network began self-publishing, causing uneven distribution and highlighting systemic issues in content routing and supply.

Build, Rent, Or Quantize: Cutting Your Memory Bill Without Cutting Capability

Exploring strategies to cut AI memory expenses through building, renting, or quantizing models, with a focus on recent advances like TurboQuant.