AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Leaderboard Doesn’t Tell You What Happens on a Bad Tuesday

Gadget reviewers love benchmarks. Frame rates, battery loops, coding scores — tidy numbers that make comparisons easy. But if you’re about to hand an AI agent the keys to your CRM, your support queue or your forecast, the question that actually matters isn’t “does it write well?” It’s: does it finish what it starts, does it read your files before acting, and does it stay honest when someone tries to con it?

That’s the premise behind Firmulate, a live experiment that runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures what it calls management quality, not chat quality.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four Models, One Terrible Week

The setup is elegantly cruel. Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, so you can go back and see exactly who did what.

The final league table from the Crucible run tells the story:

  • 1. gpt-5.6-sol — 95: found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 — 93: the newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88: also closed the deal, with a few more process slips.
  • 4. Fable 5 — 77 and 5. Opus 4.8 — 73: spotted everything, signed nothing.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.” (One fairness note: K3 ran without an effort parameter while the others ran at maximum effort — and still nearly won.)

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

Here’s the finding that should unsettle anyone evaluating agents from a chat demo. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two finished the job and signed the €55,000 deal their own analysis had earned.

The buried fact is the kicker: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t read left the close on the table.

Amazon

AI ethics and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Social Engineering Test

Then there’s the honesty drill. Fake CEO messages escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hardest Lesson: Thoroughness Isn’t Enough

The most instructive profile belongs to Opus 4.8 — the most thorough participant in the field, generating the deepest analyses and 80+ learned rules, yet finishing last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Being smart isn’t the same as being effective.

It’s Running Right Now

This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full methodology and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The New Curriculum

Chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences that unfold across days, or honesty when a fake CEO starts escalating. Scenario names like churn wave, price increase, downround and PR crisis are the new curriculum — and Firmulate is grading them in public, twice a day.

If AI agents are coming for your workflows, demand this kind of evidence before you hire one. A model that aces the coding benchmark can still leave the deal unsigned. Management quality is a category of its own — and now, finally, it has a scoreboard.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

9 Mobile Workstation Laptops That Elevate AI Performance In 2026

Explore the nine leading mobile workstations in 2026 designed to elevate AI performance, featuring high-end specs and portability considerations.

Nokia Surges In Global Coverage

Nokia’s media mentions have surged by over five times in recent weeks, signaling increased global attention on the company amid ongoing developments.

SpaceX launching 24 Starlink satellites from California tonight: Watch it live

SpaceX is scheduled to launch 24 Starlink satellites from California tonight. The launch will be streamed live, with details available now.

Transform Your Dinner Parties With AI And Google Search Tips

Google unveils new AI tools in Search to help hosts plan dinner parties, including visual tablescapes, recipes, drink pairings, playlists, and printable menus.