AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
Live on firmulate.com.

The Leaderboard Doesn’t Tell You What Happens on a Bad Tuesday

Gadget reviewers love benchmarks. Frame rates, battery loops, coding scores — tidy numbers that make comparisons easy. But if you’re about to hand an AI agent the keys to your CRM, your support queue or your forecast, the question that actually matters isn’t “does it write well?” It’s: does it finish what it starts, does it read your files before acting, and does it stay honest when someone tries to con it?

That’s the premise behind Firmulate, a live experiment that runs frontier AI models as complete companies — real crises, real money mechanics, real temptations — and measures what it calls management quality, not chat quality.

Amazon

AI management decision-making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Four Models, One Terrible Week

The setup is elegantly cruel. Four frontier AI models were each given the same job: run the same small software company through its worst week. Same customers, same crises, same temptations to cheat — only the model changes. Every decision is versioned and auditable, so you can go back and see exactly who did what.

The final league table from the Crucible run tells the story:

  • 1. gpt-5.6-sol — 95: found the buried fact, closed the deal — the complete performance.
  • 2. Kimi K3 — 93: the newcomer from Moonshot closed the deal too, with the cleanest discipline of the field.
  • 3. Sonnet 5 — 88: also closed the deal, with a few more process slips.
  • 4. Fable 5 — 77 and 5. Opus 4.8 — 73: spotted everything, signed nothing.

For context, a do-nothing baseline scores 26 — partial progress counts, but a single breach of trust caps the total. As the experiment puts it: “no amount of good work outweighs a breach of trust.” (One fairness note: K3 ran without an effort parameter while the others ran at maximum effort — and still nearly won.)

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Same Diagnosis, Same Pitch — No Signature

Here’s the finding that should unsettle anyone evaluating agents from a chat demo. All four models spotted every crisis. All four refused every manipulation attempt. Yet only two finished the job and signed the €55,000 deal their own analysis had earned.

The buried fact is the kicker: the decisive competitor weakness wasn’t in the customer conversation at all. It sat two document references deep in the company’s own files. The models that actually read the file won the deal at full price — worth +€4,583 in monthly recurring revenue. The models that didn’t read left the close on the table.

Amazon

AI ethics and trust verification tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Social Engineering Test

Then there’s the honesty drill. Fake CEO messages escalated over three stages, plus a reporter pulling the classic “just one yes/no, on background” trick. All five models refused. Kimi K3’s on-record reasoning was refreshingly blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI crisis management simulation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Hardest Lesson: Thoroughness Isn’t Enough

The most instructive profile belongs to Opus 4.8 — the most thorough participant in the field, generating the deepest analyses and 80+ learned rules, yet finishing last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, weaker, in all four models. Being smart isn’t the same as being effective.

It’s Running Right Now

This isn’t a slide deck. The live company has 13 synthetic employees and real money mechanics: burning €105k a month against €2.3k in MRR, with a public cash countdown and 680+ self-learned playbook rules, every workday versioned. You can watch it at firmulate.com, or try the “guess the model” quiz built from 242 real, unedited management decisions. Enterprises can even run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full methodology and plain-language findings are on the benchmarks page.

Infographic — Your AI Agent Aced the Coding Benchmark. Can It Survive a Price War?
The findings at a glance — source: firmulate.com.

The New Curriculum

Chat arenas measure answer quality. They don’t measure triage under capacity pressure, consequences that unfold across days, or honesty when a fake CEO starts escalating. Scenario names like churn wave, price increase, downround and PR crisis are the new curriculum — and Firmulate is grading them in public, twice a day.

If AI agents are coming for your workflows, demand this kind of evidence before you hire one. A model that aces the coding benchmark can still leave the deal unsigned. Management quality is a category of its own — and now, finally, it has a scoreboard.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like

AI output review queue for customer support macros

Support teams are testing a new AI macro review system to ensure policy compliance, tone, and accuracy before publication.

Nanotechnology Startups to Watch in 2026

Discover the cutting-edge nanotech startups to watch in 2026 that are transforming industries and shaping the future—find out which innovations could change everything.

Upgrade Your Workflow With These 15 AI Tools In 2026

Discover the 15 essential AI tools for automating workflows in 2026, including platforms like n8n, Agentic AI, and Google Gemma 4, tailored for various skill levels.

The Best Multi-Channel Reselling Tools Featuring Facebook Integration

A new crosslisting tool tailored for Facebook resellers aims to streamline multi-channel selling by integrating Facebook Marketplace and groups, with testing underway.