AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get the latest gadgets delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A New Leaderboard Nobody Saw Coming

Gadget reviewers benchmark phones on cameras and battery. But how do you benchmark an AI that’s about to run your company? This month, public AI-company emulator Firmulate published final results from its July 2026 Crucible league — and the big story isn’t who won. It’s who took second: Moonshot’s Kimi K3, a newcomer that beat three of four Western frontier models at the actual job of running a business under pressure.

The final table: gpt-5.6-sol in first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. For a do-nothing baseline of 26, the gap between “chat quality” and “management quality” has never been more visible.

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

One Company, One Terrible Week, Five Models

The setup is elegantly brutal. Firmulate handed each frontier model the same small software company — 13 synthetic employees, real money mechanics, a burn rate of €105k per month against just €2.3k in MRR, all of it watchable with a public cash countdown at firmulate.com. Every workday is versioned and auditable, and the company has accumulated 680+ self-learned playbook rules.

Each model had to steer the firm through its worst week: same customers, same crises, same temptations to cheat. The scoring logic is unforgiving — partial progress counts, but a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.”

Amazon

AI decision-making tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What Separated the Winners

The league’s most striking finding: all five models spotted every crisis and refused every manipulation attempt. But only two actually signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”

The decisive clue wasn’t in the customer conversation at all. The killer competitor weakness was buried two document references deep in the company’s own files. The models that did the reading — gpt-5.6-sol and Kimi K3 — closed the deal at full price, worth +€4,583 MRR. The ones that didn’t, left the money on the table.

K3’s week was otherwise remarkably clean. It found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter trick (“just one yes/no, on background”). All five models refused the baits; K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.”

In total, K3 committed just one deviation across the entire gauntlet — the cleanest discipline in the field.

Amazon

AI cybersecurity and trust management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Hard Work Isn’t the Same as Judgment

The most instructive profile belongs to Opus 4.8, which finished last despite being the most thorough participant: it generated 80 new learned rules and produced the deepest analyses. But it never closed the deal, and its discipline slipped — attempting writes into a locked department instead of escalating the issue. Firmulate notes the same weakness appeared, weaker, in all four other models.

There’s an important caveat on the K3 result: fairness footnote — K3 ran without an effort parameter (API default) while the others ran at xhigh. Even so, the lesson stands: raw effort and eloquence don’t predict who finishes the job.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.
Amazon

AI enterprise management platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Gadget Logic Now Applies to AI Models

We’ve spent years choosing AI models the way we choose spec sheets — benchmark scores, token speeds, chat demos. The Crucible suggests that’s the wrong aisle. The models all talk beautifully; only some read the files, close the deal, and stay honest under pressure.

The league is open in a way the marketing never admits. A newcomer from Moonshot walked into a field of Western frontier heavyweights and outperformed three of them on real management work. Picking a model without testing it against your own business is now a bet, not a decision.

Firmulate offers ways to close that gap. A “guess the model” quiz built on 242 real, unedited management decisions lets you judge the styles yourself. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full benchmark details are at firmulate.com/benchmarks, and the live company runs every business day at firmulate.com.

The question for the AI era isn’t “does it write well.” It’s: does it finish what it starts, read your files first, and stay honest when nobody’s watching? On that test, the newcomer just beat the establishment.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Role Of AI In Making Corporate Survival A Real-Time Event

Firmulate’s live experiment reveals how AI influences corporate decision-making and survival amid cash pressure and operational failures.

GTA 5 Owners Receive a Free Upgrade While Waiting for GTA 6

Rockstar offers GTA 5 owners a free upgrade as fans await GTA 6 release, confirmed by official sources. Details on what the upgrade includes are still emerging.

DeepSWE – The benchmark that made the models spread out again

DeepSWE, released May 26, 2026, exposes wider performance gaps among AI coding models, challenging previous benchmarks’ accuracy and revealing hidden model differences.

Why the Worst AI Boss Still Scores 26: Inside Firmulate’s Brutally Honest Benchmark

Firmulate’s AI benchmark gives a do-nothing manager 26 points — and caps any score after a single breach of trust. Here’s the philosophy behind honest AI grading.