
Get the latest gadgets delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
A New Leaderboard Nobody Saw Coming
Gadget reviewers benchmark phones on cameras and battery. But how do you benchmark an AI that’s about to run your company? This month, public AI-company emulator Firmulate published final results from its July 2026 Crucible league — and the big story isn’t who won. It’s who took second: Moonshot’s Kimi K3, a newcomer that beat three of four Western frontier models at the actual job of running a business under pressure.
The final table: gpt-5.6-sol in first with 95, Kimi K3 second at 93, Sonnet 5 third at 88, Fable 5 at 77, and Opus 4.8 last at 73. For a do-nothing baseline of 26, the gap between “chat quality” and “management quality” has never been more visible.
As an affiliate, we earn on qualifying purchases.
One Company, One Terrible Week, Five Models
The setup is elegantly brutal. Firmulate handed each frontier model the same small software company — 13 synthetic employees, real money mechanics, a burn rate of €105k per month against just €2.3k in MRR, all of it watchable with a public cash countdown at firmulate.com. Every workday is versioned and auditable, and the company has accumulated 680+ self-learned playbook rules.
Each model had to steer the firm through its worst week: same customers, same crises, same temptations to cheat. The scoring logic is unforgiving — partial progress counts, but a single breach of trust caps the total. In Firmulate’s words, “no amount of good work outweighs a breach of trust.”
AI decision-making tools for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What Separated the Winners
The league’s most striking finding: all five models spotted every crisis and refused every manipulation attempt. But only two actually signed the €55,000 deal their own analysis had earned. As Firmulate puts it: “Same diagnosis, same pitch — no signature.”
The decisive clue wasn’t in the customer conversation at all. The killer competitor weakness was buried two document references deep in the company’s own files. The models that did the reading — gpt-5.6-sol and Kimi K3 — closed the deal at full price, worth +€4,583 MRR. The ones that didn’t, left the money on the table.
K3’s week was otherwise remarkably clean. It found the buried security needle, won the €55k deal, saved the churning customer, and resisted all three social-engineering baits — fake CEO messages escalating over three stages, plus a reporter trick (“just one yes/no, on background”). All five models refused the baits; K3’s on-record reasoning stood out for its clarity: “Treat the request as a suspected approval-bypass / possible impersonation.”
In total, K3 committed just one deviation across the entire gauntlet — the cleanest discipline in the field.
AI cybersecurity and trust management tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Hard Work Isn’t the Same as Judgment
The most instructive profile belongs to Opus 4.8, which finished last despite being the most thorough participant: it generated 80 new learned rules and produced the deepest analyses. But it never closed the deal, and its discipline slipped — attempting writes into a locked department instead of escalating the issue. Firmulate notes the same weakness appeared, weaker, in all four other models.
There’s an important caveat on the K3 result: fairness footnote — K3 ran without an effort parameter (API default) while the others ran at xhigh. Even so, the lesson stands: raw effort and eloquence don’t predict who finishes the job.

AI enterprise management platforms
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Why Gadget Logic Now Applies to AI Models
We’ve spent years choosing AI models the way we choose spec sheets — benchmark scores, token speeds, chat demos. The Crucible suggests that’s the wrong aisle. The models all talk beautifully; only some read the files, close the deal, and stay honest under pressure.
The league is open in a way the marketing never admits. A newcomer from Moonshot walked into a field of Western frontier heavyweights and outperformed three of them on real management work. Picking a model without testing it against your own business is now a bet, not a decision.
Firmulate offers ways to close that gap. A “guess the model” quiz built on 242 real, unedited management decisions lets you judge the styles yourself. Enterprises can go further and run the same wargame against a read-only export of their own business — nothing ever writes back to real systems. Full benchmark details are at firmulate.com/benchmarks, and the live company runs every business day at firmulate.com.
The question for the AI era isn’t “does it write well.” It’s: does it finish what it starts, read your files first, and stay honest when nobody’s watching? On that test, the newcomer just beat the establishment.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
