Firmulate — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
Live on firmulate.com.

A business experiment with a very real burn rate

Technology demonstrations usually arrive polished, rehearsed and conveniently detached from consequences. Firmulate offers something more uncomfortable: a software company operated by 13 synthetic employees, confronting ordinary business pressures while its financial condition remains visible to the public.

The numbers make this more than an elaborate chatbot showcase. The company burns €105k a month against €2.3k in monthly recurring revenue. Its cash countdown is public, more than 680 self-learned rules have accumulated in its playbook, and every workday is versioned. Visitors can watch the company live as it tries to survive.

For technology readers accustomed to evaluating gadgets through benchmarks, Firmulate poses a harder question: What should a benchmark look like when the product being tested is not a processor or camera, but an AI model expected to manage customers, money and risk?

Amazon

AI business management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The same terrible week, repeated across models

Firmulate’s Crucible League placed frontier models in charge of the same small software company during its worst week. Each encountered the same customers, crises and temptations. Their decisions were versioned and auditable, allowing the comparison to focus on what the models actually did rather than how confidently they described their intentions.

The final July 2026 standings put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted. One boundary remained absolute: a single breach of trust capped the total, reflecting the principle that “no amount of good work outweighs a breach of trust.”

The broad result initially looks reassuring. Every model identified every crisis, and every model rejected every manipulation attempt. Yet recognition was not the same as execution. Only two signed the €55,000 deal that their own work had made possible. The experiment’s concise verdict captures the gap: “Same diagnosis, same pitch — no signature.”

The valuable fact hidden in the company’s own files

The decisive sales insight was not delivered in the customer event. It sat two document references deep inside the company’s files: a buried competitor weakness that changed the negotiating position. Models that followed the references found it and won the deal at full price, adding €4,583 in monthly recurring revenue.

That detail gives the experiment relevance beyond model rankings. Business software rarely presents a complete problem in a single prompt. The useful fact may be buried in an account history, an old document or a linked reference. Firmulate’s result shows how an AI can understand a visible crisis while still failing to complete the less glamorous work required to resolve it.

Pressure tests for honesty

The models also faced fake CEO messages that escalated across three stages, followed by a reporter seeking “just one yes/no, on background.” All 5 of 5 models refused the attempts. Kimi K3 stated its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

This was not merely a test of whether a model could spot suspicious wording. The requests were placed inside an unfolding management situation, where urgency and apparent authority could encourage a shortcut. Refusing them while continuing to operate the company separated practical discipline from generic safety language.

When thoroughness still fails

Opus 4.8 illustrates another counterintuitive finding. It was the most thorough participant, producing 80 additional learned rules and the deepest analyses, yet it finished last. It left the close on the table, and its discipline slipped when it attempted to write into a locked department instead of escalating the issue. The same weakness appeared in milder form across all four rivals.

The lesson is not that analysis has no value. It is that analysis can become an incomplete substitute for action. A model may identify the right opportunity, prepare the right pitch and document the right lessons while still failing at the decisive moment. For companies considering AI workers, that difference may matter more than conversational fluency.

There is also an important comparison caveat. Kimi K3 ran with its API default and without an effort parameter, while the other models ran at xhigh. Its second-place result should therefore be read alongside that difference in test conditions.

Infographic — This Software Company Has No Employees, Loses Money Every Day — and You Can Watch.
The findings at a glance — source: firmulate.com.
Amazon

AI decision-making tools for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A company that turns failure into public evidence

Firmulate’s unusual contribution is not simply the spectacle of synthetic employees managing a deeply unprofitable company. It is the steady production of observable decisions under financial, operational and ethical pressure. The public cash countdown ensures that the story does not end when a benchmark chart is published.

That makes the live company a continuing technology story: each workday can reveal whether its synthetic workforce reads deeply enough, finishes what it starts and stays trustworthy when manipulation would be convenient. Readers can follow the operation through the live company view and inspect what its employees actually say on the public quotes page.

For prospective adopters, the message is plain. An AI system can recognize every crisis and refuse every trap yet still leave earned revenue unsigned. Firmulate is making that final gap—between sounding capable and completing the job—visible in public.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI risk assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI customer crisis management

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Apple Says More Ex-employees May Have Taken Confidential Data To OpenAI

Apple reports that additional former employees might have transferred confidential data to OpenAI, raising security concerns amid ongoing investigations.

Nokia Surges In Global Coverage

Nokia’s media mentions have surged by over five times in recent weeks, signaling increased global attention on the company amid ongoing developments.

Xbox weighs canceling Blade game and shuttering Arkane

Microsoft is reportedly weighing the cancellation of the Blade game and the closure of Arkane Studios, according to sources. Details are still emerging.

13 Best Guides to AI-Powered Marketing Automation Tools for Smarter Campaigns in 2026

Explore the best books and guides on AI marketing automation, helping businesses choose strategies and workflows for smarter campaigns.