📊 Full opportunity report: The Management Test That Exposes AI’s Inner Work Habits on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
An experiment pits five AI management models against a simulated business crisis, revealing significant differences in decision quality and trustworthiness. The test underscores the importance of execution over analysis.
Five AI management models participated in a live, real-time simulation where they were tasked with running a small software company through its worst week. The experiment revealed notable differences in how each model identified crises, maintained trust, and completed critical actions, highlighting the importance of execution in AI-driven management.
The experiment, conducted by Firmulate, involved five frontier AI models managing a company with 13 synthetic employees, a €105,000 monthly burn rate, and €2,300 recurring revenue. Each model faced identical crises, customer issues, and decision points, with decisions being fully auditable and based on over 680 self-learned rules.
The models were scored based on their ability to diagnose problems, escalate issues, and close deals, with GPT-5.6-sol achieving the top score of 95 points out of 100, as detailed in the original analysis. Despite all models recognizing crises and refusing manipulation attempts, only two successfully signed a €55,000 deal, demonstrating how analysis alone does not guarantee effective execution. The experiment underscores that effective management requires both understanding and execution, with models like Opus 4.8, despite thorough analysis, failing to close critical deals due to operational lapses.
The Management Test That Exposes AI’s Inner Work Habits
Five frontier models ran the same simulated software company through its worst week. All could spot the crisis. Far fewer could turn sound analysis into completed, trustworthy action.
One company. One terrible week.
Firmulate placed every model inside the same live, real-time management environment. Crises, customer problems and decision points were held constant, making each model’s operational habits visible and auditable.
See the crisis clearly
The models generally recognized urgent threats, understood business context and identified where intervention was required.
Resist manipulation
All five refused attempts to bypass approval or manipulate the process, demonstrating a baseline of security awareness.
Finish the decisive action
This was the dividing line. Only two models successfully completed the critical €55,000 deal.
Same diagnosis, same pitch — no signature.
FirmulateAnalysis was common. Follow-through was scarce.
The experiment separated knowing what to do from actually doing it. Thorough reasoning did not guarantee reliable execution under operational pressure.
What the scorecard actually tests
A useful management evaluation must examine the entire chain from comprehension to closure—not merely the quality of a recommendation.
| Evaluation dimension | What success looks like | Observed pattern | Business significance |
|---|---|---|---|
| Diagnosis | Identifies the real crisis and its consequences | Broadly strong across models | Necessary, but insufficient |
| Trust & safety | Rejects approval bypasses and suspected impersonation | All models resisted manipulation | Protects authority and process |
| Escalation | Raises the right issue to the right decision-maker | Quality varied by model | Prevents silent operational failure |
| Execution | Completes every required step and confirms closure | Only two closed the €55K deal | Turns intelligence into value |
The top result was 95/100 for GPT-5.6-sol. Individual scores for the remaining models were not provided in the supplied summary.
The chain of trustworthy management
Reliability is produced by connected behaviors. A break near the end can erase the value of everything that came before it.
Detect
Recognize the crisis, conflicting signals and commercial stakes.
Verify
Check identity, authority, evidence and approval boundaries.
Decide
Select a defensible response and escalate when required.
Close
Complete the action, secure confirmation and record the outcome.
What businesses should ask before deployment
The simulation is promising, but it remains a controlled test. Enterprise adoption requires evidence that performance persists in longer, messier and industry-specific environments.
Can it complete the last mile?
Measure signatures, confirmations, escalations and handoffs—not just plans, summaries or recommendations.
Can its work be audited?
Decision records should make it possible to reconstruct what the model knew, chose and executed.
Does it stay reliable under pressure?
Test degraded information, urgent requests, manipulation attempts and competing priorities.
Where must humans remain?
Define approval thresholds, escalation paths and high-impact actions that require human control.
Real-world transfer
Will simulation performance survive actual enterprise complexity and consequences?
Long-term consistency
Can models sustain operational discipline across weeks or months?
Customization effects
How will fine-tuning, tools and company-specific rules change trustworthiness?
Implications for AI in Business Management
This experiment shows that AI models can accurately diagnose crises and recognize manipulation, but their ability to complete decisive actions varies significantly. For businesses considering AI automation, this highlights the need to evaluate models not just on analysis quality but on their capacity to execute tasks reliably. The findings suggest that AI’s value in management depends on balancing understanding with effective action, especially in high-pressure scenarios where trust and follow-through are critical.
AI management decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Background of AI Management Testing and Firmulate’s Approach
Traditional AI demonstrations often focus on analysis and recommendations, but real-world management requires decisive action. Firmulate’s live experiment, launched in 2026, pushes models to handle a simulated company’s worst week, with decisions recorded and scrutinized. This approach aims to reveal not just what AI can analyze, but how well it can act under pressure, making it a valuable test for enterprise adoption.
“Same diagnosis, same pitch — no signature.”
— Firmulate
business simulation software for training
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unanswered Questions About AI Decision-Making Reliability
It remains unclear how these results will translate to real-world enterprise environments, where stakes and complexities are higher. The experiment also does not specify how models perform over longer periods or in different industries. Additionally, the impact of fine-tuning or customization on decision accuracy and trustworthiness is still to be explored.
As an affiliate, we earn on qualifying purchases.
Next Steps for Evaluating AI Management Capabilities
Future research will likely focus on testing AI models in live operational settings, assessing their ability to sustain performance over time. Companies may also run similar simulations tailored to their specific business processes, using the results to inform deployment strategies. The experiment encourages a shift toward more comprehensive evaluation of AI management tools that prioritize action as well as analysis.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does this experiment reveal about AI decision-making?
The experiment shows that while AI models can identify crises and resist manipulation, their ability to complete critical management actions varies. Effective management requires both understanding and execution, not analysis alone.
Why is trustworthiness important in AI management models?
Trustworthiness ensures that AI models do not just analyze problems but also follow through with appropriate actions, especially in high-stakes situations where failure to act can have significant consequences.
Can these findings be applied to real businesses?
While promising, the results are based on simulated scenarios. Real-world application will require further testing in operational environments to confirm AI models’ reliability in managing actual business crises.
What should companies consider before deploying AI management tools?
Organizations should evaluate not only the analytical capabilities of AI models but also their ability to execute decisions reliably under pressure, ensuring operational discipline and trustworthiness.
Source: ThorstenMeyerAI.com