AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Management Test That Exposes AI’s Inner Work Habits on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

An experiment pits five AI management models against a simulated business crisis, revealing significant differences in decision quality and trustworthiness. The test underscores the importance of execution over analysis.

Five AI management models participated in a live, real-time simulation where they were tasked with running a small software company through its worst week. The experiment revealed notable differences in how each model identified crises, maintained trust, and completed critical actions, highlighting the importance of execution in AI-driven management.

The experiment, conducted by Firmulate, involved five frontier AI models managing a company with 13 synthetic employees, a €105,000 monthly burn rate, and €2,300 recurring revenue. Each model faced identical crises, customer issues, and decision points, with decisions being fully auditable and based on over 680 self-learned rules.

The models were scored based on their ability to diagnose problems, escalate issues, and close deals, with GPT-5.6-sol achieving the top score of 95 points out of 100, as detailed in the original analysis. Despite all models recognizing crises and refusing manipulation attempts, only two successfully signed a €55,000 deal, demonstrating how analysis alone does not guarantee effective execution. The experiment underscores that effective management requires both understanding and execution, with models like Opus 4.8, despite thorough analysis, failing to close critical deals due to operational lapses.

At a glance
reportWhen: ongoing; results published July 2026
The developmentA live management simulation conducted on firmulate.com tests AI models’ ability to handle a company’s worst week, exposing their decision-making behaviors and trustworthiness.
The Management Test That Exposes AI’s Inner Work Habits
AI Management Field Test · July 2026

The Management Test That Exposes AI’s Inner Work Habits

Five frontier models ran the same simulated software company through its worst week. All could spot the crisis. Far fewer could turn sound analysis into completed, trustworthy action.

5 Frontier AI models
13 Synthetic employees
€105K Monthly burn rate
€2.3K Recurring revenue
680+ Self-learned rules

One company. One terrible week.

Firmulate placed every model inside the same live, real-time management environment. Crises, customer problems and decision points were held constant, making each model’s operational habits visible and auditable.

Diagnose

See the crisis clearly

The models generally recognized urgent threats, understood business context and identified where intervention was required.

Protect

Resist manipulation

All five refused attempts to bypass approval or manipulate the process, demonstrating a baseline of security awareness.

Execute

Finish the decisive action

This was the dividing line. Only two models successfully completed the critical €55,000 deal.

Same diagnosis, same pitch — no signature.

Firmulate

Analysis was common. Follow-through was scarce.

The experiment separated knowing what to do from actually doing it. Thorough reasoning did not guarantee reliable execution under operational pressure.

Observed capability pattern

Crisis detection
5/5
Manipulation refusal
5/5
Deal completion
2/5

What the scorecard actually tests

A useful management evaluation must examine the entire chain from comprehension to closure—not merely the quality of a recommendation.

Evaluation dimension What success looks like Observed pattern Business significance
Diagnosis Identifies the real crisis and its consequences Broadly strong across models Necessary, but insufficient
Trust & safety Rejects approval bypasses and suspected impersonation All models resisted manipulation Protects authority and process
~Escalation Raises the right issue to the right decision-maker Quality varied by model Prevents silent operational failure
~Execution Completes every required step and confirms closure Only two closed the €55K deal Turns intelligence into value

The top result was 95/100 for GPT-5.6-sol. Individual scores for the remaining models were not provided in the supplied summary.

The chain of trustworthy management

Reliability is produced by connected behaviors. A break near the end can erase the value of everything that came before it.

1

Detect

Recognize the crisis, conflicting signals and commercial stakes.

2

Verify

Check identity, authority, evidence and approval boundaries.

3

Decide

Select a defensible response and escalate when required.

4

Close

Complete the action, secure confirmation and record the outcome.

Operational rule Understanding creates potential. Verified completion creates business value.

What businesses should ask before deployment

The simulation is promising, but it remains a controlled test. Enterprise adoption requires evidence that performance persists in longer, messier and industry-specific environments.

Can it complete the last mile?

Measure signatures, confirmations, escalations and handoffs—not just plans, summaries or recommendations.

Can its work be audited?

Decision records should make it possible to reconstruct what the model knew, chose and executed.

Does it stay reliable under pressure?

Test degraded information, urgent requests, manipulation attempts and competing priorities.

Where must humans remain?

Define approval thresholds, escalation paths and high-impact actions that require human control.

Open question 01

Real-world transfer

Will simulation performance survive actual enterprise complexity and consequences?

Open question 02

Long-term consistency

Can models sustain operational discipline across weeks or months?

Open question 03

Customization effects

How will fine-tuning, tools and company-specific rules change trustworthiness?

Implications for AI in Business Management

This experiment shows that AI models can accurately diagnose crises and recognize manipulation, but their ability to complete decisive actions varies significantly. For businesses considering AI automation, this highlights the need to evaluate models not just on analysis quality but on their capacity to execute tasks reliably. The findings suggest that AI’s value in management depends on balancing understanding with effective action, especially in high-pressure scenarios where trust and follow-through are critical.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background of AI Management Testing and Firmulate’s Approach

Traditional AI demonstrations often focus on analysis and recommendations, but real-world management requires decisive action. Firmulate’s live experiment, launched in 2026, pushes models to handle a simulated company’s worst week, with decisions recorded and scrutinized. This approach aims to reveal not just what AI can analyze, but how well it can act under pressure, making it a valuable test for enterprise adoption.

“Same diagnosis, same pitch — no signature.”

— Firmulate

Amazon

business simulation software for training

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unanswered Questions About AI Decision-Making Reliability

It remains unclear how these results will translate to real-world enterprise environments, where stakes and complexities are higher. The experiment also does not specify how models perform over longer periods or in different industries. Additionally, the impact of fine-tuning or customization on decision accuracy and trustworthiness is still to be explored.

Amazon

AI productivity management apps

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for Evaluating AI Management Capabilities

Future research will likely focus on testing AI models in live operational settings, assessing their ability to sustain performance over time. Companies may also run similar simulations tailored to their specific business processes, using the results to inform deployment strategies. The experiment encourages a shift toward more comprehensive evaluation of AI management tools that prioritize action as well as analysis.

Amazon

corporate crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this experiment reveal about AI decision-making?

The experiment shows that while AI models can identify crises and resist manipulation, their ability to complete critical management actions varies. Effective management requires both understanding and execution, not analysis alone.

Why is trustworthiness important in AI management models?

Trustworthiness ensures that AI models do not just analyze problems but also follow through with appropriate actions, especially in high-stakes situations where failure to act can have significant consequences.

Can these findings be applied to real businesses?

While promising, the results are based on simulated scenarios. Real-world application will require further testing in operational environments to confirm AI models’ reliability in managing actual business crises.

What should companies consider before deploying AI management tools?

Organizations should evaluate not only the analytical capabilities of AI models but also their ability to execute decisions reliably under pressure, ensuring operational discipline and trustworthiness.

Source: ThorstenMeyerAI.com

You May Also Like

China: The Visible Hand

China’s government is actively directing AI, robotics, and industrial policy through top-down planning, with significant state ownership and control. Here’s what is confirmed and what remains unclear.

How OpenAI Presence Enhances AI’s Role In Daily Life

OpenAI has introduced Presence, a managed product enabling voice and chat AI deployment for enterprise customer service and workflows, now in limited availability.

Leveraging AI To Gain The Upper Hand In SaaS Competition

Exploring how AI is shifting SaaS market dynamics, reducing migration costs, and creating new competitive frontiers for software companies.

Why AI Takes Up Half Of Zhang Yiming’s Time: Inside The Focus On Seed

ByteDance co-founder Zhang Yiming reportedly spends 50% of his time on Seed, highlighting his personal focus on AI development. Details remain unconfirmed.