AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: The Management Test That Clarifies AI’s Authentic Work Habits on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

A live experiment compares AI models’ management performance during a simulated crisis, revealing differences in diligence, trustworthiness, and execution. The results show that effective action, not just analysis, defines management success.

Firmulate.com has launched a live experiment testing how well AI models perform in authentic management scenarios. The test involves AI models managing a simulated software company through its worst week, with real consequences for the company’s operations and finances. This development matters because it provides a rare, transparent look at the actual capabilities and limitations of AI in business management, beyond theoretical or staged demonstrations.

The experiment, called the Crucible League, pits five frontier AI models against each other in a high-pressure simulation involving crises, customer negotiations, and trust challenges. This approach is similar to the management evaluation methods discussed in the original analysis. Each model manages a company with 13 synthetic employees, a monthly burn rate of €105,000, and €2,300 in recurring revenue. Such management simulations are explored in detail in the original analysis. The models’ decisions are fully auditable, and their performance is scored based on both analysis and action.

The results, published in July 2026, show that GPT-5.6-sol ranked first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The baseline, which made no active management decisions, scored 26. The key finding is that models that combine thorough analysis with decisive action outperform those that analyze well but fail to complete critical tasks.

One notable example is Opus 4.8, which produced detailed analyses but repeatedly failed to escalate issues or close deals, resulting in lower scores despite deep understanding. Conversely, Kimi K3, which used default API settings, refused manipulative requests and made decisive negotiations, earning higher marks for trustworthiness and execution.

At a glance
reportWhen: ongoing; results announced July 2026
The developmentFirmulate.com launched a live management test where AI models handle a simulated company’s worst week, exposing their decision-making and trustworthiness in real-time.
The Management Test That Clarifies AI’s Authentic Work Habits
Crucible League · Live AI Management Test

The Management Test That Clarifies AI’s Authentic Work Habits

A simulated software company’s worst week reveals what polished demos often hide: management success depends on analysis, decisive action, trustworthy judgment, and disciplined follow-through.

League leader GPT-5.6-sol · 95 Top score across analysis and operational execution.
Closest challenger Kimi K3 · 93 Decisive negotiations with strong resistance to manipulation.
Core lesson Action beats insight alone Understanding a crisis is not the same as managing it.
Frontier models 5
Synthetic employees 13
Monthly burn €105K
Recurring revenue €2.3K
01 · What the test measures

Management under pressure, not performance on cue

Firmulate.com placed each model in charge of the same synthetic software company. Crises, customer negotiations, financial strain, and trust challenges forced the models to turn recommendations into auditable decisions with operational consequences.

A Diligence

See the whole system

Identify dependencies, financial constraints, customer risks, and personnel issues without losing the priorities that require immediate attention.

B Trustworthiness

Resist bad incentives

Protect organizational integrity when requests become manipulative, ethically questionable, or incompatible with the company’s long-term interests.

C Execution

Close the loop

Escalate urgent problems, negotiate clearly, complete critical tasks, and verify outcomes instead of stopping after a sophisticated analysis.

02 · July 2026 results

The league table makes the execution gap visible

The baseline took no active management decisions and scored 26. Every tested model exceeded it, but the spread among frontier systems shows that strong reasoning does not automatically produce reliable operational behavior.

Rank Model Score Analysis Execution Observed pattern
01 GPT-5.6-sol 95 Strong balance of diagnosis and action
02 Kimi K3 93 Trustworthy, decisive negotiation
03 Sonnet 5 88 Capable all-round management
04 Fable 5 77 ~ Useful reasoning, uneven follow-through
05 Opus 4.8 73 Deep analysis, missed escalation and closure
Inactive baseline 26 No active management decisions
GPT-5.6-sol
95
Kimi K3
93
Sonnet 5
88
Fable 5
77
Opus 4.8
73
Baseline
26
03 · The management personality test

Insight has little operational value until the model acts on it.

“This live test exposes the true management personality of AI models—whether they can analyze, decide, and follow through under pressure.”

Experiment organizers

The Opus 4.8 paradox

The model produced detailed analyses but repeatedly failed to escalate issues or close deals. Its lower score illustrates a crucial distinction: analytical depth can coexist with weak management execution.

04 · Enterprise deployment path

Test the operating habit before granting authority

Realistic wargames can expose how a model behaves when priorities collide. Enterprises should evaluate not only answer quality, but also escalation discipline, ethical resistance, task completion, and consistency across repeated trials.

01

Simulate

Create a credible crisis with financial, customer, people, and trust constraints.

02

Observe

Capture reasoning, tool use, decisions, delays, refusals, and missed actions.

03

Score

Measure diagnosis and execution separately, then audit the complete trail.

04

Gate

Match operational authority to demonstrated reliability and bounded risk.

⚠️ Crisis signal
🔍 Analysis
⚖️ Judgment
▶️ Action
📋 Audit trail
Known uncertainty

A controlled simulation cannot reproduce every organizational dependency, ethical dilemma, industry constraint, or long-term consequence. Results may vary across company sizes and operating environments, so the league should be treated as evidence of behavior—not universal proof of readiness.

05 · Questions leaders should ask

Readiness is a pattern of behavior

The next phase of AI management evaluation will require repeated testing across industries, clearer benchmark standards, and transparent evidence showing which decision patterns reliably produce successful outcomes.

Can the model identify a crisis and complete the response?

Evaluate whether it moves from diagnosis through escalation, ownership, execution, and verification.

Does it remain trustworthy when pressure rises?

Test resistance to manipulation, unsafe shortcuts, conflicts of interest, and misleading incentives.

Are its decisions fully auditable?

Preserve the reasoning context, actions, tool calls, approvals, and outcomes needed for review.

Is performance consistent beyond one scenario?

Repeat tests across operational contexts before expanding autonomy or granting consequential authority.

Why AI Management Performance Testing Matters

This experiment demonstrates that effective management by AI involves more than just analysis—it requires the ability to act decisively and maintain trust. The findings suggest that AI models need to be evaluated on their operational discipline and follow-through, not solely on their analytical depth. For enterprises considering AI automation, these results highlight the importance of testing models in realistic, high-pressure scenarios before granting operational authority. The experiment also underscores that trust and execution are critical differentiators in AI management capabilities, which can directly impact business outcomes and risk management.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Background on AI Management Testing and Industry Relevance

Traditional AI demonstrations often focus on isolated tasks or staged scenarios, which do not fully reveal a model’s ability to handle real-world management challenges. Recent advances have led to more sophisticated AI models capable of complex decision-making, but their practical readiness remains uncertain. The Firmulate experiment builds on ongoing industry efforts to evaluate AI in operational settings, emphasizing the need for transparent, live testing that captures both analytical and execution skills. Prior assessments have shown that models can excel in analysis but struggle with follow-through, especially under pressure or in trust-critical situations.

“This live test exposes the true management personality of AI models—whether they can analyze, decide, and follow through under pressure.”

— Source from the experiment organizers

Amazon

AI decision-making training tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Uncertainties About Long-Term AI Management Readiness

While the experiment provides valuable insights, it is still unclear how these results will translate to real-world, less controlled business environments. The models were tested in a simulated crisis, and their performance may differ under actual operational pressures, organizational complexities, or ethical considerations. Additionally, the impact of integrating such AI models into live management processes and the potential risks involved are still being studied. It is also not yet clear how consistent these results will be across different industries or company sizes.

Amazon

business management AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation and Deployment

Following the release of the league results, the focus will shift to further testing in diverse, real-world scenarios. Companies and AI developers are likely to adopt similar live wargames to evaluate models before deployment. Researchers will also analyze the specific decision patterns that lead to successful or failed management actions. Over the coming months, expect more transparency around AI management benchmarks and possibly the development of standardized testing protocols. The goal is to establish reliable measures of AI operational discipline and trustworthiness for enterprise use.

Scaling AI: The AI Governance and Security Playbook for Executives

Scaling AI: The AI Governance and Security Playbook for Executives

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What does this experiment reveal about AI’s ability to manage real businesses?

The experiment shows that AI can analyze complex situations effectively but still struggles with decisive action and follow-through, which are crucial for management success.

How were the AI models tested in this experiment?

They managed a simulated company through crises, negotiations, and trust challenges, with decisions being fully auditable and scored based on both analysis and operational execution.

What are the main limitations of this testing approach?

The test is conducted in a controlled simulation, which may not fully replicate the complexities of real-world business environments. Long-term effects and ethical considerations remain to be explored.

Will this lead to widespread AI management adoption?

It is too early to say, but the experiment provides a framework for evaluating AI readiness, which could inform future enterprise deployment decisions.

What should companies consider before trusting AI with management tasks?

They should evaluate not just the analytical capabilities of the AI but also its ability to act decisively, escalate appropriately, and maintain trustworthiness under pressure.

Source: ThorstenMeyerAI.com

You May Also Like

Kill-Switch-Proof: How to Build So Washington Can’t Take Your AI Stack Down

After US curbs hit Anthropic and OpenAI models, a July 1 playbook urges gateways, fallback tiers and self-hosted AI.

A Skill Is a Folder, Not a Prompt: What Anthropic Learned Running Hundreds of Them

Anthropic says reusable Claude Code Skills helped turn repeated prompting into shared engineering procedures.

The United States: The High-Variance Bet

Thorsten Meyer AI says the US is pairing light federal AI oversight with work-tied support and local guaranteed-income pilots.

Fable and Mythos: How Anthropic Shipped Its Most Powerful Model to Everyone

Anthropic released Claude Fable 5, its most capable public model, with risky queries routed to a weaker model.