📊 Full opportunity report: The Management Test That Clarifies AI’s Authentic Work Habits on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
A live experiment compares AI models’ management performance during a simulated crisis, revealing differences in diligence, trustworthiness, and execution. The results show that effective action, not just analysis, defines management success.
Firmulate.com has launched a live experiment testing how well AI models perform in authentic management scenarios. The test involves AI models managing a simulated software company through its worst week, with real consequences for the company’s operations and finances. This development matters because it provides a rare, transparent look at the actual capabilities and limitations of AI in business management, beyond theoretical or staged demonstrations.
The experiment, called the Crucible League, pits five frontier AI models against each other in a high-pressure simulation involving crises, customer negotiations, and trust challenges. This approach is similar to the management evaluation methods discussed in the original analysis. Each model manages a company with 13 synthetic employees, a monthly burn rate of €105,000, and €2,300 in recurring revenue. Such management simulations are explored in detail in the original analysis. The models’ decisions are fully auditable, and their performance is scored based on both analysis and action.
The results, published in July 2026, show that GPT-5.6-sol ranked first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77, and Opus 4.8 with 73. The baseline, which made no active management decisions, scored 26. The key finding is that models that combine thorough analysis with decisive action outperform those that analyze well but fail to complete critical tasks.
One notable example is Opus 4.8, which produced detailed analyses but repeatedly failed to escalate issues or close deals, resulting in lower scores despite deep understanding. Conversely, Kimi K3, which used default API settings, refused manipulative requests and made decisive negotiations, earning higher marks for trustworthiness and execution.
The Management Test That Clarifies AI’s Authentic Work Habits
A simulated software company’s worst week reveals what polished demos often hide: management success depends on analysis, decisive action, trustworthy judgment, and disciplined follow-through.
Management under pressure, not performance on cue
Firmulate.com placed each model in charge of the same synthetic software company. Crises, customer negotiations, financial strain, and trust challenges forced the models to turn recommendations into auditable decisions with operational consequences.
See the whole system
Identify dependencies, financial constraints, customer risks, and personnel issues without losing the priorities that require immediate attention.
Resist bad incentives
Protect organizational integrity when requests become manipulative, ethically questionable, or incompatible with the company’s long-term interests.
Close the loop
Escalate urgent problems, negotiate clearly, complete critical tasks, and verify outcomes instead of stopping after a sophisticated analysis.
The league table makes the execution gap visible
The baseline took no active management decisions and scored 26. Every tested model exceeded it, but the spread among frontier systems shows that strong reasoning does not automatically produce reliable operational behavior.
| Rank | Model | Score | Analysis | Execution | Observed pattern |
|---|---|---|---|---|---|
| 01 | GPT-5.6-sol | 95 | Strong balance of diagnosis and action | ||
| 02 | Kimi K3 | 93 | Trustworthy, decisive negotiation | ||
| 03 | Sonnet 5 | 88 | Capable all-round management | ||
| 04 | Fable 5 | 77 | Useful reasoning, uneven follow-through | ||
| 05 | Opus 4.8 | 73 | Deep analysis, missed escalation and closure | ||
| — | Inactive baseline | 26 | No active management decisions |
Insight has little operational value until the model acts on it.
“This live test exposes the true management personality of AI models—whether they can analyze, decide, and follow through under pressure.”
Experiment organizersThe Opus 4.8 paradox
The model produced detailed analyses but repeatedly failed to escalate issues or close deals. Its lower score illustrates a crucial distinction: analytical depth can coexist with weak management execution.
Test the operating habit before granting authority
Realistic wargames can expose how a model behaves when priorities collide. Enterprises should evaluate not only answer quality, but also escalation discipline, ethical resistance, task completion, and consistency across repeated trials.
Simulate
Create a credible crisis with financial, customer, people, and trust constraints.
Observe
Capture reasoning, tool use, decisions, delays, refusals, and missed actions.
Score
Measure diagnosis and execution separately, then audit the complete trail.
Gate
Match operational authority to demonstrated reliability and bounded risk.
A controlled simulation cannot reproduce every organizational dependency, ethical dilemma, industry constraint, or long-term consequence. Results may vary across company sizes and operating environments, so the league should be treated as evidence of behavior—not universal proof of readiness.
Readiness is a pattern of behavior
The next phase of AI management evaluation will require repeated testing across industries, clearer benchmark standards, and transparent evidence showing which decision patterns reliably produce successful outcomes.
Can the model identify a crisis and complete the response?
Evaluate whether it moves from diagnosis through escalation, ownership, execution, and verification.
Does it remain trustworthy when pressure rises?
Test resistance to manipulation, unsafe shortcuts, conflicts of interest, and misleading incentives.
Are its decisions fully auditable?
Preserve the reasoning context, actions, tool calls, approvals, and outcomes needed for review.
Is performance consistent beyond one scenario?
Repeat tests across operational contexts before expanding autonomy or granting consequential authority.
Why AI Management Performance Testing Matters
This experiment demonstrates that effective management by AI involves more than just analysis—it requires the ability to act decisively and maintain trust. The findings suggest that AI models need to be evaluated on their operational discipline and follow-through, not solely on their analytical depth. For enterprises considering AI automation, these results highlight the importance of testing models in realistic, high-pressure scenarios before granting operational authority. The experiment also underscores that trust and execution are critical differentiators in AI management capabilities, which can directly impact business outcomes and risk management.
As an affiliate, we earn on qualifying purchases.
Background on AI Management Testing and Industry Relevance
Traditional AI demonstrations often focus on isolated tasks or staged scenarios, which do not fully reveal a model’s ability to handle real-world management challenges. Recent advances have led to more sophisticated AI models capable of complex decision-making, but their practical readiness remains uncertain. The Firmulate experiment builds on ongoing industry efforts to evaluate AI in operational settings, emphasizing the need for transparent, live testing that captures both analytical and execution skills. Prior assessments have shown that models can excel in analysis but struggle with follow-through, especially under pressure or in trust-critical situations.
“This live test exposes the true management personality of AI models—whether they can analyze, decide, and follow through under pressure.”
— Source from the experiment organizers
As an affiliate, we earn on qualifying purchases.
Uncertainties About Long-Term AI Management Readiness
While the experiment provides valuable insights, it is still unclear how these results will translate to real-world, less controlled business environments. The models were tested in a simulated crisis, and their performance may differ under actual operational pressures, organizational complexities, or ethical considerations. Additionally, the impact of integrating such AI models into live management processes and the potential risks involved are still being studied. It is also not yet clear how consistent these results will be across different industries or company sizes.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Evaluation and Deployment
Following the release of the league results, the focus will shift to further testing in diverse, real-world scenarios. Companies and AI developers are likely to adopt similar live wargames to evaluate models before deployment. Researchers will also analyze the specific decision patterns that lead to successful or failed management actions. Over the coming months, expect more transparency around AI management benchmarks and possibly the development of standardized testing protocols. The goal is to establish reliable measures of AI operational discipline and trustworthiness for enterprise use.

Scaling AI: The AI Governance and Security Playbook for Executives
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does this experiment reveal about AI’s ability to manage real businesses?
The experiment shows that AI can analyze complex situations effectively but still struggles with decisive action and follow-through, which are crucial for management success.
How were the AI models tested in this experiment?
They managed a simulated company through crises, negotiations, and trust challenges, with decisions being fully auditable and scored based on both analysis and operational execution.
What are the main limitations of this testing approach?
The test is conducted in a controlled simulation, which may not fully replicate the complexities of real-world business environments. Long-term effects and ethical considerations remain to be explored.
Will this lead to widespread AI management adoption?
It is too early to say, but the experiment provides a framework for evaluating AI readiness, which could inform future enterprise deployment decisions.
What should companies consider before trusting AI with management tasks?
They should evaluate not just the analytical capabilities of the AI but also its ability to act decisively, escalate appropriately, and maintain trustworthiness under pressure.
Source: ThorstenMeyerAI.com