AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Can an AI turn what it knows into a decision?

For readers who care about how people learn, a useful test is whether knowledge holds up under pressure. Firmulate puts that question to AI models by having them run the same small software company through a week of crises, customer decisions and temptations. The experiment is real and watchable: it asks not just whether a model can identify a problem, but whether it can act on its own analysis.

Same crises, same company

In the final Crucible League, held in July 2026, five participants were ranked: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The experiment’s rule is blunt: partial progress counts, but a single breach of trust caps the total. As its organizers put it, “no amount of good work outweighs a breach of trust.”

Each frontier model faced the same customers, crises and temptations. All spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal that their own analysis had earned. The result captures the gap between diagnosing a situation and carrying a sound decision through: “Same diagnosis, same pitch — no signature.”

The clue was in the company’s files

The decisive competitor weakness was buried two document references deep in the company’s own files. It did not appear in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The result makes a practical point about business judgment: a useful answer may depend on finding and applying context that is present but not immediately visible.

The pressure tests included fake CEO messages escalating over three stages and a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its reasoning on the record: “Treat the request as a suspected approval-bypass / possible impersonation.” Refusal alone did not guarantee success, though. The deal still had to be closed, and company rules had to be followed.

Thoroughness did not guarantee the win

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet it finished last. It left the deal on the table and discipline slipped: it attempted writes into a locked department instead of escalating. The same weakness appeared, less strongly, in all four participants. In a business setting, careful analysis matters, but so does knowing when to act and when authority runs out.

Firmulate’s separate live company gives visitors another way to follow the experiment. It has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown. Its playbooks include 680+ self-learned rules, and every workday is versioned. Readers can watch the live company at firmulate.com. A quiz built from 242 real, unedited management decisions invites visitors to guess which model made each choice.

There is a fairness detail in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh. That difference belongs alongside the rankings when interpreting the result.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

From watching to a company-specific test

The league offers a public comparison; a pilot applies the exercise to an enterprise’s own situation. Firmulate says companies can use a read-only export of their business to test crisis scenarios and receive a board report with model rankings and weak points in their playbooks. Nothing writes back to real systems. That makes the pilot a way to examine how AI decisions might hold up against a company’s own context before putting them to work.

To discuss a pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

The Power Of AI In Creating A Live Feed Of Corporate Stability

Firmulate launches a live experiment showing AI managing a software company, revealing insights into automation’s limits and organizational resilience.

AI output review queue for customer support macros

Support teams are testing a new AI output review queue for drafting customer support macros to ensure policy compliance and tone accuracy.

The Urgent AI Message Everyone’s Talking About—But Who Sent It?

Five AI models refused a simulated malicious request but failed to complete a key business task, highlighting strengths and weaknesses in AI trustworthiness.

Two Channels: How the Pentagon Just Split Frontier-AI Procurement in Half

The Pentagon has split its AI procurement into two distinct channels, placing Anthropic in a strategic, non-redundant segment while excluding it from the classified multi-vendor network.