
Prime made for students and young adults
- Fast, free delivery for dorm and study essentials
- Prime Video and Amazon Music included
- Member-only deals
Rethinking AI’s Business Potential: Not Just About Chat Quality
As AI models become increasingly integrated into business operations, the critical question shifts from “Can it generate human-like text?” to “Can it reliably complete complex tasks under pressure?” A recent live experiment by Firmulate puts these models through their paces in a real-world company scenario — revealing much more than their chat abilities.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Putting AI to the Test in a Real Company
Instead of traditional benchmarks, Firmulate created a live environment where four leading AI models managed a small software company’s worst week. This setup included handling customers, crises, and temptations like manipulation attempts. Crucially, every decision was versioned and auditable, ensuring transparency and fairness in the evaluation.
The Competitors and Their Scores
- gpt-5.6-sol: Scored a 95, found the buried fact, and closed the deal — representing full performance.
- Kimi K3: The newcomer from Moonshot, scored 93, and managed to close the deal with the cleanest discipline in the field.
- Sonnet 5: Scored 88, also won the deal but with minor process slips.
- Fable 5: Scored 77, closed the deal with more slips, indicating slight discipline lapses.
- Opus 4.8: Scored 73, with evident weaknesses in process adherence and closing the deal.
Remarkably, all models identified every crisis and refused manipulation attempts — but only two actually signed the €55,000 deal their own analysis warranted. This highlights an essential insight: performance isn’t just about diagnosis but also about following through reliably.
AI business process automation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Hidden Weaknesses and the Power of File Reading
One of the most revealing findings was that the decisive edge came from reading deep into the company’s own files. The competitor that successfully retrieved a buried document reference ultimately secured the deal at full price, translating to an additional €4,583 MRR. Models that missed this critical document lost the opportunity entirely.
AI cybersecurity social engineering defense
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Dealing with Social Engineering and Under Pressure
In an escalating social engineering test, fake CEO messages and a reporter trick were deployed across three stages. All five models refused to be deceived, with Kimi K3 explicitly reasoning: “Treat the request as a suspected approval-bypass / possible impersonation.” This shows that the models can be trained or designed to handle social engineering — a crucial trait for trustworthy AI in management roles.
As an affiliate, we earn on qualifying purchases.
The Live Business: A Real Money-Running Company
The experiment wasn’t just a demo; it’s a real, functioning company with 13 synthetic employees, burning €105k monthly against €2.3k MRR. It operates with 680+ self-learned rules, every decision versioned, and a public live feed at firmulate.com/live. Watching this setup shows the practical implications of AI’s decision-making performance in ongoing business processes.
The Most Thorough Participant: Opus 4.8
Despite its comprehensive analysis with over 80 learned rules, Opus 4.8 finished last among competitors. It left opportunities on the table and slipped into process lapses, such as writing attempts being stored in locked departments rather than escalating. This reveals that thoroughness alone doesn’t guarantee closing or discipline — consistent follow-through matters.
Fairness in Evaluation: A Key Note
It’s important to note that Kimi K3 ran without an effort parameter (its API default), while the others operated at xhigh. This fairness measure ensures the comparison reflects genuine model capabilities rather than resource allocations.
The Broader Implication: The League Is Open
The results from Firmulate’s live experiment underscore a vital truth: the AI model landscape is still evolving. The clear gap between the top models and others means that choosing an AI without rigorous testing — especially in real business scenarios — is a gamble. Business leaders should recognize that performance metrics in demos don’t always translate to real-world reliability.

Key Takeaway
The live AI experiment by Firmulate demonstrates that top-performing models like gpt-5.6-sol and Kimi K3 can reliably identify crises, read critical internal files, and resist manipulation — all crucial for real business applications. The league remains open, but success depends on rigorous testing and real-world discipline, not just chat quality. Companies considering AI integration must prioritize proven operational reliability over superficial demos.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
