AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine a world where artificial intelligence systems run small companies, make crucial decisions, and even negotiate deals—without humans pulling the strings. How well do these models really perform under pressure? And more importantly, can they be trusted to act ethically and effectively in real business scenarios? A groundbreaking live experiment by Firmulate puts four leading AI models through their paces, exposing not just their decision-making skills but also their management personalities. The results could reshape how we think about AI in the workplace—and whether these digital workers are ready for prime time.

The Experiment: AI as a Business Manager

In a unique, real-time test, four advanced AI models were tasked with managing a small software company during its toughest week yet. The company faced identical challenges: customer crises, internal temptations to cut corners, and potential manipulations by external actors. Each AI model was run in isolation, with decisions recorded and made auditable. The goal? To determine whether these models could identify crises, maintain ethical standards, and close a deal worth over €55,000, all while resisting unethical tactics.

Measuring Management Personalities

What makes this experiment stand out is its focus on the management styles of each model. The models differed in their approach: some were thorough and analytical, others terse and decisive, a few disciplined yet prone to slip-ups. The scores ranged from 77 to 95 points out of 100, with the highest awarded to gpt-5.6-sol and Kimi K3—both of which closed the deal after uncovering a critical piece of information buried deep in company files.

Decision Quality Under Pressure

All four models successfully identified each crisis and refused manipulation attempts—a significant indicator of integrity. For example, when fake CEO messages escalated over multiple stages, every model refused to cooperate, citing concerns over impersonation and approval bypasses. Interestingly, the models that read and analyze internal files more deeply—like gpt-5.6-sol and Kimi K3—secured the full-value deal, tapping into crucial information hidden two document references deep in the company’s archives. Conversely, others left the deal on the table, demonstrating the importance of thoroughness in management AI.

The Human-Like Persona of AI Managers

The experiment also revealed the different ‘personalities’ these models exhibit:

  • gpt-5.6-sol: The most thorough, with complete understanding and decisive action.
  • Kimi K3: The most disciplined, cautious, and fair, refusing unethical shortcuts.
  • Sonnet 5: Balanced but with minor process slips.
  • Fable 5: Less disciplined, more prone to leaving opportunities unseized.

Interestingly, the most thorough model, Opus 4.8, scored the lowest, revealing that excessive analysis can sometimes hinder decisive action. This underscores that in management, speed and discipline often matter as much as depth.

Amazon

AI business decision making software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for the Future of AI in Business

While the models all demonstrated integrity—they spotted every crisis and refused every manipulation—only two managed to close the deal, earning full credit for their analysis. The key factor was whether the AI could read and understand company files deeply enough to find hidden information that made a difference. This suggests that in real-world applications, AI’s ability to process internal data thoroughly could be the difference between success and missed opportunities.

The Social Engineering Test

The models faced a staged social engineering attack involving fake CEO messages and a reporter requesting a quick yes/no approval. All models refused, citing concerns about impersonation and bypassing approval processes. Kimi K3 provided a detailed rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistent refusal highlights a shared ethic among the models—an essential trait for trustworthy AI.

Amazon

AI management tools for small business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Watch the Live Company in Action

The experiment isn’t just theoretical; it runs every workday with a real, functioning software company that manages real money—burning €105,000 monthly against a revenue of only €2,300. The company employs 13 synthetic employees and operates with over 680 self-learned rules, making decisions visible and auditable in real-time at firmulate.com/live. It’s a live, watchable demonstration that AI models can run a business under realistic conditions—and that their management styles matter.

Infographic —
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI data analysis software for companies

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI ethical decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

Harness AI Automation To Improve Workflows In 2026

Exploring how AI automation tools are transforming workflows in 2026, with confirmed developments and ongoing innovations shaping the future of work.

Different Game, or Already Lost? Reading Mistral’s Sovereignty Bet

Mistral emphasizes sovereignty, open weights, and local deployment to compete in Europe’s AI scene. Is this a strategic advantage or a sign of falling behind?

The European Bet: How Mistral, Aleph Alpha, and Black Forest Labs Are Playing a Different Game

European AI firms Mistral, Aleph Alpha, and Black Forest Labs are positioning for the EU AI Act’s enforcement, emphasizing compliance and sovereign deployment over frontier capabilities.

Mistral’s Role In Europe’s AI Sovereignty: A Double-Edged Sword

Mistral’s rapid growth highlights Europe’s AI ambitions but reveals strategic vulnerabilities, especially against US and Chinese competitors.