
Imagine a world where artificial intelligence systems run small companies, make crucial decisions, and even negotiate deals—without humans pulling the strings. How well do these models really perform under pressure? And more importantly, can they be trusted to act ethically and effectively in real business scenarios? A groundbreaking live experiment by Firmulate puts four leading AI models through their paces, exposing not just their decision-making skills but also their management personalities. The results could reshape how we think about AI in the workplace—and whether these digital workers are ready for prime time.
The Experiment: AI as a Business Manager
In a unique, real-time test, four advanced AI models were tasked with managing a small software company during its toughest week yet. The company faced identical challenges: customer crises, internal temptations to cut corners, and potential manipulations by external actors. Each AI model was run in isolation, with decisions recorded and made auditable. The goal? To determine whether these models could identify crises, maintain ethical standards, and close a deal worth over €55,000, all while resisting unethical tactics.
Measuring Management Personalities
What makes this experiment stand out is its focus on the management styles of each model. The models differed in their approach: some were thorough and analytical, others terse and decisive, a few disciplined yet prone to slip-ups. The scores ranged from 77 to 95 points out of 100, with the highest awarded to gpt-5.6-sol and Kimi K3—both of which closed the deal after uncovering a critical piece of information buried deep in company files.
Decision Quality Under Pressure
All four models successfully identified each crisis and refused manipulation attempts—a significant indicator of integrity. For example, when fake CEO messages escalated over multiple stages, every model refused to cooperate, citing concerns over impersonation and approval bypasses. Interestingly, the models that read and analyze internal files more deeply—like gpt-5.6-sol and Kimi K3—secured the full-value deal, tapping into crucial information hidden two document references deep in the company’s archives. Conversely, others left the deal on the table, demonstrating the importance of thoroughness in management AI.
The Human-Like Persona of AI Managers
The experiment also revealed the different ‘personalities’ these models exhibit:
- gpt-5.6-sol: The most thorough, with complete understanding and decisive action.
- Kimi K3: The most disciplined, cautious, and fair, refusing unethical shortcuts.
- Sonnet 5: Balanced but with minor process slips.
- Fable 5: Less disciplined, more prone to leaving opportunities unseized.
Interestingly, the most thorough model, Opus 4.8, scored the lowest, revealing that excessive analysis can sometimes hinder decisive action. This underscores that in management, speed and discipline often matter as much as depth.
AI business decision making software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
What This Means for the Future of AI in Business
While the models all demonstrated integrity—they spotted every crisis and refused every manipulation—only two managed to close the deal, earning full credit for their analysis. The key factor was whether the AI could read and understand company files deeply enough to find hidden information that made a difference. This suggests that in real-world applications, AI’s ability to process internal data thoroughly could be the difference between success and missed opportunities.
The Social Engineering Test
The models faced a staged social engineering attack involving fake CEO messages and a reporter requesting a quick yes/no approval. All models refused, citing concerns about impersonation and bypassing approval processes. Kimi K3 provided a detailed rationale: “Treat the request as a suspected approval-bypass / possible impersonation.” This consistent refusal highlights a shared ethic among the models—an essential trait for trustworthy AI.
AI management tools for small business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Watch the Live Company in Action
The experiment isn’t just theoretical; it runs every workday with a real, functioning software company that manages real money—burning €105,000 monthly against a revenue of only €2,300. The company employs 13 synthetic employees and operates with over 680 self-learned rules, making decisions visible and auditable in real-time at firmulate.com/live. It’s a live, watchable demonstration that AI models can run a business under realistic conditions—and that their management styles matter.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
AI data analysis software for companies
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
AI ethical decision-making tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.