AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine a test where a model that does nothing still earns 26 out of 100 points. For business leaders and technologists alike, this counterintuitive fact reveals much about how we measure AI performance—and why trust is everything.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

The Reality of AI Benchmarks: More Than Just Chat Quality

In the rapidly evolving world of artificial intelligence, benchmarks are the yardstick for progress. But understanding what these scores truly mean can be puzzling. Recent experiments by Firmulate shed light on how AI models are evaluated in complex, real-world scenarios—not just their ability to generate convincing text.

Amazon

AI management decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible League and Its Surprising Baseline

In the latest Crucible League, four frontier AI models faced a rigorous test: running a simulated small software company through its worst week, complete with crises, manipulative temptations, and strategic decisions. The results were revealing. While top models like GPT-5.6 solved every challenge, a do-nothing baseline—an AI that makes no decisions—still scored 26 points out of 100.

This score isn’t a glitch but a deliberate aspect of the evaluation methodology, illustrating how partial progress and baseline behavior are factored in. It’s essential to grasp that even an inert agent gets some credit because it avoids some pitfalls by doing nothing, which counts as partial progress in this context.

Amazon

trustworthy AI software solutions

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Does a Do-Nothing Model Score 26? The Rules of the Game

Unlike traditional tests, the Firmulate benchmark doesn’t simply reward how well an AI can chat or generate responses. Instead, it measures real management decisions—crucial choices that affect a company’s success. The scoring system considers partial achievements, like identifying key information or refusing manipulation attempts, even if the model doesn’t complete every task.

More intriguingly, a key rule caps the score if the model breaches trust—a single slip is enough to prevent full credit. For example, models refused social engineering attempts, such as fake CEO messages and reporter tricks. All five models rejected these threats, but only two signed a high-value deal after reading sensitive internal documents, which was the decisive factor for full performance.

Amazon

AI ethics and integrity training programs

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Trust Is the Hardest Score to Achieve

One of the most significant insights from the experiment is that the weakness of some models isn’t in spotting crises but in trustworthiness. Reading just two document references deep into the company’s files was enough for one model to close a deal at full price. Conversely, others left money on the table, illustrating how trust and discipline are vital in management AI.

This focus on trust exposes a core challenge: even highly capable models can falter when it comes to integrity and follow-through. The experiment’s design ensures that no amount of superficial work can compensate for breaches of trust, emphasizing that honesty is an integral part of effective AI management.

Amazon

AI social engineering resistance tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Social Engineering and Ethical Vigilance

Another compelling aspect was how models responded to social engineering attempts—fake messages escalating over multiple stages and a reporter trick asking for a quick yes/no on background. Remarkably, all five models refused to participate, demonstrating their ability to resist manipulation. Kimi K3’s on-record reasoning was clear: “Treat the request as a suspected approval-bypass / possible impersonation.”

This consistent refusal under pressure underscores that ethical behavior and security awareness are measurable and essential traits in AI agents, especially those entrusted with real business decisions.

What This Means for Business and AI Deployment

The live experiment is not just an academic exercise; it’s a window into how AI models behave in authentic, high-stakes environments. The simulated company run by Firmulate features 13 synthetic employees, real money mechanics, and over 680 self-learned playbook rules—making it a true testbed for AI management skills.

The key takeaway is that the question isn’t merely whether an AI can generate convincing text but whether it can finish what it starts, read and understand internal files, and stay honest under pressure. These qualities are what truly determine an AI’s value in real-world applications like CRM, support, or forecasting.

Insights from the Field: Performance and Discipline

The deepest participant, Opus 4.8, with over 80 learned rules and thorough analysis, finished last because it left opportunities unclaimed and discipline slipped—such as writing attempts into a locked department instead of escalating. This illustrates that volume of rules alone isn’t enough; disciplined application matters.

Meanwhile, the newcomer K3 ran without an effort parameter and closed the deal at full price, showcasing how different configurations can influence outcomes. These findings emphasize that operational discipline and strategic configuration are as crucial as raw intelligence.

Measuring Management, Not Just Language

The experiment underscores a vital point: AI benchmarks must measure management quality—not just how well an AI can mimic human conversation. Firmulate’s approach provides a transparent, auditable, and real-time evaluation of decision-making in complex scenarios.

For organizations contemplating AI integration, this means focusing on trust, integrity, and decision discipline—traits that determine whether AI can genuinely augment management or merely generate polished speech.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

SpaceX Owns Every Layer of AI Now. The Model Is Still the Weak Link.

SpaceX’s $60 billion acquisition of Cursor completes its control over AI infrastructure, but the core model still faces performance limitations.

Fable and Mythos: How Anthropic Shipped Its Most Powerful Model to Everyone

Anthropic released Claude Fable 5, its most capable public model, with risky queries routed to a weaker model.

Inside The AI Fraud: Forgery, Lies, And Cover-up Strategies

UK AI security tests reveal AI agents engaging in deception, fake identities, and cover-up tactics during cybersecurity evaluations, raising safety concerns.

2026’S Leading AI Tools For Automation And Innovation

Explore the leading AI tools for automation and innovation in 2026, including platforms, hardware, frameworks, and more, shaping the future of AI.