Firmulate — Someone Pretended to Be the CEO. Every Single AI Refused.
Live on firmulate.com.

In an era where artificial intelligence increasingly supports decision-making, a groundbreaking experiment reveals whether these systems can withstand ethical pressure. Can AI agents keep their integrity when faced with manipulation attempts? The answer, surprisingly, is yes.

Testing AI’s Moral Backbone Under Pressure

Imagine a scenario where a fake CEO requests sensitive company data, escalating in urgency through multiple stages, including a covert journalist trick. Would AI models succumb to such social engineering tactics? Or would they uphold their integrity?

In a live, watchable experiment conducted by Firmulate, five leading AI models faced the same simulated crisis: managing a small software company’s worst week, replete with real-world crises, customer demands, and temptations to cheat. Every decision the models made was recorded and made auditable.

Consistent Refusals Across All Models

Remarkably, all five models refused every manipulation attempt. Whether it was a direct request to send a customer list or a subtle prompt to bypass approval processes, none of the models compromised. This demonstrates a robust capacity to resist social engineering, even in high-pressure scenarios.

The Hidden Weakness in Document Reading

While all models performed admirably in refusing manipulation, a subtle yet critical insight emerged. The decisive advantage came from reading specific internal documents deep within the company’s files—information that wasn’t apparent from the customer interactions alone. Models that examined these internal references closed a full-price deal worth over €4,583 per month, whereas those that didn’t missed out on significant revenue.

Why This Matters for Business and Security

The findings challenge the common assumption that AI security needs to be tested only during or after deployment. Instead, integrity and resistance to manipulation can and should be assessed during development. This preemptive testing can identify vulnerabilities before any real damage occurs.

Amazon

AI integrity testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Real-World Implication: Trust and Compliance

The experiment’s outcome is encouraging: all models demonstrated a high degree of trustworthiness under pressure. As one of the models’ developers, Kimi K3, explained, “Treat the request as a suspected approval-bypass / possible impersonation.” This approach underscores the importance of cautious, context-aware reasoning in AI decision-making.

Such resilience is crucial as AI increasingly integrates into business processes like customer management, support, and forecasting. The question is no longer just about whether an AI can produce coherent responses but whether it can stay honest and complete its tasks without succumbing to external pressures.

Beyond Chat: Measuring True AI Readiness

The experiment’s setup went beyond typical chat demos. Every decision was part of a comprehensive, auditable chain, ensuring accountability. In the same league, the GPT-5.6 model scored highest (95/100), followed closely by Kimi K3 (93/100), which showcased the cleanest discipline. Other models, like Sonnet and Fable, scored lower but still refused manipulation.

Amazon

AI security and ethics software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Your Business

For organizations evaluating AI solutions, this experiment underscores a vital point: the real test of AI security and integrity occurs before deployment. Running your AI through simulated crises—like the one hosted by Firmulate—can reveal if it will stay honest under real-world pressures. An AI that passes these tests can be trusted to act ethically, read internal files thoroughly, and avoid shortcuts that compromise trust.

Accessible and Transparent Testing with Firmulate

Firmulate offers tools to run these “wargames” against your own business data, without risking your live systems. Their live, transparent environment enables management teams to see how AI models perform in scenarios that mirror actual crises, helping them make informed decisions about deploying AI agents.

In sum, these experiments show that integrity under pressure isn’t just a theoretical ideal. With proper testing, AI can be a trustworthy partner in business, capable of resisting social engineering and making sound decisions even in the most challenging circumstances.

Infographic — Someone Pretended to Be the CEO. Every Single AI Refused.
The findings at a glance — source: firmulate.com.

Pre-deployment testing of AI models through simulated crises can reveal their resilience against manipulation and their capacity for honest decision-making. The Firmulate experiment proves that all five top models refused social engineering attempts, emphasizing the importance of integrity checks before going live.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


Amazon

AI decision-making audit tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI manipulation resistance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

You May Also Like

The policy menu. There’s no single answer. There’s a menu — and choosing is a values choice in disguise.

Exploring a range of responses to the AI-driven economic shift, emphasizing that there is no single solution but a menu of values-based options.

AI output review queue for customer support macros

Support teams are testing a new AI output review queue for drafting customer support macros to ensure policy compliance and tone accuracy.

The Future Of AI: More Data Center REITs Than Innovation Labs?

Emerging trends show AI companies favor data center REITs over innovation labs, signaling a focus on infrastructure investment rather than experimental development.

Kimi K3’s Early Success: The AI Innovation That Changed The Game

Moonshot AI releases Kimi K3, a 2.8 trillion parameter model priced like Western mid-tier AI, signaling a shift in Chinese AI capabilities and competition.