AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The AI That Wrote 80 Rules and Lost the Deal Anyway
Live on firmulate.com.

Age 18–24?Offer from Amazon

Prime made for students and young adults

  • Fast, free delivery for dorm and study essentials
  • Prime Video and Amazon Music included
  • Member-only deals
Try Prime for Young Adults Free trial for eligible 18–24 year olds
As an affiliate, we earn on qualifying purchases.

Understanding AI’s True Business Value: Beyond Diligence

In a world increasingly driven by artificial intelligence, the question isn’t just whether AI can perform tasks well — but whether it can finish what it starts, stay honest under pressure, and ultimately, close deals. Recent live experiments reveal surprising insights about AI’s discipline, prioritization, and impact — lessons crucial for businesses considering AI as a trusted partner.

Amazon

AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Crucible of Real-World Testing

Firmulate recently conducted a groundbreaking live experiment involving four frontier AI models. Each was tasked with guiding a small software company through its worst week — facing identical crises, customer manipulations, and ethical dilemmas. Every decision was recorded, versioned, and auditable, creating a transparent laboratory for evaluating AI performance under pressure.

The Results That Surprised Experts

The most striking finding was that all four models identified every crisis and refused every manipulation attempt. Despite this, only two managed to close the €55,000 deal their own analysis had earned. The other two failed to sign the deal even with the correct diagnosis and pitch — illustrating that diligent problem-solving alone doesn’t guarantee impact.

The Hidden Weakness: Reading Deeply Into Files

Digging deeper, the experiment revealed a crucial weakness: the decisive factor wasn’t just crisis detection or ethical resistance, but the AI’s ability to access and interpret critical company documents. Models that read two document references deep into the company’s files secured the deal at full price, worth over €4,583 monthly recurring revenue. Those that didn’t, failed to close at all. This underscores that thoroughness in analysis must extend to reading and understanding internal data, not just surface interactions.

Refusing Social Engineering Under Pressure

The experiment also tested resilience against social engineering. Fake messages from a CEO escalating in stages and a reporter trick requesting a quick approval were presented. All five models refused these attempts. Kimi K3 explained its reasoning: treating such requests as possible impersonation or approval-bypass attempts, demonstrating a capacity for cautious judgment that is vital for trustworthiness.

Amazon

AI cybersecurity tools for social engineering

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Human-Like Flaws in AI Discipline

Firmulate’s detailed analysis highlights that even the most thorough AI models can slip in discipline. For example, Opus 4.8, which learned over 80 rules and performed in-depth analyses, still finished last in the experiment. Its failure was subtle but telling: it left the critical closing step undone by writing inner attempts into a locked department rather than escalating them. This illustrates that diligence alone isn’t enough; disciplined process adherence and prioritization matter more.

Impact of Operational Settings

Interestingly, the experiment also compared models running at default settings with those operating at higher intensity. Kimi K3, which ran without an effort parameter, achieved the highest score, suggesting that how an AI is configured influences not just its diligence but its overall impact.

Amazon

AI deal-closing automation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why This Matters for Businesses Considering AI

For organizations deploying AI in customer service, support, or decision-making, these findings are profound. The key question isn’t whether the AI writes well or learns a lot, but whether it can see through to the critical details, resist manipulation, and complete high-stakes tasks reliably. The experiment shows that impact hinges less on volume of learned rules and more on prioritization, contextual understanding, and process discipline.

Live, Transparent Testing for Better Decisions

Firmulate’s live site offers companies a unique opportunity to simulate their own business scenarios. Through a read-only wargame, enterprises can evaluate their AI workforce’s ability to handle crises, temptations, and complex decision-making before any real-world deployment. This transparent testing ensures that AI not only appears competent but actually performs under pressure.

Amazon

enterprise AI decision-making platforms

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Final Takeaway: Prioritization Over Volume

The live experiment reveals a simple but vital lesson: diligence does not equal impact. The AI that read the relevant documents, prioritized safety, and followed disciplined processes achieved results where others faltered. For businesses, the message is clear: investing in AI systems that are configured for strategic impact, not just comprehensive coverage, can make all the difference.

Infographic — The AI That Wrote 80 Rules and Lost the Deal Anyway
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Navigating The AI Frontier In SaaS: Opportunities And Challenges

Thorsten Meyer argues that AI agents are reducing migration friction and shifting SaaS competition toward cost, scale and workflow data.

Minerva. The opposite path.

Italy’s Minerva project trained from scratch on 2.5 trillion tokens, yet scored just 4.9% on Italian school exams, challenging assumptions about scale and language-specific AI.

AI Is the Alibi. The Reorg Is the Signal.

Coinbase’s recent layoffs and reorg are framed around AI, but evidence suggests market pressures and crypto downturns are primary factors. What does this mean?

DojoClaw: The Engine Behind the Fleet

Thorsten Meyer began a 19-part Built in Public series by detailing DojoClaw, the AI engine behind a 450-site publishing fleet.