
In an era where AI is touted as the ultimate multitasker, it’s tempting to believe that more rules and deeper analyses will secure the deal. But a recent live experiment challenges that assumption, revealing that diligence alone won’t cut it in the real world of decision-making under pressure.
Listen free for 30 days with Audible
Thousands of audiobooks and originals — cancel anytime.
As an affiliate, we earn on qualifying purchases.
The Live Experiment: Putting AI to the Test in a Simulated Business Crisis
Firmulate, a company specializing in AI management simulation, orchestrated a groundbreaking experiment to evaluate how different AI models handle complex, high-stakes scenarios. Four frontier models—gpt-5.6-sol, Kimi K3, Sonnet 5, and Fable 5—were tasked with running a small software firm through its most challenging week. The stakes? Real money, real crises, and the temptation to manipulate for personal gain.
All models faced the same conditions: the same customers, the same crises, and the same internal temptations, including a fake CEO message escalating over three stages and a reporter trick asking for a secret approval. Every decision made by the AI was recorded and verifiable, ensuring a transparent comparison of their approaches and integrity.
As an affiliate, we earn on qualifying purchases.
The Results: Diligence vs. Impact
Remarkably, all four models identified every crisis and refused every attempt at manipulation. Yet, only two managed to close the critical €55,000 deal that their own analyses had earned. The other two, despite thorough diagnosis and correct pitches, left the deal on the table—failing to act decisively when it mattered most.
The key issue was discipline. The Opus 4.8 profile, which was the most thorough participant—learning over 80 rules and conducting deep analyses—still finished last. Its downfall? A slip into complacency, where decision attempts were documented into a locked department instead of escalating appropriately. The same pattern, albeit less pronounced, appeared across all models.
As an affiliate, we earn on qualifying purchases.
The Hidden Weakness: Reading Between the Lines
Crucially, the decisive advantage went to models that went beyond surface-level judgments. Those that read two document references deep into the company’s files were able to discover a buried fact vital to securing the deal—adding over €4,500 in monthly recurring revenue (MRR). This underscores a vital insight: thoroughness isn’t just about volume; it’s about relevance and focus.
As an affiliate, we earn on qualifying purchases.
The Human Element: Trust and Integrity
The models’ refusal of social engineering attempts was uniform. All five models refused to sign off on fake CEO messages or background requests, with Kimi K3 explicitly reasoning that such requests resemble possible impersonation. This demonstrates that AI can be reliably disciplined to reject manipulative tactics designed to bypass controls—if the rules are embedded correctly.
As an affiliate, we earn on qualifying purchases.
The Takeaway: Impact Over Volume
In a world increasingly driven by AI decisions, the lesson is clear: more rules and deeper analyses do not automatically translate into better outcomes. The most thorough model, Opus 4.8, still finished last because it failed to prioritize key actions and allowed discipline to slip—highlighting an enduring truth: diligence alone is insufficient. Prioritization, focus, and disciplined escalation are the real differentiators.
For companies deploying AI, the message is straightforward: it’s not about how much your models analyze, but whether they can act decisively and trustworthily when stakes are high. The live experiment at firmulate.com/live exemplifies this in real time, where AI-driven decision-making is put through its paces with visible, auditable processes.

The experiment underscores a vital lesson: in high-stakes environments, impact and prioritization trump sheer diligence. AI models must learn to focus on key issues and escalate appropriately—otherwise, even the most thorough analysis remains useless if the close is lost on overlooked detail or discipline lapses.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Flea & tick season Picks
flea and tick prevention
As an affiliate, we earn on qualifying purchases.