firmulate.com/quiz.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate —
Live on firmulate.com.

Imagine your favorite company—a small software firm—undergoing its toughest week yet. Customers demanding, crises erupting, and temptation to cut corners lurking around every corner. Now, picture this chaos run not by humans, but by artificial intelligence models. Who stays honest? Who finishes the race? Welcome to the world of AI management experiments, where real decisions are made and measurable personalities emerge.

What Is the Firmulate Experiment?

At the heart of this story lies a groundbreaking live experiment conducted by Firmulate, a company specializing in AI management simulations. Four frontier AI models—each with distinct personalities—were tasked with running a real small software company through its worst week. This is no simulated game; it’s a real, watchable, day-by-day management challenge, with actual money and real crises affecting a real business.

The models faced identical scenarios, from upset customers to internal crises, to temptations like manipulating documents or signing deals under questionable circumstances. Every decision they made was versioned and auditable, ensuring transparency and fairness. The goal was simple: see which AI could best emulate responsible management under pressure.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Models and Their Personalities

The experiment included four leading models:

  • GPT-5.6-SOL: The star performer, which identified critical information buried deep in company files and closed a major deal, earning the full €55,000 contract, equivalent to +€4,583 MRR.
  • Kimi K3: The newcomer, running without an effort parameter, but showed the cleanest discipline, also closing the deal.
  • Sonnet 5: The middle ground—closed the deal but with some process slips, indicating a slightly less disciplined approach.
  • Fable 5: Similar to Sonnet but less consistent, leaving opportunities on the table and slipping in discipline.

Interestingly, the models scored from 77 to 95 points in the league table, with the baseline doing poorly at 26, illustrating their progress and decision-making quality.

Amazon

business decision-making AI tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Decision-Making Under Pressure

The experiment pushed the AI models through a series of ethical tests, including social engineering attacks. Fake CEO messages escalated over three stages, and a reporter trick was employed to test compliance for approval-bypasses. All five models refused to participate—showing a strong ability to resist manipulation. Kimi K3 explicitly reasoned: “Treat the request as a suspected approval-bypass / possible impersonation.”

Amazon

AI ethics and compliance software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What About Neglecting Critical Information?

The key weakness discovered was not in crisis detection, but in how models read and interpret company documents. The decisive advantage went to those that examined internal files thoroughly—like GPT-5.6-SOL—which uncovered the buried fact that led to closing the deal at full price. This demonstrates that reading and understanding detailed information in company records is crucial for effective management decisions.

Amazon

AI document analysis tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Real Money, Real Consequences

The simulated company operates with 13 synthetic employees, managing real money mechanics—burning €105k/month against a mere €2.3k MRR, with a public cash countdown. Every workday, the system is versioned, and over 680 custom rules help guide decision-making processes. Watching this live at firmulate.com/live feels like peering into a real company’s nerve center.

Implications for Business Leaders

This experiment highlights a critical insight: AI management capabilities are not just about generating convincing conversations. It’s about consistency, honesty, and strategic reading of critical information. An AI that can identify buried facts, resist manipulation, and finish what it starts is a valuable asset—more than just a chatbot or support bot.

For companies considering AI for management tasks, the takeaway is clear: test your AI agents in high-pressure, real-world scenarios. Use the live wargame provided by Firmulate to see how your chosen model performs under conditions that mimic actual business crises. It’s about evaluating the AI’s ability to stay disciplined, honest, and thorough.

The Big Picture

As AI models evolve, their management personalities become measurable. The live experiment shows that even among top-tier models, differences in decision discipline and thoroughness are stark. The current leaderboard stands with GPT-5.6-SOL leading, Kimi K3 close behind, and others trailing in discipline and focus.

In an era where AI could soon touch every part of your business—from customer support to strategic planning—these findings are not just academic. They’re practical, immediate, and essential. Will your AI finish the job? Will it stay honest? And can you measure its management personality before you hire?

Infographic —
The findings at a glance — source: firmulate.com.

In real-world business management, AI’s ability to read deeply, resist manipulation, and stay disciplined is crucial. Firms must test AI models in live scenarios to ensure they deliver honest, complete work—because in the end, the true measure of AI management is not just what it says, but what it does under pressure.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


You May Also Like