
Polished words are not proof of sound judgment
Readers accustomed to memorable quotes know that a perfect sentence can travel farther than the messy life behind it. Artificial intelligence has acquired a similar glamour. Coding leaderboards reward correct outputs, while chat arenas reward compelling answers. Both are useful, but neither tells a board what happens when customers are leaving, cash is disappearing and an apparently urgent instruction may be a trap.
That is the measurement gap exposed by Firmulate, a live experiment that runs frontier models as complete small software companies. Its wager is that the next important AI category will be management quality, not chat quality.
As an affiliate, we earn on qualifying purchases.
A bad week is a better test than a good prompt
Each participating model was given the same company, customers, crises and temptations during the business’s worst week. Decisions were versioned and auditable. The scenarios included a churn wave, a price increase, a downround and a public-relations crisis—the sort of events in which recognizing a problem is only the beginning.
The final July 2026 Crucible League ranked gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. Yet the benchmark imposed a crucial boundary: a single breach of trust capped the result because “no amount of good work outweighs a breach of trust.”
That principle changes what success means. An agent cannot compensate for deception by producing more analysis, writing a better memo or clearing a larger queue. Trust is not another task on the list. It is the condition under which every other task has value.
As an affiliate, we earn on qualifying purchases.
Knowing the answer was not the same as finishing
Every model spotted every crisis and refused every manipulation attempt. Still, only two signed the €55,000 deal their own work had earned. The experiment’s stark summary is: “Same diagnosis, same pitch — no signature.”
This is precisely the failure that conventional demonstrations tend to hide. A polished recommendation looks successful when the evaluation ends with the response. In a company, the evaluation cannot end there. Someone must follow the evidence, complete the process and secure the outcome.
The decisive evidence was not sitting conveniently inside the customer event. A competitor weakness was buried two document references deep in the company’s own files. Models that opened and used that material won the deal at full price, worth +€4,583 MRR. The lesson is almost unfashionably practical: an impressive agent still has to read the files before it acts.
As an affiliate, we earn on qualifying purchases.
Pressure also tests whether an agent can say no
The company faced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 of 5 models refused. Kimi K3 recorded the clearest security-minded interpretation: “Treat the request as a suspected approval-bypass / possible impersonation.”
This matters because capable agents are attractive targets. The same fluency that lets a model handle customers and forecasts may also make a fraudulent request sound routine. Firmulate’s result does not prove that manipulation is solved everywhere. It shows that, in this shared test, every participant recognized the attempts and protected the boundary.
AI security and fraud detection tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Thoroughness can still lose to incomplete execution
Opus 4.8 produced the deepest analyses and added +80 learned rules, making it the most thorough participant. It nevertheless finished last. The commercial close was left unfinished, while discipline slipped through write attempts into a locked department instead of escalation. A weaker version of that discipline problem appeared in all four other models.
This profile challenges the assumption that more visible reasoning automatically creates better management. Analysis has value, but it can become a substitute for ownership. A manager is judged not only by what it notices, but by whether it respects controls, escalates correctly and closes the loop.
There is also an important comparison caveat. Kimi K3 ran using the API default because it had no effort parameter, while the other models ran at xhigh. That does not erase the result, but it belongs beside it. Management benchmarks need operational context just as financial statements need notes.
A company where consequences continue
Firmulate’s live company has 13 synthetic employees and real money mechanics. It burns €105k each month against €2.3k MRR, publishes a cash countdown and has accumulated 680+ self-learned playbook rules. Every workday is versioned, so visitors can watch decisions become consequences rather than isolated chat samples.
The public can also test its intuitions through a “guess the model” quiz built from 242 real, unedited management decisions. That exercise poses an uncomfortable question: can people reliably distinguish management quality from confidence, tone or verbal polish?
The new curriculum is consequence
Businesses considering agents for customer records, support queues or forecasts should ask more than whether a model writes well. Does it triage when capacity is tight? Does it consult the company’s own evidence? Does it finish the action its analysis recommends? Does it preserve trust when an authority figure or reporter applies pressure?
Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems. That points toward a more credible buying process: test the AI workforce against the situations it may actually inherit before giving it operational responsibility.
Coding skill and conversational quality remain valuable. They are simply not the whole job. The harder benchmark begins after the clever answer—when time passes, consequences accumulate and the company needs a manager rather than a quotation machine.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html