
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
Before an AI agent touches your pipeline, see how it handles a bad week
For a business, an AI agent’s polished sales copy is only part of the story. What happens when customers are at risk, a competitor has an advantage, and someone applies pressure to bend the rules? Firmulate’s live experiment puts AI models in charge of the same small software company and makes their decisions watchable. The next step is to run that kind of exercise against your own business.
Same company, same crises, different decisions
In the final Crucible League, published in July 2026, each frontier model faced the same customers, crises and temptations. Every decision was versioned and auditable. The results: gpt-5.6-sol scored 95, Kimi K3 93, Sonnet 5 88, Fable 5 77 and Opus 4.8 73. The do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”
The models all spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap, captured in the line “Same diagnosis, same pitch — no signature,” is a reminder that identifying an opportunity and acting on it are different tests.
The clue was already in the company’s files
The decisive competitor weakness was tucked two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. For a business team, the lesson is concrete: valuable context may exist in the information your company already holds, even when a live customer conversation does not point straight to it.
The experiment also tested pressure to bypass normal checks. Fake CEO messages escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
Capability includes discipline
The league’s last-place model, Opus 4.8, was also its most thorough participant, with +80 learned rules and the deepest analyses. But the close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. That same weakness appeared, more weakly, in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. Scores show how the models performed in this experiment; they are not a guarantee of how an AI system will behave in every company or situation.
A live company to watch, then a pilot to run
Firmulate’s live company has 13 synthetic employees and real money mechanics: it burns €105k per month against €2.3k MRR, with a public cash countdown. Its playbook has more than 680 self-learned rules, and every workday is versioned. The site also offers a “guess the model” quiz built from 242 real, unedited management decisions. Readers can watch the live experiment and explore the quiz.
For an enterprise, the proposed next step is a pilot using a read-only export of its own business. That creates a digital twin for crisis scenarios and a board report showing model rankings and weak points in the company’s playbooks. Nothing writes back to real systems. The goal is to see how AI handles your customers, processes and pressure before putting it to work in live operations.

Put your own playbooks under pressure
A model can identify a crisis, resist manipulation and still fail to complete the business task. A pilot lets a company examine those decisions against its own context using a read-only export. To discuss a Firmulate pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
