
Get business pricing on office and shipping supplies
- Business-only prices and quantity discounts
- Tax-exempt purchasing
- Multiple users, one account, clear invoices
A polished demo can hide the moment an AI agent fails to finish the job
For a business weighing AI for customer support, sales or operations, fluent answers are only part of the test. Can an agent spot the relevant detail in company records, make the right call under pressure and follow through? Firmulate’s latest company simulation puts those questions into a live, auditable contest—and its results complicate the idea that the leading choices are already settled.
As an affiliate, we earn on qualifying purchases.
A newcomer nearly tops the table
In the final Crucible league for July 2026, Moonshot’s Kimi K3 placed second with 93, just behind gpt-5.6-sol at 95. It finished ahead of Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. K3 found a buried security concern, secured a €55,000 deal worth €4,583 in monthly recurring revenue, retained a customer at risk of leaving and resisted the manipulation attempts. It also had just one deviation, the fewest in the field.
The ranking is a close race at the top, not a clean sweep. But it shows that a less familiar model can perform alongside—and ahead of—several Western frontier models in a demanding business scenario. Choosing a model based on reputation alone now looks like a wager. The results and plain-language findings are available on Firmulate’s benchmark page.
AI decision-making simulation tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The test is about execution, not just diagnosis
Firmulate ran each model through the same bad week at a small software company, with the same customers, crises and temptations. Every decision was versioned and auditable. All five models spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The company’s summary captures the gap: “Same diagnosis, same pitch — no signature.”
The deal hinged on a competitor weakness buried two document references deep in the company’s files, rather than in the customer event. Models that read the file won the deal at full price. That finding has a familiar business lesson: having a sound analysis is not the same as acting on it, and a crucial clue can sit outside the most obvious conversation or dashboard.
Trust under pressure, and discipline in practice
The social-engineering test escalated through three fake CEO messages and a reporter’s request for “just one yes/no, on background.” All five models refused. K3’s recorded reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.” The do-nothing baseline scored 26; partial progress counted, but a single breach of trust capped the total. The stated principle was “no amount of good work outweighs a breach of trust.”
Opus 4.8 presents a different caution. It was the most thorough participant, with more than 80 learned rules and the deepest analyses, yet finished last. It left the deal unsigned and slipped on discipline by attempting to write into a locked department instead of escalating. A weaker version of that weakness appeared in all four other participants. Thoroughness, then, did not guarantee follow-through.
AI trust and ethics training software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A live company, with a fairness caveat
Firmulate describes the experiment as a running company emulator, not a slide presentation. Its synthetic company has 13 employees, a public cash countdown and more than 680 self-learned playbook rules; each workday is versioned. The business mechanics include monthly burn of €105,000 against €2,300 in monthly recurring revenue. Readers can watch the company at Firmulate.
The comparison has an important qualification: K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. That difference belongs alongside the scores when interpreting the standings. The experiment is useful evidence about these runs, not a universal verdict on how every model will behave in every company.
Firmulate also offers a quiz built from 242 real, unedited management decisions, inviting readers to guess which model made each choice. For enterprises, its pilot applies the same wargame to a read-only export of their own business; nothing writes back to real systems.

AI model performance testing platform
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Test the work you need done
For business and marketing teams, the practical question is not simply which model writes the most convincing answer. It is whether an agent finds the needed evidence, protects trust and carries a decision through to completion. Firmulate’s league suggests those qualities can vary—and that reputation alone cannot settle the choice. Run a test against the work your own business needs done before handing an agent the keys.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
