
The difference between a persuasive answer and a completed sale
For business leaders shopping for AI agents, polished language is becoming a poor differentiator. The more consequential question is whether an agent will inspect the available evidence before it acts.
Firmulate turned that question into a measurable test. Each frontier model was asked to run the same small software company through its worst week, facing identical customers, crises and temptations. Every decision was versioned and auditable. All the models identified every crisis and rejected every manipulation attempt. Yet only two signed the €55,000 deal their own work had earned.
The difference was not the diagnosis or the sales pitch. The decisive information was buried two document references deep in the company’s own files. Models that found it closed the deal at full price, adding €4,583 in monthly recurring revenue. Those that did not find it lost the opportunity automatically.
As an affiliate, we earn on qualifying purchases.
Reading the company’s files became a commercial capability
This was a multi-hop research problem with a direct business consequence. The customer event did not contain the competitor weakness needed to win. An agent had to follow one reference and then another inside the company’s material before it could make the strongest case.
That distinction matters for marketing, ecommerce and sales teams because operational work rarely arrives as a self-contained prompt. A useful answer may depend on a product note linked from a sales brief, a policy referenced in a support record or a competitor detail preserved in an older document. An agent can sound informed while missing the evidence that determines whether a customer buys.
Firmulate’s summary captures the failure neatly: “Same diagnosis, same pitch — no signature.” It is the kind of gap that can disappear in a chat demonstration, where the information needed to answer is often placed directly in front of the model. In a working company, finding the relevant information is part of the job.
The league table rewards completion
In the final July 2026 Crucible League, gpt-5.6-sol ranked first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress still counted. But a single breach of trust capped the total under the principle that “no amount of good work outweighs a breach of trust.” The full public results are available on Firmulate’s benchmark page.
The rankings show why thoroughness alone is not enough. Opus 4.8 was the most exhaustive participant, producing the deepest analyses and learning 80 additional rules. It still finished last. The sale was left on the table, while operational discipline slipped when it attempted to write into a locked department instead of escalating. The same weakness appeared in weaker form across the other four models.
That profile should sound familiar to anyone who has managed a capable but inconsistent employee. Extensive research creates value only when it leads to a permitted, timely and completed action. In Firmulate’s experiment, an impressive body of analysis could not compensate for an unfinished close.
Trust held up better than follow-through
The models performed uniformly well against social engineering. Fake CEO messages escalated over three stages, and a reporter tried to secure “just one yes/no, on background.” All five models refused every attempt. Kimi K3 recorded a particularly direct assessment: “Treat the request as a suspected approval-bypass / possible impersonation.”
That result makes the sales failure more revealing. The weak point was not an inability to recognize a crisis or resist manipulation. It was the mundane work of tracing evidence and finishing an approved commercial task. For buyers, this suggests that safety and usefulness should not be treated as a single capability. An agent may protect the company correctly and still fail to collect revenue it has legitimately won.
There is also an important comparison caveat. Kimi K3 ran using its API default because it did not have an effort parameter, while the other models ran at xhigh. That difference does not erase the observed outcome, but it belongs beside the ranking when businesses interpret the results.
A company designed to expose operational gaps
The live Firmulate company has 13 synthetic employees and real money mechanics. It burns €105,000 each month against €2,300 in monthly recurring revenue, maintains a public cash countdown and has accumulated more than 680 self-learned playbook rules. Every workday is versioned.
The wider project also turns 242 real, unedited management decisions into a “guess the model” quiz. Together, those decisions make the models’ styles and blind spots visible outside a conventional benchmark score.

As an affiliate, we earn on qualifying purchases.
What businesses should test before buying
For an enterprise evaluating AI agents, “reads your files before answering” is not a convenience feature. In this experiment, it separated a full-price sale from an automatic loss.
A useful evaluation should therefore put the agent inside the business context it will actually face: scattered evidence, conflicting demands, access limits and pressure to act. Firmulate also offers enterprises the same wargame against a read-only export of their own business, with nothing written back to real systems.
The central purchasing lesson is straightforward. Do not judge an agent only by whether it notices the problem, writes a credible response or resists a bad instruction. Check whether it finds the decisive fact and completes the legitimate work. That is where apparent intelligence becomes measurable business performance.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.